Electronic component and operation control method of electronic component

By dividing the computing unit into a computing domain cluster and monitoring the status parameters in real time, personalized operation control of electronic components is realized, poor timeliness is solved, and the energy efficiency and stability of high-performance computing is improved.

CN120386442AActive Publication Date: 2025-07-29INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510872425.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-07-29
Estimated Expiration
2045-06-26

AI Technical Summary

Technical Problem

Existing electronic components have poor timeliness in operation control, especially in high-performance computing, lack of coordinated optimization of performance, power consumption, temperature and reliability, resulting in high power consumption and heat dissipation problems, making it difficult to adapt to sudden load changes.

Method used

The computing unit is divided into a computing domain cluster. Each computing domain is an independent voltage frequency adjustment domain. The status parameters are monitored in real time through a distributed sensor network. The control components perform personalized adjustments when the threshold conditions are met. The machine learning engine is used to predict load changes and dynamic resource allocation.

Benefits of technology

Improves the timeliness and flexibility of operating control of electronic components, optimizes the collaborative management of performance, power consumption and temperature, ensures equipment stability and extends battery life.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120386442A_ABST
    Figure CN120386442A_ABST
Patent Text Reader

Abstract

The invention discloses an electronic component and an operation control method of the electronic component, and relates to the technical field of computer hardware, the electronic component comprises a control component and a plurality of computing units, and the plurality of computing units are divided into computing domain clusters; in the computing domain cluster, different computing domains comprise different computing units; the control component is used for monitoring a group of state parameters of the activated computational domains under the condition that K activated computational domains exist in the computational domain cluster; in the K activated computational domains, under the condition that the parameter value of the specified state parameter in the group of state parameters meets the corresponding threshold condition, the first computational domain is subjected to operation parameter adjustment according to the adjustment mode corresponding to the specified state parameter, so that the operation parameter adjustment can be performed on the computing unit in different regions; the problem of poor timeliness of operation control of electronic components in related technologies is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer hardware technologies, and particularly to an electronic component and a method for controlling the operation of the electronic component. Background Art

[0002] In order to improve the efficiency of executing compute-intensive tasks, electronic components can be used for task acceleration. The above-mentioned compute-intensive tasks can be tasks in the field of Artificial Intelligence (AI for short). However, high-performance computing is often accompanied by high power consumption and heat dissipation problems, especially in mobile devices and edge computing scenarios. Therefore, it is necessary to adjust the operating parameters of the electronic components.

[0003] In related technologies, the methods for adjusting the operating parameters of electronic components mainly rely on Dynamic Voltage and Frequency Scaling (DVFS for short) and clock gating technologies. However, the electronic components in related technologies are usually based on static strategies or simple feedback control, and there is a problem of poor timeliness in operation control. Summary of the Invention

[0004] This application provides an electronic component and a method for controlling the operation of the electronic component, so as to at least solve the problem of poor timeliness in operation control of the electronic component in related technologies.

[0005] This application provides an electronic component, including a control component and a plurality of computing units. The plurality of computing units are divided into computing domain clusters, and the computing domains in the computing domain clusters are sets of computing units that independently adjust operating parameters; in the computing domain clusters, one computing domain includes some of the plurality of computing units, and different computing domains include different computing units; wherein, the control component is configured to monitor a set of state parameters of the activated computing domains when there are K activated computing domains in the computing domain clusters, where K is a positive integer greater than or equal to 1; in the case of a first computing domain in which the parameter value of a specified state parameter in the set of state parameters exists and satisfies the corresponding threshold condition among the K activated computing domains, adjust the operating parameters of the first computing domain according to the adjustment method corresponding to the specified state parameter.

[0006] The present application also provides a method for controlling the operation of an electronic component. The electronic component includes a plurality of computing units, and the plurality of computing units are divided into computing domain clusters. A computing domain in the computing domain cluster is a set of computing units that independently adjust operation parameters. In the computing domain cluster, one computing domain includes some of the plurality of computing units, and different computing domains include different computing units. The method includes: when there are K activated computing domains in the computing domain cluster, monitoring a set of state parameters of the activated computing domains, where K is a positive integer greater than or equal to 1; when there is a first computing domain in the K activated computing domains where the parameter value of a specified state parameter in the set of state parameters meets the corresponding threshold condition, adjusting the operation parameters of the first computing domain according to the adjustment method corresponding to the specified state parameter.

[0007] Through the present application, since the electronic component includes a plurality of computing units, the plurality of computing units are divided into computing domain clusters, a computing domain in the computing domain cluster includes some of the plurality of computing units, and a computing domain in the computing domain cluster is a set of computing units that independently adjust operation parameters. In the computing domain cluster, one computing domain includes some of the plurality of computing units, and different computing domains include different computing units, it is thus possible to separately monitor and adjust the computing domains in the computing domain cluster. Among them, the control component is used to monitor a set of state parameters of the activated computing domains when there are K activated computing domains in the computing domain cluster, where K is a positive integer greater than or equal to 1; when there is a first computing domain in the K activated computing domains where the parameter value of a specified state parameter in the set of state parameters meets the corresponding threshold condition, adjusting the operation parameters of the first computing domain according to the adjustment method corresponding to the specified state parameter. Therefore, the problem of poor timeliness in the operation control of the electronic component in the related art can be solved, and the technical effects of improving the timeliness and flexibility of the operation control of the electronic component can be achieved. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] In order to more clearly illustrate the embodiments of the present application, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0009] Figure 1 It is a schematic structural diagram of an optional electronic component provided by an embodiment of the present application.

[0010] Figure 2 It is a schematic diagram of an optional electronic component provided by an embodiment of the present application.

[0011] Figure 3 Another optional schematic diagram of an electronic component provided by an embodiment of the present application.

[0012] Figure 4 A schematic flowchart of an optional operation control method for an electronic component provided by an embodiment of the present application. Detailed implementation manners

[0013] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0014] It should be noted that in the description of the present application, the terms "including", "comprising" or any other variants thereof are intended to cover non-exclusive inclusions, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects, rather than to describe a specific order or sequence.

[0015] In order to enable those skilled in the art of the present technology to better understand the solution of the present application, the present application will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.

[0016] According to one aspect of the embodiments of the present application, an electronic component is provided. As Figure 1 shown, the above-mentioned electronic component 101 includes a control component 1011 and a plurality of computing units 1012. The plurality of computing units 1012 are divided into computing domain clusters, and the computing domain 1013 in the computing domain cluster is a set of computing units 1012 that independently adjust operation parameters; in the computing domain cluster, one computing domain 1013 includes some of the plurality of computing units 1012, and different computing domains 1013 include different computing units 1012.

[0017] [[ID=ID=24]]The electronic component in this embodiment can be applied to the field of computer hardware technology and can be applied to scenarios where an electronic component needs to process computing tasks and perform operation control on the electronic component.

[0018] With the rapid development of artificial intelligence technology, electronic components play an increasingly important role in computationally intensive tasks. They are generally used to provide powerful parallel computing performance to accelerate the training and inference processes of AI models while maintaining low latency. They can be designed as hardware components specifically for the inference and training of deep learning models, and their structures usually include multiple dedicated computing units, high-bandwidth memory, high-speed interconnection networks, and complex power and clock management systems. They are suitable for processing various AI tasks, including but not limited to image classification, natural language processing, speech recognition, and recommendation systems. However, high-performance computing is often accompanied by high power consumption and heat dissipation problems, especially in mobile and edge device scenarios, which pose significant challenges. Mobile devices have limited heat dissipation capabilities due to size constraints, and edge computing devices may face more extreme working environments, such as high humidity and temperature. Based on this, it is necessary to control the operation of electronic components, and this operation control not only concerns extending the battery life of the device but also directly affects the performance stability and user experience of the device.

[0019] In these scenarios, operation control needs to find the best balance among performance, power consumption, temperature, and cost, and the related electronic components have limitations in this regard. The operation control of related electronic components often relies on dynamic voltage and frequency scaling and clock gating techniques, lacking the timeliness of operation control.

[0020] Here, dynamic voltage and frequency scaling is a widely used operation control technology. It adjusts the working voltage and frequency in real time to adapt to the current workload. When the load is low, the voltage and frequency are reduced to reduce dynamic power consumption; when the load is high, the voltage and frequency are increased to ensure sufficient computing performance. However, DVFS adjustment is usually global, that is, the voltage and frequency within the entire electronic component are adjusted simultaneously, which results in low efficiency of different internal functional modules when the load is uneven. For example, computationally intensive modules may be limited in performance due to overall frequency reduction, while non-computationally intensive modules may operate at high voltages, wasting energy. In addition, DVFS adjustment relies on feedback control and takes time to respond to load changes, which may lead to performance loss and unnecessary function waste in the case of rapid load changes.

[0021] Here, clock gating technology is another common operation control technology. By pausing the clock signal at times when computation is not required, it reduces the power consumption of the restricted circuit. For example, when the Central Processing Unit (CPU) is in an idle state, some clock signals can be turned off until a new task arrives and then reactivated. However, clock gating usually decides whether to turn on or off the clock based on preset static rules rather than real-time load conditions, lacking timeliness. In addition, frequent clock gating switches can lead to performance delays, especially during task switching, and there may be a performance degradation.

[0022] In addition, the operation control of electronic components in related technologies is also carried out based on heterogeneous computing resource scheduling and Near-threshold Computing (NTC) mode. Here, heterogeneous computing refers to using multiple types of processors to work together in a system to achieve higher computing efficiency and energy efficiency ratio. In electronic components, heterogeneous computing architectures are very common because different computing tasks may be more suitable for the processing methods of different types of processors. However, most current heterogeneous resource scheduling algorithms are based on preset rules or simple load balancing strategies. For example, high-computation-intensive tasks are assigned to the Graphics Processing Unit (GPU), and low-load tasks are assigned to the CPU. Although this technology improves the resource utilization efficiency to a certain extent, in the face of the diversity and uncertainty of AI tasks, static rules may not be able to make timely and optimal resource allocations, and moreover, such scheduling algorithms often ignore the correlations of key factors such as real-time performance, temperature, and power consumption. For example, when the temperature in the GPU area is too high, continuing to assign high-load tasks to it may exacerbate the overheating problem, affecting the system stability and service life.

[0023] Here, near-threshold computing is another operation control technology. Since power consumption is proportional to the square of the voltage, by reducing the operating voltage of the electronic component to an operating point close to the transistor threshold voltage, the power consumption can be significantly reduced. However, when the transistor operates near the threshold voltage, due to factors such as process variations, temperature changes, and environmental noise, timing errors are likely to occur, resulting in inaccurate data processing. To solve this problem, a relatively large voltage margin is often reserved within a safe range, and a more conservative strategy is adopted, such as setting a higher voltage lower limit or frequently performing voltage recovery. This can improve the system reliability but also reduces the effect of energy efficiency optimization.

[0024] It can be seen that the operation control of electronic components in related technologies is usually based on static strategies or simple feedback control, with high adjustment delays and difficulty in adapting to sudden load changes or fast dynamic scenarios. Traditional DVFS and scheduling technologies cannot perform fine-grained adjustments to different modules inside electronic components, and usually only focus on a single dimension (such as performance or temperature). There is a lack of coordinated optimization of performance, power consumption, temperature and reliability, which can easily lead to local bottlenecks or resource waste. There is also a lack of predictive ability for load changes, and it is impossible to adjust voltage and frequency in advance, resulting in adjustment lag and performance loss. Even if the NTC mode is applied, due to the lack of an efficient error detection and recovery mechanism, a large voltage margin needs to be reserved, which limits the room for energy efficiency improvement. That is, the electronic components in related technologies have the problem of poor timeliness of operation control.

[0025] In order to at least partially solve the above technical problems, in this embodiment, an electronic component with a hierarchical adaptive energy efficiency management architecture is proposed. The computing units on the electronic component can be divided into multiple computing domain clusters. The computing domains in the computing domain cluster can include some computing units from multiple computing units. The computing domains in the computing domain cluster can be a set of computing units that independently perform operation control; in the computing domain cluster, a computing domain includes some computing units from multiple computing units, and different computing domains include different computing units, that is, the computing units in the electronic component can be divided into different areas, each of which can be an independent dynamic voltage and frequency adjustment domain, which means that they can independently adjust their operating voltage and frequency to adapt to different load requirements, thereby improving the accuracy and flexibility of the adjustment, and ensuring that different computing modules (such as CPU, GPU) inside the electronic component can achieve the optimal energy efficiency ratio according to their specific computing tasks.

[0026] Optionally, the division of computing units may be pre-set.

[0027] Optionally, each computing domain can also include a distributed sensor network. These sensors are used to monitor the status of the computing domain in real time. Each domain can have 2-4 sensors, or even more sensors. The data sampling rate of these sensors is as high as 1-10kHz, which can capture changes at the microsecond level and provide accurate information for dynamic energy efficiency management.

[0028] Optionally, each computing domain may include one or more computing units. For example, multiple computing units may be included as processing units (PUs), including PU0, PU1, PU2, ... PUm. The above computing units may be CPUs, GPUs, or other types of computing units. Each computing unit may adjust its working mode according to the configuration of the computing domain in which it is located to achieve optimal energy efficiency.

[0029] Here, different computational domains may include different computational units, i.e., a computational unit does not belong to two computational domains simultaneously. Controlling the operation of a computational domain can be to control the operation of the computational units in this area.

[0030] Optionally, the electronic component may include a control component. The control component may include a central arbiter, a central scheduler, etc. The electronic component can be initialized during the power-on startup phase, including: power supply initialization of the electronic component, activation of the central arbiter, loading of the firmware program, etc. Here, the central arbiter can be a module in the electronic component for coordinating and managing the operation control strategy of the electronic component, and it can be responsible for key functions such as task scheduling, performance monitoring, and resource allocation. After confirming the power stability based on the Power_ON signal, the central arbiter can be immediately activated to start the initialization sequence and get ready for the following.

[0031] Optionally, during the system initialization phase, the sensor network can be calibrated. For example, the central arbiter can be used to send Analog-to-Digital Converter (ADC) calibration instructions to the sensor network distributed across the entire electronic component. In response to this instruction, the ADC can perform zero-input reference measurement and gain coefficient adjustment.

[0032] Optionally, during the system initialization phase, key registers (such as a clock divider, an interrupt controller, etc.) can also be initialized. Here, the clock divider is used to adjust the clock frequencies of various parts within the electronic component to ensure that all components can operate at the expected frequencies, and the interrupt controller is used to control various internal and external events and interrupt requests of the electronic component. By initializing the key registers, it can be ensured that the system can respond to important events in a timely manner and efficiently coordinate the execution of various tasks.

[0033] Optionally, after the sensor completes self-calibration, the calibration result can be stored in the local register. This is because the sensor network may be spread across the entire chip. To reduce the latency and power consumption of data transmission, locally storing the calibration result can ensure quick access when needed and immediate application to data processing.

[0034] Optionally, the electronic component can read and load the firmware program from a Read-Only Memory (ROM). The firmware can include the code and data necessary for the startup and initialization of the electronic component.

[0035] In this embodiment, the initialization of the machine learning engine can also be performed to accurately predict the task load and participate in dynamic regulation in subsequent processes. This process can include reading model weights, accelerator configuration, and lightweight self-check. Among them, the reading of model weights can be that the central arbiter reads the pre-trained neural network model weights from a non-volatile memory (Electrically Erasable Programmable Read-Only Memory, abbreviated as EEPROM). The accelerator configuration can refer to loading the model weights into a static random access memory (Static Random-Access Memory, abbreviated as SRAM) and performing address mapping so that the accelerator can efficiently access these data. In addition, hyperparameters such as the time step of the long short-term memory network (Long Short-Term Memory, abbreviated as LSTM) network can also be set to prepare the prediction model for operation. The lightweight self-check refers to verifying whether the key hardware components of the machine learning engine (such as the matrix multiplication and accumulation unit) are working properly.

[0036] Optionally, as an independent voltage and frequency control region, the compute domain can be configured to a default low-power safe state at startup to prepare for upcoming computing tasks.

[0037] Optionally, the central arbiter can also be used to send configuration packets to all compute domains via the power management bus to set the initial low-power state. For example, the voltage is set to 0.6V and the frequency is set to 800MHz, which can not only ensure the fast startup of the acceleration device but also reduce the power consumption during startup.

[0038] Optionally, once all modules are initialized, the system can enter the ready state to prepare for receiving and processing computing tasks.

[0039] In this embodiment, the control component can be used to monitor a set of state parameters of the activated compute domains when there are K activated compute domains in the compute domain cluster, where K is a positive integer greater than or equal to 1. Here, the number of K can be variable, that is, the number of activated and working compute domains in the compute domain cluster can be variable. There can be dormant compute domains being activated, or activated compute domains being frozen into a dormant state.

[0040] Optionally, each computing domain in the computing domain cluster can have a set of status parameters, that is, parameters reflecting the current operating conditions of the electronic component or its computing unit, which can be used for monitoring and diagnosing the status. These parameters can include, but are not limited to, temperature, power consumption, computing load (such as instructions per cycle), cache hit rate, etc. The sensor network distributed on each computing domain in the electronic component can be used to monitor the above parameters. For example, temperature sensors can be used to measure the local temperature of the computing domain, and hardware performance counters can be used to measure the computing power load and other performance metrics of the computing domain.

[0041] Optionally, in the computing domain architecture, a star-shaped power supply network can be adopted, which means that each computing domain can have its own power supply path, from the central power supply point to the voltage regulation unit of each domain, forming a radial network. This design reduces the complexity of the power supply network, supports faster voltage switching, and at the same time, the voltage regulation of each domain can be carried out independently, improving the flexibility of regulation.

[0042] Optionally, decoupling capacitors can be integrated in each computing domain to suppress voltage transient noise and ensure power supply quality.

[0043] Optionally, in the case of the first computing domain among the K activated computing domains where the parameter value of a specified status parameter in a set of status parameters satisfies the corresponding threshold condition, the control component can be used to adjust the operating parameters of the first computing domain according to the adjustment method corresponding to the specified status parameter. Among them, each status parameter can have its specific threshold condition, and the setting of these conditions can be preset depending on factors such as the physical characteristics of the electronic component, the requirements of the current workload, and the system's operation control strategy. For example, the temperature threshold can be set to 85°C to prevent the corresponding component from overheating; while the instructions per cycle threshold may be dynamically adjusted according to changes in process, voltage, and temperature.

[0044] Optionally, operating parameters refer to parameters that can directly control the working state of the electronic component or the computing unit, and their adjustment can directly affect status parameters such as power consumption, performance, and temperature. For example, operating parameters can include voltage and frequency.

[0045] Optionally, when there is a first computing domain among the K activated computing domains and the specified parameter in a set of its status parameters satisfies the preset threshold condition, it indicates that the operating state of this computing domain needs to be intervened. For example, if the temperature of a certain domain exceeds the safety threshold, or its power consumption reaches the predefined upper limit, then the corresponding control strategy will be triggered.

[0046] Optionally, the operation control is usually determined based on specified state parameters that meet threshold conditions. For example, if the temperature is too high, heat dissipation can be achieved by reducing the frequency and voltage of the computing domain, or by migrating part of the load to other computing domains with lower temperatures. If the power consumption is too high, it may be necessary to reduce computing tasks through scheduling or modify the computing mode of the computing unit.

[0047] Once it is determined that the first computing domain needs to be adjusted, the control component can intervene according to the corresponding operation parameter adjustment method. This process can be guided by a Machine Learning (ML) engine or a central arbiter. For example, if the ML engine predicts that the load of a certain computing domain will increase, it can pre-adjust the operation parameters by adjusting the voltage and frequency in advance to avoid sudden increases in performance bottlenecks and thermal stress.

[0048] Through such a monitoring and control mechanism, the computing domain cluster can maintain an optimal operating state under changing workload conditions. On the one hand, it ensures the stability and lifespan of electronic components by preventing overheating and excessive power consumption; on the other hand, through precise resource scheduling and dynamic adjustment, it maximizes the overall performance of the system while ensuring the response speed and accuracy requirements of computing tasks. The operation control under this mechanism is not limited to a single computing domain, but can perform thermal load balancing and dynamic resource allocation across computing domains, forming a synergistically optimized whole.

[0049] Through the embodiments provided in this application, the control component is used to monitor a set of state parameters of the activated computing domains when there are K activated computing domains in the computing domain cluster, where K is a positive integer greater than or equal to 1; in the case of a first computing domain in which the parameter value of the specified state parameter in the set of state parameters satisfies the corresponding threshold condition among the K activated computing domains, the operation parameters of the first computing domain are adjusted according to the adjustment method corresponding to the specified state parameter, which can solve the problem of poor timeliness of operation control in electronic components in related technologies and achieve the technical effects of improving the timeliness and flexibility of operation control of electronic components.

[0050] In an exemplary embodiment, the electronic component further includes at least one of the following:

[0051] A voltage regulation unit configured for each computing domain in the computing domain cluster to regulate the voltage of the corresponding computing domain, where the operation parameters of the computing domains in the computing domain cluster include the voltage of the computing domains in the computing domain cluster, and the voltage regulation unit includes at least one of the following: a low dropout regulator, a switched capacitor converter;

[0052] A local clock tree synthesis component, an asynchronous bridge, and digital phase-locked loops respectively configured for computing domains in a computing domain cluster, wherein the local clock tree synthesis component is used to distribute the global clock signal of an electronic component to different computing domains in the computing domain cluster to form clock domains corresponding to different computing domains in the computing domain cluster, the asynchronous bridge establishes a data transmission path between different clock domains, and the digital phase-locked loop is used to adjust the frequency of the corresponding computing domain. The operating parameters of the computing domains in the computing domain cluster include the frequencies of the computing domains in the computing domain cluster.

[0053] Optionally, each region in the computing domain cluster is a region that can perform independent voltage / frequency control, that is, each computing domain can have its own power management and clock frequency to meet the energy efficiency requirements under different computing tasks or load conditions. In the computing domain cluster, the operating parameters of each computing domain include but are not limited to its operating voltage and frequency. By dynamically adjusting these parameters, the computing efficiency and energy efficiency can be adjusted. For example, under light load conditions, the power consumption can be reduced by lowering the voltage; under heavy load conditions, the computing performance can be improved by increasing the voltage and frequency.

[0054] Optionally, the electronic component may include voltage regulation units respectively configured for the computing domains in the computing domain cluster. Here, the voltage regulation unit is a hardware component responsible for adjusting the voltage of the computing domain. The voltage regulation unit may include a Low Dropout Regulator (LDO for short), which can be used to provide a stable operating voltage for the computing domain, stably convert the input voltage (such as the global 1.8V) into a voltage level suitable for the load of the computing domain (0.4 - 1.2V) to meet the power supply requirements of different computing tasks while reducing power consumption and heat generation.

[0055] Optionally, the voltage regulation unit may also include a switched-capacitor converter. Here, the switched-capacitor converter is a circuit that uses the combination of capacitors and switches to achieve voltage conversion. It can adjust the voltage without an external power supply but by using the charge and discharge cycle of the internal power supply of the electronic component.

[0056] Optionally, the electronic component may further include a local clock tree component for distributing the global clock signal of the electronic component to different computing domains in the computing domain cluster, forming clock domains corresponding to different computing domains in the computing domain cluster, which can be dynamically adjusted according to the task load requirements without affecting the operation of other computing domains. Each computing domain may be integrated with a Digital Phase Locked Loop (DPLL) for generating and adjusting the operating frequency, enabling the computing units within the domain to operate at different frequencies and being able to dynamically adjust the frequency according to the load within the domain. Exemplarily, the frequency adjustment range may be stepped from 10 MHz to 100 MHz, providing flexible clock management capabilities for the computing domain.

[0057] Optionally, to ensure cross-domain data synchronization, the electronic component may further include an asynchronous bridge, such as a dual-clock domain First-In-First-Out (FIFO) buffer. This kind of bridge can perform distortion-free data transmission between different clock frequencies. The depth of the FIFO is at least 8, thereby ensuring data integrity and timing correctness, and maintaining data synchronization and correctness even when exchanging data between domains with different frequencies.

[0058] Through this embodiment, by means of the local clock tree synthesis component, asynchronous bridge in the electronic component, and the digital phase-locked loops and voltage regulation units respectively configured for the computing domains in the computing domain cluster, the voltage and frequency of each computing domain can be independently adjusted, improving the flexibility and reliability of the operation control of the electronic component.

[0059] In an exemplary embodiment, the control component is connected to the computing units within the computing domains in the computing domain cluster through a Network-on-Chip (NoC).

[0060] Here, the Network-on-Chip (NoC) is a high-speed, low-latency communication network that can be used inside large-scale integrated circuits. It can replace the traditional bus architecture and provide a more flexible and reliable data transmission mechanism. The NoC may include multiple routing nodes and links, capable of supporting parallel communication between multiple computing units.

[0061] Optionally, the topology of the NoC may be a regular grid, ring, tree, or a more complex hybrid topology. The choice of topology may depend on the number, location of the computing domains, and data transmission requirements. For example, a grid topology can be used for a large number of uniformly distributed computing units, and a tree topology can be used for a computing domain cluster with a hierarchical structure.

[0062] Optionally, through the NoC, the control component can send control signals to the computing domain and each computing unit within the computing domain to achieve operation controls such as dynamic task allocation and voltage / frequency adjustment. For example, in the case of sudden high load, the control component can send commands to reduce the frequency and voltage to the computing units in the computing domain cluster through the NoC to reduce power consumption and temperature.

[0063] Through this embodiment, the control component and the computing units within the computing domain in the computing domain cluster are connected through the on-chip interconnection network, which can improve the efficiency and flexibility of data transmission.

[0064] In an exemplary embodiment, a set of state parameters includes the instructions per cycle, and the threshold condition corresponding to the instructions per cycle is that the value of the instructions per cycle is less than or equal to the first instruction number threshold;

[0065] The control component is further configured to adjust the operation parameters of the first computing domain according to the adjustment method corresponding to the specified state parameter, including:

[0066] When the specified state parameter includes the instructions per cycle, perform a first adjustment operation on the first computing domain to adjust the operation parameters of the first computing domain, where the first adjustment operation includes: lowering the frequency of the first computing domain and lowering the voltage of the first computing domain.

[0067] Here, the instructions per cycle (InstructIons Per Cycle, abbreviated as IPC) refers to the number of instructions that can be executed within one clock cycle, which can be the instructions per cycle of the computing domain or the instructions per cycle of the computing unit. In this embodiment, a set of state parameters can include the instructions per cycle. A high IPC value means that the computing domain can efficiently process more instructions, thereby achieving higher computing performance. Generally, higher frequencies and voltages can support more instruction executions, but at the same time, they also bring higher power consumption and heat generation.

[0068] Optionally, a set of state parameters can include multiple metrics including IPC to comprehensively monitor and evaluate the operating state of the electronic component. In this set of parameters, the threshold condition of IPC can be set as the first instruction number threshold. When the value of IPC continuously falls below this threshold, it indicates that the computing load of the electronic component is low, and the current frequency and voltage settings may exceed the actual requirements, resulting in unnecessary power consumption waste.

[0069] Optionally, when the specified status parameter includes the number of instructions per cycle, the control component can be used to perform a first adjustment operation on the first computing domain to adjust the operating parameters of the first computing domain. Here, the first adjustment operation can include: lowering the frequency of the first computing domain and lowering the voltage of the first computing domain. Here, by reducing the clock frequency of the computing domain, the number of instructions executed in each clock cycle can be reduced, thereby reducing power consumption. In addition, a lower voltage also means lower power consumption.

[0070] Optionally, in the operation control strategy, the performance of the computing domain can be measured in the form of TOPS / W (tera operations per second per watt), which can reflect the computing ability of electronic components under unit power consumption. The computing domain scoring formula can be performance R / (power consumption × temperature), so as to balance performance and power consumption and consider the influence of temperature. In this computing domain scoring, although a high IPC value represents high performance, if the power consumption and temperature are also high, the advantage of high IPC will be weakened. Therefore, the goal of operation control can be to minimize power consumption and temperature by adjusting frequency and voltage on the premise of meeting performance requirements, so as to optimize the performance of the entire system.

[0071] In this embodiment, by dynamically monitoring the IPC and comparing it with a preset threshold, the low-load state can be intelligently identified, and the frequency and voltage can be lowered in a timely manner, avoiding energy waste when the computing demand is low and improving the overall energy efficiency of the system.

[0072] Optionally, both the setting of the IPC threshold and the adjustment strategy of the frequency / voltage can be dynamic and can be flexibly adjusted according to different application environments and computing loads. This operation control strategy is not only applicable to bursty computing tasks but also suitable for high-load scenarios with long-term operation, ensuring that the system can maintain the best working state under various working conditions.

[0073] Through this embodiment, by monitoring the number of instructions per cycle and performing corresponding operating parameter adjustments when the number of instructions per cycle is lower than the corresponding threshold, the timeliness and reliability of the operation control of electronic components can be improved.

[0074] In an exemplary embodiment, the electronic component further includes: a hardware performance counter for real-time monitoring of the number of instructions per cycle of K activated computing domains.

[0075] Here, the hardware performance counter (Hardware Performance Counter, abbreviated as HPC) can be used to collect and statistically analyze the performance index data of each unit in the electronic component in real time. The unit can be a computing unit or a computing domain.

[0076] Optionally, hardware performance counters can be used to monitor, including but not limited to, instructions per cycle, cache hit rate, branch hit rate, instruction type mix, data flow pattern, etc.

[0077] In this embodiment, the hardware performance counters can be used to continuously monitor the IPC values of K activated computing domains, where K refers to the number of computing domains that are currently active, i.e., the number of activated computing domains. By monitoring the IPC of these domains in real time, it is possible to quickly identify which computing domains are executing instructions efficiently and which computing domains have relatively light loads. Furthermore, when the IPC of a certain computing domain continuously drops below a preset threshold, the operating frequency and voltage of this computing domain can be reduced.

[0078] Similarly, if the IPC value of a certain computing domain is high, it indicates that this computing domain is in a high-pressure working state, which may require additional cooling or voltage adjustment to prevent overheating and excessive power consumption.

[0079] Through this embodiment, by using the hardware performance counters to monitor the instructions per cycle of the activated computing domains in real time, the immediate working status of each computing domain inside the electronic component can be obtained in real time, so as to be able to quickly respond to changes and implement corresponding operation control measures, improving the timeliness of the operation control of the electronic component.

[0080] In an exemplary embodiment, a set of status parameters includes the computing domain temperature, and the threshold condition corresponding to the computing domain temperature is that the temperature value of the computing domain temperature is greater than or equal to the first temperature threshold;

[0081] The control component is further configured to, when the specified status parameter includes the computing domain temperature, perform a second adjustment operation on the first computing domain to adjust the operating parameters of the first computing domain, where the second adjustment operation includes at least one of the following: lowering the frequency of the first computing domain and performing voltage compensation on the second computing domain, where the second computing domain is an activated computing domain that is adjacent to the first computing domain and does not meet the threshold condition corresponding to the computing domain temperature; migrating the task executed by the first computing domain to an activated computing domain that does not meet the threshold condition corresponding to the computing domain temperature.

[0082] In this embodiment, the control component can be used to perform adjustment operations on the computing domains based on the monitoring results of the computing domain temperature. Here, the computing domain temperature can be used to evaluate the thermal status of each computing domain (i.e., different computing regions of the electronic component). Once the temperature of a certain computing domain reaches or exceeds the predefined temperature threshold, it means that this computing domain has entered an overheated state, which not only affects its computing performance but may also affect the stability and lifespan of the hardware.

[0083] Optionally, the first temperature threshold can be preset or modified as needed.

[0084] Optionally, when it is detected that the temperature of the first computational domain has reached or exceeded a preset temperature threshold, a corresponding second adjustment operation can be triggered to optimize the overall energy efficiency and thermal distribution, and avoid performance degradation or hardware damage caused by excessive temperature.

[0085] Optionally, the second adjustment operation may include: reducing the frequency of the first computational domain and performing voltage compensation on the second computational domain. Here, reducing the frequency can directly achieve the effects of reducing power consumption and heat generation. Through this adjustment operation, the power consumption of the first computational domain can be quickly reduced, its temperature can be lowered, and overheating can be avoided. After reducing the frequency and voltage of the first computational domain, in order to maintain the performance stability of the overall system, voltage compensation can be performed on the second computational domain that is adjacent to the first computational domain and whose temperature has not reached the overheat threshold. This is because due to the electrical connection between computational domains, the frequency reduction operation may indirectly affect its neighboring computational domains. Specifically, the frequency reduction of the first computational domain results in a decrease in its power consumption and current demand, which may cause a change in the current distribution of the overall power network, and further increase the voltage drop (abbreviated as IR Drop) effect in the area where the neighboring computational domain is located. Here, IR Drop refers to the problem that the power supply voltage decreases at various points inside the electronic component due to changes in the resistance and current inside the electronic component. To counteract this effect, while reducing the frequency of the first computational domain, the voltage of the second computational domain can be moderately increased. By increasing the supply voltage of the second computational domain, the local voltage drop of the power network caused by the frequency reduction of the first computational domain can be offset, thereby maintaining the stable operation and high-performance output of the second computational domain.

[0086] Optionally, the second computational domain may only include activated computational domains.

[0087] Optionally, multiple high-temperature computational domains with temperatures exceeding the preset temperature threshold, that is, multiple hotspots, can be identified. For each identified hotspot, the temperature conditions around it can be further analyzed, and adjacent points within a certain distance from the hotspot and with temperatures also higher than the threshold can be merged into the same cluster to form an adjustment area.

[0088] Optionally, the second adjustment operation may further include: migrating some computational tasks on the first computational domain to other computational domains with lower temperatures, and further controlling the heat source through reallocation of tasks to achieve thermal equilibrium.

[0089] Through this embodiment, by taking corresponding adjustment operations when detecting overheating of the computational domain, the reliability and stability of the operation control of the electronic component can be improved.

[0090] In an exemplary embodiment, the control component is also used to lower the frequency of the first calculation domain by a specified frequency value, wherein the specified frequency value is the product of the first coefficient and the specified temperature difference value, and the specified temperature difference value is the difference between the temperature value of the calculation domain temperature of the first calculation domain and the first temperature value; and to increase the voltage of the second calculation domain by a specified voltage value, wherein the specified voltage value is the product of the second coefficient, the specified difference value and the specified voltage compensation value, and the specified difference value is the difference obtained by subtracting the ratio of the temperature value of the calculation domain temperature of the second calculation domain to the temperature value of the calculation domain temperature of the first calculation domain from 1.

[0091] In this embodiment, a high-temperature computing domain (i.e., a computing domain whose temperature exceeds a first temperature threshold) may be subjected to a gradient frequency reduction, as shown in formula (1):

[0092] Δf = -k×(T_now - T_max) (1)

[0093] Wherein, Δf refers to the specified frequency value to be lowered, k refers to the first coefficient, T_now refers to the temperature value of the calculation domain temperature of the first calculation domain, and T_max refers to the first temperature value, that is, the first temperature threshold.

[0094] Optionally, the specified frequency value refers to a dynamic value calculated based on the first coefficient and the specified temperature difference. The first coefficient can be a value predetermined by experiment, which is used to reflect the strength of the relationship between temperature and frequency adjustment. The larger the first coefficient, the more obvious the impact of temperature change on frequency adjustment, and vice versa. The specified temperature difference refers to the difference between the actual temperature of the first calculation domain and the first temperature threshold. That is, the more the calculation domain temperature exceeds the threshold, the greater the frequency reduction amplitude, and the amplitude of the frequency reduction is proportional to the degree of temperature exceeding the limit, thereby ensuring that the response speed of thermal control matches the degree of temperature increase, avoiding the risk of overheating.

[0095] Optionally, while lowering the frequency, the supply voltage of the computational domain can be adjusted according to a dynamic voltage frequency scaling (DVFS) table to ensure stable circuit operation at different frequencies while further reducing power consumption. The DVFS table can include recommended voltages for different frequencies. For example, when the frequency drops from 2.0 GHz to 1.5 GHz, the supply voltage can also be reduced from 1.0 V to 0.8 V. This voltage-following strategy ensures that the voltage adjustment matches the frequency change during frequency reduction, maintaining efficient and reliable circuit operation.

[0096] Optionally, when performing cold zone compensation, that is, performing voltage compensation on the second calculation domain, voltage compensation may be performed based on formula (2):

[0097] ΔV_cold =α× (1 - (T_cold / T_hot)) × ΔV_max (2)

[0098] Where ΔV_cold is the specified voltage value, α is the second coefficient, T_cold is the temperature value of the second calculation domain, T_hot is the temperature value of the first calculation domain, (1 - (T_cold / T_hot)) is the specified difference, and ΔV_max is the specified voltage compensation value.

[0099] Here, the second coefficient can be a value predetermined through experiments, which can be used to reflect the correlation between temperature changes and voltage compensation effects, that is, the degree of influence of temperature differences on the voltage compensation amount. The specified difference is defined as 1 minus the difference between the temperature of the second calculation domain and the temperature of the first calculation domain. It can be used to represent the difference between the hot and cold temperatures of the second calculation domain relative to the first calculation domain. The larger the temperature difference, the greater the need for compensation. The specified voltage compensation value can be an upper limit value used to limit the compensation amount to avoid new problems caused by excessive voltage.

[0100] Through the above-mentioned anti-gradient regulation algorithm, it can be achieved that when the temperature difference between the cold zone and the high temperature domain is large, the cold zone voltage compensation will be more significant to offset the voltage loss caused by the increase in hot spot current in the power supply network and ensure the normal operation of the cold zone circuit.

[0101] Through this embodiment, by performing gradient frequency reduction on the high-temperature calculation domain and compensating the voltage of the adjacent cold zone, the local heat source can be finely adjusted. Through dynamic adjustment and precise control, efficient control is achieved under complex and changing load and environmental conditions, ensuring performance optimization while reducing energy waste.

[0102] In an exemplary embodiment, the electronic component also includes a distributed sensor network; wherein, the sensor network is used to detect the temperature of the activated computing domain separately when a set of state parameters includes the computing domain temperature, and obtain sensor data corresponding to the activated computing domain; the control component is also used to parse the sensor data corresponding to the activated computing domain, and obtain the actual temperature value of the activated computing domain at the current moment; based on the actual temperature value of the activated computing domain at the current moment, the current power value of the activated computing domain and the ambient temperature value at the current moment, the temperature value of the activated computing domain at the next moment of the current moment is predicted, wherein the current power value of the activated computing domain is the power value of the activated computing domain at the current moment, and the temperature value of the computing domain temperature of the activated computing domain is the predicted temperature value of the activated computing domain at the next moment.

[0103] In related technologies, electronic components can integrate temperature, voltage, and current sensors for real-time monitoring of component status. However, this sensor data is usually only used for simple threshold alarms or global regulation and is not deeply integrated with dynamic operation regulation strategies.

[0104] In this embodiment, the electronic component further includes a distributed sensor network. Among them, sensors can be distributed in each computing domain for real-time monitoring of temperature changes within the computing domain. Correspondingly, the sensor network can detect the temperature in each activated computing domain respectively to obtain sensor data corresponding to the activated computing domain.

[0105] The control component can be used to analyze the sensor data corresponding to the activated computing domain to obtain the actual temperature value of the activated computing domain at the current moment. Among them, the sensor data corresponding to the activated computing domain is the sensor data obtained by the sensor network detecting the activated computing domain respectively. That is, by analyzing the data sent by the sensor network, it can be analyzed and decoded to obtain the accurate temperature value of each activated computing domain at the current time point.

[0106] Optionally, in addition to detecting the temperature within the computing domain, the current ambient temperature can also be considered. This is because the thermal condition of the modules in the electronic component is not only affected by internal computing activities, but the ambient temperature is also an important factor. Incorporating the ambient temperature helps to establish a more accurate temperature model, especially for edge computing devices whose working environment may be more variable.

[0107] Optionally, based on real-time temperature, power, and ambient data, the control component can predict the temperature value of the activated computing domain at the next time point and use it as the computing domain temperature of the activated computing domain, so as to take control operations such as reducing the frequency in advance and transferring the load before the temperature actually exceeds the threshold, thereby avoiding performance degradation and hardware damage caused by overheating.

[0108] Through this embodiment, the actual temperature and predicted temperature of the computing domain can be obtained through the sensor data obtained based on the distributed sensor network, which can improve the forward-looking and flexibility of the operation control of the electronic component.

[0109] In an exemplary embodiment, the control component is further used to determine the sum result of the product of the actual temperature value of the activated computing domain at the current moment and the thermal inertia coefficient, the product of the current power value of the activated computing domain and the power-temperature conversion coefficient, and the product of the ambient temperature value at the current moment and the specified heat exchange rate as the predicted temperature value of the activated computing domain at the next moment, where the specified heat exchange rate is the heat exchange rate between the electronic component and the environment where the electronic component is located.

[0110] In this embodiment, the control component can be used to determine the predicted temperature value of the activated computing domain at the next moment as the sum of the product of the actual temperature value of the activated computing domain at the current moment and the thermal inertia coefficient, the product of the current power value of the activated computing domain and the power-temperature conversion coefficient, and the product of the ambient temperature value at the current moment and the specified heat exchange rate, as shown in Equation (3):

[0111] T_new = A×T_prev + B×Power_map + C×Ambient (3)

[0112] Among them, T_new represents the new temperature state, T_prev is the temperature at the previous moment, Power_map represents the power distribution of each part in the electronic component, and Ambient is the ambient temperature. The coefficient matrices A, B, and C respectively reflect the thermal inertia coefficient, power-temperature conversion coefficient of the electronic component, and heat exchange rate between the environment and the electronic component (i.e., the environmental coupling coefficient).

[0113] Among them, the thermal inertia coefficient A can be used to represent the change speed of the electronic component in the face of temperature changes. The larger the A coefficient, the slower the response of the electronic component to temperature changes; the power-temperature conversion coefficient B can be used to reflect the temperature rise degree caused by unit power consumption per unit time; the environmental coupling coefficient C can be used to measure the ability of the electronic component to exchange heat with the surrounding environment, that is, the influence intensity of the ambient temperature on the internal temperature of the electronic component.

[0114] Optionally, the temperature, power, and ambient temperature of the computing domain can be monitored in real time, and then these three parameters are respectively multiplied by the corresponding coefficient matrices A, B, and C, and finally the results are summed to obtain the predicted temperature at the next moment. This process can be updated every 100 μs, ensuring the real-time and accuracy of temperature prediction.

[0115] Optionally, in the later stage of the electronic component design (i.e., post-silicon testing), the temperature distribution can be recorded using an infrared thermal imager under different power consumption conditions by heating the resistor, and the three coefficient matrices A, B, and C can be finely calibrated to ensure the accuracy and effectiveness of the prediction model.

[0116] Through this embodiment, temperature prediction can be performed by combining the thermal prediction model with the actual temperature, power value, ambient temperature, and thermal coefficients of the electronic component, which can improve the real-time and accuracy of temperature prediction, and further improve the accuracy of operation control.

[0117] In an exemplary embodiment, the sensor network includes a set of sensing nodes arranged in a grid layout, and one computing domain in the computing domain cluster corresponds to some of the sensing nodes in the set of sensing nodes;

[0118] The control component is also used to obtain the sensor data detected by the sensing nodes corresponding to K activated computing domains in the form of address polling, and obtain the sensor data corresponding to the K activated computing domains.

[0119] Optionally, the sensor network can adopt a uniform grid layout, with 1 sensing node integrated per square millimeter. One computing domain in the computing domain cluster can correspond to some of the sensing nodes in the sensing node set. The sensor can be a high-precision sensor (e.g., with a thermal gradient detection accuracy of ±0.1°C / mm).

[0120] Optionally, the above-mentioned sensing nodes can include: a PN junction diode temperature sensor for temperature detection, with a nominal accuracy within ±0.5°C; a micro Hall sensor for current detection, with a measurement range covering 0 - 100 mA and a resolution of 1 mA, so as to accurately measure the current flowing through each computing domain; a 12-bit successive approximation register (SAR) ADC integrated on the chip for converting the analog signals of temperature and current into digital signals. The sampling rate of the ADC can be 1 kHz (i.e., sampling 1000 times per second) to ensure the real-time nature of the data, and an effective number of bits (ENOB) ≥ 10 bits means that the actual effective number of bits of the ADC under the influence of noise is high enough and the data is reliable.

[0121] Optionally, the control component can interact with the sensor network in the form of address polling. In the address polling, the control component can send data requests to each sensing node in the sensor network in a predefined order, requesting them to return the current monitoring data.

[0122] Optionally, the communication between the control component and the sensor network generally follows a certain standardized protocol, such as the Inter-Integrated Circuit (I2C).

[0123] Optionally, the control component can send requests to each sensing node at a fixed time interval (such as 100 μs) and wait for a response. This time interval can ensure that the above requests can cover all the sensing nodes of the activated computing domains without affecting the overall system performance.

[0124] Optionally, after receiving the sensor data, the control component can perform preliminary parsing and preprocessing on the data for formulating the dynamic adjustment strategy.

[0125] Optionally, the control component can send data requests only to the sensing nodes corresponding to the activated computing domains. After receiving the requests, the sensing nodes corresponding to the activated computing domains can return data such as the current temperature, voltage, and current to the control component via the I2C bus.

[0126] Through this embodiment, by the control component periodically polling the sensing nodes corresponding to the activated computing domains, the real-time performance and reliability of data collection can be improved.

[0127] In an exemplary embodiment, the electronic component further includes a temperature sensor; wherein, the temperature sensor is used for periodically detecting the ambient temperature;

[0128] The temperature sensor or the control component is further used for calculating the temperature difference between the detected ambient temperature value and the recorded reference ambient temperature value; in the case where the temperature difference between the detected ambient temperature value and the recorded reference ambient temperature value is greater than or equal to the temperature difference threshold, the product of the detected ambient temperature value and a specified coefficient plus a second temperature value is determined as the updated first temperature threshold, and the recorded reference ambient temperature value is updated to the detected ambient temperature value.

[0129] In this embodiment, the temperature sensor can be configured to periodically (e.g., every minute or when there is a major system state change) detect the ambient temperature, so as to obtain the real-time temperature change in the operating environment of the electronic component.

[0130] Optionally, after receiving the latest ambient temperature value from the temperature sensor, it can be compared with the pre-recorded reference ambient temperature value, and the difference between the two temperature values is calculated. Here, the reference ambient temperature value can be the ambient temperature at the previous detection or the initial ambient temperature recorded during system initialization. This calculation process can be performed by the control component or autonomously by the temperature sensor.

[0131] Similar to the foregoing embodiments, the temperature threshold can be dynamically updated and adjusted. Here, a temperature difference threshold can be set to determine whether the current ambient temperature change is sufficient to affect the operation control strategy of the electronic component. If the difference between the detected ambient temperature value and the reference value is greater than or equal to this threshold (e.g., ΔT≥5°C), it can be considered that the environmental conditions have changed significantly and the temperature threshold in the electronic component needs to be adjusted. Specifically, the new first temperature threshold (the temperature upper limit for triggering the adjustment operation) can be defined as the product of the detected ambient temperature value and a specified coefficient plus a fixed second temperature value. Here, the specified coefficient reflects the sensitivity of the ambient temperature to the internal temperature of the electronic component, and the second temperature value can represent the temperature offset when the electronic component operates normally under ideal cooling conditions, as shown in formula (4):

[0132] T_max = Base value + β × Ambient temperature (4)

[0133] Wherein, T_max is the first temperature threshold, and β is the specified coefficient.

[0134] Optionally, after the temperature threshold is adjusted, the newly detected ambient temperature value can also replace the original reference ambient temperature value as the comparison benchmark for ambient temperature detection in the next cycle.

[0135] By adjusting the temperature threshold, its performance output can be maximized under different environmental conditions while maintaining within a safe energy consumption range. For example, in a low-temperature environment, electronic components may be able to withstand a higher internal temperature.

[0136] Through this embodiment, by dynamically updating the first temperature threshold according to the ambient temperature, the accuracy and timeliness of temperature judgment can be improved, thereby improving the accuracy and timeliness of the operation control of electronic components.

[0137] In an exemplary embodiment, the control component is further configured to, in the case where there is a first computing task to be allocated, select a third computing domain for executing the first computing task from the K activated computing domains according to the task label of the first computing task and the operation characteristic data of the K activated computing domains, wherein the task label of the first computing task is used to indicate the task attributes of the first computing task and the execution constraint conditions of the first computing task; allocate the first computing task to the third computing domain for the first computing unit of the third computing domain to execute the first computing task.

[0138] In this embodiment, when there is a first computing task to be allocated, that is, when a new computing task enters the central task scheduler queue in the control component and waits for resource allocation, the most suitable computing domain can be selected based on the task requirements and the running status of the currently activated computing domains to execute the task, so as to achieve the maximization of efficiency and the optimization of performance.

[0139] Optionally, the task requirements can be determined based on the task label of the computing task. Here, the task label is a part of the task descriptor (Task Descriptor, abbreviated as TD), which can include task attributes and execution constraint conditions. For example, the TD of an AI inference request may include max_power_allowed (maximum allowed power consumption), required precision (such as single-precision floating point, half-precision floating point, or 8-bit integer), and task type (such as image classification, natural language processing, etc.).

[0140] Optionally, the central scheduler can parse the task label of the first computing task to understand the task type, precision requirements, and execution constraint conditions.

[0141] Similar to the foregoing embodiments, the operating state parameters of all the activated computing domains on the electronic component can be monitored in real time, including their current temperature and load status, etc. Furthermore, the central scheduler can calculate the adaptability score of each computing domain according to the operating characteristic data of the computing domain. For example, the computing domain score (performance² / power consumption × temperature), the load balancing situation, and whether the current power consumption of the computing domain meets the execution constraint conditions of the task can be calculated.

[0142] Optionally, the central scheduler can continuously monitor the real-time Process-Voltage-Temperature (PVT) data of the electronic component. These data reflect the current operating environment and status of the electronic component and can be used to evaluate the available computing resources on the electronic component.

[0143] Optionally, after evaluating all K activated computing domains, the computing domain with the highest score and best meeting the requirements of the first computing task can be selected as the third computing domain. If the scores of multiple domains are the same, other priorities (such as temperature, power consumption, or availability of computing resources) can be used for selection.

[0144] Optionally, once the third computing domain is determined, the central scheduler can allocate the first computing task to the first computing unit of this computing domain. Subsequently, the first computing unit executes the task while following the execution constraint conditions specified in the task label and can continuously monitor its operating state to ensure the consistency of energy efficiency and performance.

[0145] When a task arrives, the central scheduler adds the new computing task to the task queue and starts resource allocation according to the task label and real-time PVT data. This step is the starting point of the dynamic operation control strategy, where the task label and precision requirements (single-precision floating-point / half-precision floating-point / 8-bit integer) play a guiding role to ensure that the task is allocated to the computing unit most suitable for its needs. For example, a high-load, low-precision image detection task may be directed to a Neural Processing Unit (NPU) in 8-bit integer mode because the NPU can provide a high energy efficiency ratio in 8-bit integer mode and meet the precision requirements of the task at the same time.

[0146] Through this embodiment, by allocating tasks based on the task label of the task and the operating characteristic data of the activated computing domains when the task arrives, reasonable allocation of computing resources can be achieved, avoiding resource waste and excessive power consumption, and improving the flexibility of the system.

[0147] In an exemplary embodiment, a task level prediction engine runs on a control component. The task level prediction engine is used to predict the task level of a computing task corresponding to an input task label based on the input task label and input historical performance data. Among them,

[0148] The control component is further configured to input the task label of the first computing task and the historical performance data of the activated computing domain into the task level prediction engine, and obtain the task level of the first computing task predicted by the task level prediction engine. The task level of the first computing task is used to characterize the computing power requirement for executing the first computing task. From the K activated computing domains, select an activated computing domain that matches the task level of the first computing task to obtain the third computing domain.

[0149] Here, the task level prediction engine refers to a machine learning component integrated in the central scheduler, which can be used to predict the task level of a computing task corresponding to an input task label based on the input task label and input historical performance data. For example, the prediction engine can analyze the execution conditions of past tasks similar to the attributes of the first computing task (such as accuracy requirements, power consumption constraints, etc.), including their load patterns, power consumption characteristics, execution time, and temperature changes. Through these historical data, the prediction engine can build a model for predicting the behavior pattern and computing power requirement of new tasks on different computing domains.

[0150] Optionally, the control component can parse the task descriptor of the first computing task to understand its attributes and execution constraint conditions, and use the TD information of the task and the historical performance data of the matching activated computing domain as inputs and pass them to the task level prediction engine. The task level prediction engine can predict the computing power requirement level of the first computing task based on the input data. The prediction result can also reflect the expected performance of the task on different computing domains and can be used to guide resource allocation.

[0151] Optionally, after completing the task level prediction, the control component can select a computing domain that best matches the predicted task level from the K activated computing domains as the third computing domain. The selection criteria may include the current state of the computing domain, historical performance, and the computing power requirement predicted by the task level prediction. Finally, the scheduler assigns the first computing task to the third computing domain and adjusts the operating parameters (such as voltage and frequency) of the computing domain according to the predicted performance requirements and execution constraints of the task to optimize the efficiency and energy efficiency ratio of task execution.

[0152] Through this embodiment, predicting the computing power requirement of a task through the task level prediction engine can improve the accuracy of task allocation and reduce resource waste.

[0153] In an exemplary embodiment, one of the multiple computing units allows a computing task to be executed in at least one precision mode. The higher the computing precision corresponding to the precision mode, the more floating-point bits are used to execute the computing task; the task attribute of the first computing task includes the precision requirement of the first computing task, and the precision requirement of the first computing task is used to indicate the lowest precision mode required to execute the first computing task.

[0154] The control component is further configured to, when there are at least two candidate computing domains in the K activated computing domains whose allowed precision modes match the precision requirement of the first computing task, calculate the computing domain score of the candidate computing domains according to the performance, power consumption, and temperature of the candidate computing domains, where the computing domain score of the candidate computing domain is positively correlated with the performance of the candidate computing domain and negatively correlated with the power consumption and temperature of the candidate computing domain, and the performance of the candidate computing domain is the number of instructions per cycle of the candidate computing domain.

[0155] Determine the candidate computing domain with the highest computing domain score among the at least two candidate computing domains as the third computing domain.

[0156] In modern heterogeneous computing systems, such as AI electronic components, computing units that can support multiple precision modes are widely adopted to meet the requirements of different computing tasks. Among them, the selection of precision modes has a direct impact on power consumption, performance, and operating status. Although high-precision computing (such as single-precision floating-point) provides higher accuracy, it also consumes relatively more resources (including power consumption and computing time). Lower-precision computing (such as 8-bit integers), while sacrificing some accuracy, achieves a higher performance ratio and is especially suitable for scenarios with high real-time requirements but slightly larger tolerance for accuracy, such as inference.

[0157] In this embodiment, each computing unit can support one or more precision modes, such as single-precision floating-point mode, half-precision floating-point mode, and 8-bit integer mode. The higher the precision, the more floating-point bits are used when executing a computing task, and the more accurate the computing process is, but the corresponding computing resource consumption and power consumption are also greater.

[0158] Specifically, the following precision modes may exist:

[0159] Single-precision floating-point mode (2'b00): Adopts 32-bit floating-point calculation and is suitable for scenarios with extremely high precision requirements, such as scientific computing tasks like model training.

[0160] Half-precision floating-point mode (2'b01): Uses 16-bit floating-point calculation and can achieve a better balance between precision and energy efficiency, and is often used in mixed-precision training.

[0161] 8-bit integer mode (2'b11): Uses 8-bit fixed-point calculation, with the highest computing performance, mainly applied to inference scenarios such as image classification and object detection.

[0162] Taking the reconfigurable multiplier as an example, different calculation methods can be selected according to the value of the precision_mode function:

[0163] When precision_mode is 2'b00, the fp32_mult function is called for 32-bit floating-point multiplication calculation.

[0164] When precision_mode is 2'b01, first perform 16-bit floating-point multiplication fp16_mult on the lower 16 bits of the input, and then convert the result to 32-bit floating-point.

[0165] When precision_mode is 2'b11, perform 8-bit fixed-point multiplication int8_mult on the lower 8 bits of the input, and pad zeros in the high bits to extend to 32 bits.

[0166] To further optimize the area of electronic components, the reconfigurable multiplier can reuse some hardware resources among the three modes, such as the mantissa processing unit. This means that when switching between different precision modes, the multiplier can process data with different bit widths through the same circuit block, and only needs to be adjusted through internal configuration to adapt to different computing requirements. In addition, when switching modes, the system latency can be less than 10 clock cycles. This short latency is mainly due to the update of hardware configuration and ensuring that there is no residual data of the previous mode in the computing pipeline. Therefore, the pipeline needs to be cleared before and after switching to avoid data confusion.

[0167] Exemplarily, different precision modes of the reconfigurable multiplier correspond to different power consumption levels and applicable scenarios:

[0168] Single-precision floating-point mode: Provides the highest precision and is suitable for scientific computing.

[0169] Half-precision floating-point mode: Provides medium precision and is suitable for mixed-precision training scenarios.

[0170] 8-bit integer mode: Provides the lowest precision, but the highest energy efficiency ratio, and is suitable for inference tasks such as image classification and object detection.

[0171] Optionally, as part of the task descriptor, the precision requirement of the first computing task can indicate the lowest precision mode required to execute the task. For example, an image recognition task may only require 8-bit integer precision, while a scientific computing task may require single-precision floating-point precision.

[0172] Optionally, the control component is further configured to, when there is only one candidate computing domain among the K activated computing domains whose allowed precision mode matches the precision requirement of the first computing task, determine this computing domain as the third computing domain; when there are at least two candidate computing domains among the K activated computing domains whose allowed precision mode matches the precision requirement of the first computing task, the computing domain scores of the candidate computing domains can be calculated based on the performance, power consumption, and temperature of the candidate computing domains.

[0173] Optionally, for each candidate computing domain, its computing domain score can be calculated. This score can comprehensively consider the performance of the computing domain (such as instructions per cycle), power consumption, and temperature. The computing domain temperature here can be the actual temperature of the computing domain or the predicted temperature of the computing domain at the next moment. This is not limited in this embodiment.

[0174] Optionally, the scoring formula can reflect the positive correlation between performance and the computing domain score, as well as the negative correlation between power consumption and temperature and the computing domain score, as shown in formula (5):

[0175] Computing domain score = (Performance^2) / (Power consumption * Temperature) (5)

[0176] Wherein, the performance is measured by instructions per cycle (IPC). IPC can reflect the computing efficiency of the computing domain, while the power consumption and temperature represent the energy consumption and thermal state of the computing domain.

[0177] Optionally, among at least two candidate computing domains, the candidate computing domain with the highest computing domain score can be determined as the third computing domain, so as to ensure that the task can be allocated to the computing resources with the optimal energy efficiency ratio while meeting the precision requirement, thereby achieving the balance between performance and energy efficiency.

[0178] Through this embodiment, by selecting the computing domain of the task based on the precision requirement and the computing domain score, the balance between performance and energy efficiency can be achieved, the utilization rate of resources can be improved, and the system performance can be optimized.

[0179] In an exemplary embodiment, one of the multiple computing units is allowed to execute a computing task in at least one precision mode. The higher the computing precision corresponding to the precision mode, the more floating-point digits are used to execute the computing task; the task attribute of the first computing task includes the precision requirement of the first computing task, and the precision requirement of the first computing task is used to indicate the lowest precision mode required to execute the first computing task; the third computing domain includes at least two types of computing units;

[0180] A third computing domain is used to screen computing units from the computing units in the third computing domain according to preset screening conditions. When at least two candidate computing units are screened out, according to the performance of the candidate computing units, the power consumption of the candidate computing units, and the temperature of the candidate computing units, calculate the computing unit score of the candidate computing units. Among at least two candidate computing units, when the precision mode allowed for the candidate computing unit with the highest computing unit score matches the precision requirement of the first computing task, determine the candidate computing unit with the highest computing unit score as the first computing unit; among at least two candidate computing units, when the precision mode allowed for the candidate computing unit with the highest computing unit score does not match the precision requirement of the first computing task, determine the specified computing unit in the third computing domain as the first computing unit; wherein, the preset screening conditions include one of the following: when the temperature of the electronic component is greater than or equal to the second temperature threshold, screen the computing units with a temperature lower than the third temperature threshold; screen the computing units that meet the execution constraint conditions; screen the computing units whose allowed precision mode matches the precision requirement of the first computing task.

[0181] Wherein, the computing unit score of the candidate computing unit is positively correlated with the performance of the candidate computing unit, and negatively correlated with the power consumption of the candidate computing unit and the temperature of the candidate computing unit. The performance of the candidate computing unit is the number of instructions per cycle of the candidate computing unit.

[0182] Similar to the foregoing embodiments, one of multiple computing units is allowed to execute a computing task in at least one precision mode. The higher the computing precision corresponding to the precision mode, the more floating-point bits are used to execute the computing task; the task attribute of the first computing task includes the precision requirement of the first computing task, and the precision requirement of the first computing task is used to indicate the lowest precision mode required to execute the first computing task; the third computing domain includes at least two types of computing units, that is, each computing unit supports executing a computing task in at least one precision mode. The high-precision mode (such as single-precision floating-point) uses more floating-point bits and provides higher mathematical precision, which is suitable for application scenarios with extremely high accuracy requirements. The low-precision mode (such as 8-bit integer) uses fewer floating-point bits. Although the precision is lower, it can significantly improve the energy efficiency ratio and reduce power consumption in scenarios such as inference.

[0183] Optionally, when selecting a computing unit for executing the first computing task, computing units can be screened out from the third computing domain according to preset screening conditions. The preset screening conditions can include:

[0184] 1. When the temperature of the electronic component is greater than or equal to the second temperature threshold, calculate units with a temperature lower than the third temperature threshold are screened. Here, both the second temperature threshold and the third temperature threshold are preset and are used to automatically exclude high-power consumption computing cores when the overall temperature of the electronic component is too high to ensure the stability of the electronic component.

[0185] 2. Calculate units that meet the execution constraint conditions are screened out, that is, calculate units that perform well in terms of power consumption and temperature, to optimize the overall system energy efficiency.

[0186] 3. Calculate units whose allowed precision modes match the precision requirements of the first computing task are screened out to ensure that the calculate units can meet the precision requirements of the task.

[0187] After at least two candidate calculate units are screened out, the calculate unit score of each candidate calculate unit can be calculated. The calculation formula can be as shown in formula (5). Among them, the calculate unit score of the candidate calculate unit is positively correlated with the performance of the candidate calculate unit, and negatively correlated with the power consumption and temperature of the candidate calculate unit. The performance of the candidate calculate unit is the number of instructions per cycle of the candidate calculate unit.

[0188] Optionally, each calculate unit can provide real-time power consumption data through a built-in current sensor, and at the same time use distributed thermistors to collect temperature data at a sampling rate of 1 kHz to provide a basis for scheduling decisions.

[0189] Optionally, under the condition of meeting the precision requirements, the candidate calculate unit with the highest calculate unit score will be determined as the first calculate unit.

[0190] Optionally, if the calculate unit with the highest calculate unit score does not match the precision requirements of the first computing task, then a calculate unit preset in the third computing domain and matching the precision requirements can be selected as the first calculate unit for executing the first computing task.

[0191] Through this embodiment, the dynamic screening and selection strategy of calculate units based on precision requirements and execution constraints can dynamically find a balance between meeting the task precision requirements and optimizing the overall energy efficiency, ensuring that when processing diverse computing tasks, it can continuously provide a high-performance and low-power consumption operating state and optimize the system performance.

[0192] In an exemplary embodiment, one of the multiple calculate units is allowed to execute a computing task using at least one precision mode. The higher the computing precision corresponding to the precision mode, the more floating-point bits are used to execute the computing task;

[0193] The fourth computational domain among the K activated computational domains is used to extract the to-be-executed second computational task from the computational task queue of the fourth computational domain. The second computational task has at least two computational stages, and the precision requirements of the second computational task are different in different computational stages among the at least two computational stages. The precision requirement of the second computational task is used to indicate the lowest precision mode required to execute the second computational task. In different computational stages of the second computational task, at least one computational unit in the fourth computational domain executes the second computational task in different precision modes, so that the precision mode adopted by the computational unit in the fourth computational domain matches the precision requirement of the computational stage of the second computational task.

[0194] In the fields of high-performance computing and AI acceleration, computational tasks often contain multiple stages, and the computational precision requirements may be different in each stage. The diversity of such requirements requires the ability to flexibly adjust the precision mode of its computational units to adapt to the requirements of the task in different stages and achieve the best energy efficiency ratio at the same time.

[0195] In this embodiment, each computational unit, such as a CPU, GPU, etc., can execute computational tasks in different precision modes.

[0196] Optionally, the fourth computational domain among the K activated computational domains can be used to extract the to-be-executed second computational task from the computational task queue of the fourth computational domain. The second computational task has at least two computational stages, and the precision requirements of the second computational task are different in different computational stages among the at least two computational stages. The precision requirement of the second computational task is used to indicate the lowest precision mode required to execute the second computational task. For example, in an image processing task, the preprocessing stage may only require relatively low-precision calculations (such as 8-bit integers) for image scaling and cropping; while in the feature extraction or classification stage, higher-precision calculations (such as single-precision floating-point) may be required to ensure the accuracy and robustness of the model.

[0197] Optionally, the precision requirement of each computational stage can be used to indicate the lowest precision mode required to execute that stage.

[0198] Optionally, in each different computational stage of the second computational task, at least one computational unit in the fourth computational domain executes the second computational task in different precision modes, and can execute the task in a precision mode that matches the precision requirement of that stage. Thus, the precision mode can be dynamically adjusted in different computational stages, reducing unnecessary power consumption and computational resource consumption while ensuring the accuracy of the computational results in each stage.

[0199] For example, taking the continuous face detection of a mobile AI camera as an example, the computing mode can be dynamically switched according to the requirements of different stages: in the initial stage, since high precision is required in this stage, the system uses the GPU to run a half-precision floating-point model to ensure the accuracy of face positioning. At this time, the computing unit score is 15 TOPS / W; in the stable stage, that is, when the detection results of 10 consecutive frames are consistent, it indicates that low-precision computing can also meet the requirements in the current scenario, and it can be automatically switched to the NPU to run an 8-bit integer quantization model, and the computing unit score is increased to 42 TOPS / W; when an exception handling is triggered, such as when the detection confidence drops by more than 5%, it can immediately fallback to the half-precision floating-point mode and trigger a recalculation of the accuracy-power efficiency weight to ensure the accuracy of the detection.

[0200] Through this embodiment, by matching the most suitable precision mode for different computing stages, the computing efficiency and accuracy of each stage can be ensured, avoiding performance bottlenecks caused by insufficient precision or resource waste, and achieving efficient utilization of resources and maximization of energy efficiency.

[0201] In an exemplary embodiment, the activated computing domain is configured with multiple load levels, and a load level prediction engine runs on the control component. The load level prediction engine is used to predict the load level of the computing domain corresponding to the input prediction reference data in M future cycles based on the input prediction reference data. The prediction reference data includes the following data of the computing domains in the computing domain cluster in N historical cycles: instruction mix, cache miss rate, and task queue depth of the computing task queue, where M is a positive integer greater than or equal to 1, and N is a positive integer greater than or equal to 1; among them,

[0202] The control component is further configured to obtain the prediction reference data of the activated computing domain, input the prediction reference data of the activated computing domain into the load level prediction engine, obtain the load level of the activated computing domain predicted by the load level prediction engine in M future cycles, and pre-adjust the voltage and frequency of the activated computing domain based on the predicted load level of the activated computing domain in M future cycles.

[0203] In this embodiment, the control component can obtain the operation metrics of the activated computing domain in the most recent N cycles as prediction reference data. These metrics can include: instruction mix, which is used to reflect the proportion and combination of different types of instructions (such as integer, floating-point, memory access) to evaluate the computing requirements and load characteristics of the computing domain; cache miss rate, which is used to indicate the efficiency of cache access. A higher cache miss rate means frequent data exchange, and a higher frequency may be required to compensate for the resulting latency; the task queue depth of the computing task queue, which is used to represent the current state of the task queue. The larger the queue depth, the more tasks to be processed, and the higher the load of the computing domain may be.

[0204] Optionally, the selection of the value of N can depend on the characteristics of the task and the system's response requirements for the historical period. Generally, N is a positive integer greater than or equal to 1, which means that the reference data covers at least one complete cycle to capture the dynamic changes in recent operations.

[0205] Optionally, the control component can be used to input the collected prediction reference data into the load level prediction engine. This engine can be based on machine learning models such as recurrent neural networks, long short-term memory networks, or other time series analysis models, which are not limited in this embodiment. It can be used to predict the load level of the computing domain in the next M cycles. M is also a positive integer greater than or equal to 1, representing the prediction time window.

[0206] Optionally, the load level prediction engine can be a quantized 2-layer LSTM network running on a dedicated tensor accelerator. The quantized LSTM network can convert data types such as weights and activation functions from floating-point numbers to fixed-point numbers to reduce the computational complexity and memory occupancy while maintaining the network's prediction ability.

[0207] Optionally, the number of cycles in the future that the load level prediction engine can predict can be determined in advance according to experiments.

[0208] Optionally, according to the prediction results output by the load level prediction engine, the voltage and frequency of the activated computing domain can be pre-adjusted. For example, the voltage or frequency can be pre-adjusted 50 - 100 clock cycles in advance. Through pre-adjustment, the voltage and frequency can be increased before the actual load increases to prevent performance degradation. Correspondingly, when it is predicted that the load will decrease, the voltage and frequency can also be reduced in advance to save power consumption.

[0209] Similar to the previous embodiments, traditional DVFS responds after the load changes, with a certain delay. The prediction-based pre-adjustment can significantly reduce this delay and improve the overall response speed and stability of the system.

[0210] Optionally, based on the prediction results, the voltage can be adjusted through the Power Management Unit (PMU for short). It can follow the predefined correspondence table of load level and voltage frequency (as shown in Table 1) to ensure that the computing requirements are met while minimizing power consumption.

[0211] Table 1

[0212]

[0213] In addition to voltage regulation, the frequency can also be finely tuned through a digital phase-locked loop to adapt to the predicted load level and ensure that the computing unit operates in the most suitable state.

[0214] Through this embodiment, by pre-adjusting the voltage and frequency of the computing domain based on the load level prediction engine, the hysteresis and resource allocation problems in dynamic voltage and frequency adjustment can be solved, and the timeliness and flexibility of the operation control of electronic components are improved.

[0215] In an exemplary embodiment, the control component is further configured to control the activated computing domain to enter the near-threshold computing mode when the number of instructions per cycle of the activated computing domain is less than or equal to the second instruction number threshold and the temperature of the activated computing domain is less than or equal to the fourth temperature threshold; in the near-threshold computing mode, based on the operating state data of the activated computing domain, adjust the voltage offset of the activated computing domain, where the operating state data of the activated computing domain is used to indicate the temperature of the activated computing domain, the voltage of the activated computing domain, and the process deviation coefficient of the activated computing domain, and the voltage offset of the activated computing domain is obtained by summing the product of the difference between the temperature of the activated computing domain and the specified temperature and the third coefficient, the product of the difference between the voltage of the activated computing domain and the specified voltage and the fourth coefficient, and the product of the process deviation coefficient and the fifth coefficient.

[0216] Similar to the foregoing embodiment, in the hierarchical adaptive energy efficiency management architecture, a near-threshold computing mode (i.e., near-threshold voltage mode) can be introduced. In this mode, the operating voltage of the computing unit is very close to the threshold voltage of the transistor. Therefore, the voltage needs to be carefully controlled to avoid timing errors caused by factors such as process deviation and temperature fluctuations.

[0217] Optionally, the number of instructions per cycle and the temperature of the activated computing domain can be continuously monitored. When the IPC is less than or equal to the third instruction number threshold (i.e., the computing load of the current computing domain is low and there is enough room to reduce the operating voltage without affecting performance) and the temperature is less than or equal to the third temperature threshold (ensuring that the computing domain operates in a cooled state and reducing the risk of timing errors caused by temperature rise), the computing domain can be controlled to enter the near-threshold computing mode to further reduce power consumption.

[0218] Optionally, in the near-threshold computing mode, the operating voltage of the computing domain can be fine-tuned according to the real-time operating state data of the computing domain to achieve the best energy efficiency ratio while ensuring the reliability of the calculation. Here, the operating state data can mainly include: temperature, voltage, and process deviation coefficient.

[0219] Optionally, a voltage offset (delta_V) can be calculated based on the above parameters to guide the voltage fine-tuning. The calculation formula of the voltage offset is shown in formula (6):

[0220] (6)

[0221] Among them, k1, k2, and k3 are the third coefficient, the fourth coefficient, and the fifth coefficient respectively, representing the influence degrees of temperature, voltage, and process deviation on the voltage offset amount.

[0222] Optionally, by dynamically adjusting the voltage of the computing domain according to the calculated voltage offset amount, a suitable operating point can be found on the boundary between safety and energy efficiency.

[0223] Optionally, after adjustment, the timing state of the computing domain can be re-detected to verify the stability of the system. If no timing violation or error is detected within the next N cycles (N is usually set to 10 or more), the system state is considered stable and the task can continue to be executed under the current voltage and computing mode. If errors are continuously detected, further voltage adjustment can be triggered or the system can be reverted to a safer operating mode (such as the operating mode under standard voltage) to avoid data corruption or computing failure.

[0224] Through this embodiment, by introducing the near-threshold computing mode and performing fine voltage adjustment in this mode, extremely low-power operation can be achieved while ensuring the correct execution of computing tasks, reducing resource consumption.

[0225] In an exemplary embodiment, the electronic component further includes: a master latch and a shadow latch, where the triggering edges of the master latch and the shadow latch are different; among them,

[0226] The master latch and the shadow latch are used to sample data of the activated computing domain when the activated computing domain enters the near-threshold computing mode.

[0227] The control component is further used to sample data of the activated computing domain through the master latch and the shadow latch, and detect timing errors of the activated computing domain, where the triggering edges of the master latch and the shadow latch are different; in the case of detecting a timing error in the activated computing domain, continuously increase the voltage of the activated computing domain until the error instruction is successfully executed, where the error instruction is the instruction executed by the activated computing domain corresponding to the timing error in the activated computing domain; or, control the activated computing domain to exit the near-threshold computing mode.

[0228] In this embodiment, in the near-threshold computing mode, real-time monitoring of data sampling performed on the computing domain can be achieved through the comparison mechanism of the master latch and the shadow latch. Here, the triggering edges of the master latch and the shadow latch are different. For example, the master latch (positive-edge triggered) and the shadow latch (negative-edge triggered) can be triggered by the rising edge and the falling edge of the clock pulse respectively within the same clock cycle to sample the data signal. Under normal circumstances, since they sample the same signal simultaneously, the output results should be the same. However, if after the sampling is completed, the output results of the master latch and the shadow latch are compared and the difference exceeds a pre-set tolerance (e.g., ±3%), it can be considered that a timing violation has occurred, that is, Error = 1, indicating that the timing consistency of data sampling has been disrupted.

[0229] Optionally, the Razor flip-flop can detect the inconsistency between the outputs of the master latch and the shadow latch to monitor and judge the occurrence of timing errors in real time.

[0230] Optionally, once a timing error is detected, actions can be taken to prevent the further propagation of the error. For example, the pipeline can be immediately stalled to freeze the execution of the current and subsequent instructions, preventing the results of incorrect instructions from being passed to the next stage of the computing process and avoiding the spread of errors.

[0231] Optionally, during the pipeline stall, the execution of subsequent instructions will be suspended until it is confirmed that the error has been corrected.

[0232] Optionally, it is possible to roll back to the last correct state, that is, use the correct instruction copy pre-stored in the buffer and re-execute to ensure the continuity and correctness of the calculation.

[0233] In this embodiment, the required voltage compensation amount (delta_V) can be calculated based on the current PVT parameters to restore the timing consistency of the transistor. This adjustment process is shown in Equation (7):

[0234] (7)

[0235] where k1, k2, and k3 are compensation coefficients related to temperature, voltage, and process deviation coefficients respectively.

[0236] Through the above dynamic voltage regulation process, the voltage can be gradually adjusted to a stable operating point to avoid the recurrence of timing errors. For example, the voltage can be gradually increased in steps of 10 mV / cycle until the incorrect instruction can be successfully executed.

[0237] Optionally, if consecutive voltage regulation attempts fail to stabilize, or timing errors are frequently detected within a period of time, it can be determined that the risk of operating in NTC mode is too high. At this time, the voltage of the activated computing domain can be controlled to rise to a safer level, and the near-threshold computing mode can be exited and returned to the normal operating mode to ensure the normal operation of the computing platform under all conditions.

[0238] In this embodiment, by combining a hardware-accelerated load prediction engine with Razor trigger protection, sub-millisecond-level response can be achieved. This architecture can achieve high-precision energy efficiency regulation capabilities through a hardware-accelerated key control loop (prediction engine and error recovery mechanism) and a hierarchical power supply design. Predictive regulation can reduce the latency penalty compared to traditional DVFS, and ensure reliability in NTC mode through a detection circuit, balancing the recovery speed and energy efficiency with the voltage regulation step size.

[0239] Through this embodiment, by detecting timing errors in real time in the near-threshold computing mode and performing error recovery through a dynamic voltage regulation strategy, when facing an unrecoverable situation, the near-threshold computing mode is exited in a timely manner and switched to a safer operating state, effectively avoiding the risks of data corruption and computing failure.

[0240] As an optional embodiment, different computing domains in the computing domain cluster are configured with different activation levels. The higher the activation level, the higher the computing power of the computing domain. The highest activation level among the activation levels of the computing domains in the computing domain cluster is the first activation level.

[0241] The control component is further configured to continuously monitor the following information of K activated computing domains: the number of instructions per cycle, and the task queue depth of the computing task queue, where the highest activation level among the activation levels of the K activated computing domains is the second activation level; when the second activation level is not the first activation level, when the number of instructions per cycle of the activated computing domain at the second activation level is greater than a third instruction number threshold and the task queue depth of the computing task queue of the activated computing domain at the second activation level is greater than a specified depth threshold, activate the computing domain at the third activation level in the computing domain cluster, where the third activation level is the next activation level after the second activation level.

[0242] In this embodiment, different computing domains in the computing domain cluster can be configured with different activation levels. The higher the activation level, the higher the computing power of the computing domain. The highest activation level among the activation levels of the computing domains in the computing domain cluster is the first activation level.

[0243] Exemplarily, the computing domain corresponding to the lowest activation level can be set as the zeroth activation level domain, that is, the L0 domain. The L0 domain can be in a normally active state, always powered, and responsible for basic control, sensor data acquisition and simple tasks (such as background synchronization); the computing domains corresponding to the secondary activation level and higher activation levels can be set as L1, L2, L3...Ln domains, respectively, where the L1-Ln domains can be turned off and can be sorted according to the system topology.

[0244] For example, Figure 2 As shown, the computing units on the electronic components can be divided into computing domain clusters, and the computing domains in the computing domain clusters can be configured as different activation levels, for example, the zeroth activation level domain, the second activation level domain, the third activation level domain, etc. The above computing domains can all be connected to the control components (including input / output interfaces, double data rate memory controllers, general processors, and central arbiters) through the on-chip interconnection network.

[0245] Optionally, the activation level of the computing domain refers to the level of computing power that can be provided when the computing domain is activated. Generally, the higher the activation level, the stronger the computing power of the computing domain, but the corresponding power consumption and heat may also increase accordingly.

[0246] Optionally, in the initial state, only the L0 domain and necessary input / output (I / O) domains may be active, and other domains are power-gated (ie, the supply voltage is 0 and the clock is turned off).

[0247] Optionally, if the L0 domain cannot meet the load and computing power requirements of the task, primary activation can be performed to activate the L1 (corresponding to the second activation level) domain. At this time, the central arbitrator can send an activation signal to the power controller of the L1 domain. The power controller can release the power gating (for example, turn on the LDO, and the voltage rises from 0V to 0.7V, which takes about 200ns) and release the clock gating (for example, the DPLL is locked to 1.0GHz, which takes 150ns). When the L1 domain has been activated, the control component can migrate the basic tasks to the L1 domain.

[0248] For example, the activation order of computed fields can be as follows Figure 3 As shown, when the trigger condition is met, the first activation level domain (i.e., L1 domain) is activated, and the load is continuously monitored to determine whether it meets the demand. If it is still insufficient, the next level of computing domain is activated, and this process is repeated until the load is balanced.

[0249] Optionally, performance indicators of the K activated computing domains may be monitored in real time, including the number of instructions per cycle and the task queue depth of the computing task queue. These indicators reflect the computing efficiency and task load of the computing domains.

[0250] Optionally, when the second activation level is not the first activation level (i.e., there is a computing domain where the highest level has not been activated yet), it is possible to determine whether the activated computing domains at the second activation level meet the following two conditions based on the performance metrics of the K activated computing domains monitored in real time:

[0251] 1. The number of instructions per cycle is greater than the third instruction number threshold, indicating that the computing resources of the computing domain are approaching saturation and may require more computing power support.

[0252] 2. The task queue depth of the computing task queue is greater than the specified depth threshold, meaning that there are a large number of tasks waiting to be processed in the task queue and the current computing power is not sufficient to respond quickly.

[0253] Optionally, when the above conditions are met, it is possible to further activate the computing domains at the third activation level in the computing domain cluster. Here, the third activation level is the next activation level after the second activation level, i.e., a computing domain with stronger computing power.

[0254] Optionally, if the load of the computing domain drops below the preset threshold, it is also possible to initiate a downgrading process to reallocate the computing tasks to a computing domain with lower power consumption to save energy.

[0255] Similarly, when it is necessary to activate a higher-level computing domain (such as the L2 or L3 domain), the PMU can also quickly adjust the voltage and frequency to a predetermined value, such as setting the L2 domain to 0.9V / 2.0GHz.

[0256] Optionally, it is possible to pre-define the voltage and frequency settings corresponding to different load levels through a look-up table method to minimize the power consumption overhead during the adjustment process. For example, the table defined here can be the same as Table 1 in the foregoing embodiment.

[0257] Optionally, when the computing task is migrated, to ensure the continuity of the task, it is possible to make the computing domain load the complete computing context before receiving the task, including but not limited to the current state of the task, the processed data, and the cache information, so as to continue execution from the interruption point of the previous processing unit.

[0258] Optionally, when the IPC of a certain computing domain is continuously lower than the preset threshold (for example, 0.3 indicates light load or idle) for more than N cycles (for example, 1000 cycles), it is possible to trigger the sleep process of this computing domain. At this time, the control component can migrate the unfinished tasks on this computing domain to a computing domain with a low load level, then freeze the clock signal, save the state of the critical registers on the computing domain to the reserved registers for quickly restoring the state during subsequent activation, and apply power gating to reduce the voltage of the computing domain to 0V to enter the deep sleep state, thereby reducing the power consumption of the system.

[0259] Through this embodiment, through the computing domain dynamic expansion strategy based on the activation level, the computing domain can be dynamically activated, achieving effective utilization of resources while meeting the computing requirements, and improving the flexibility and adaptability of the system.

[0260] As an optional embodiment, the computing tasks executed by the electronic component are configured with corresponding task priorities, and the task priorities of the computing tasks executed by the electronic component correspond to the activation levels of the computing domains in the computing domain cluster;

[0261] The control component is further configured to, in the case that there is a third computing task to be allocated, and when the activation levels of the K activated computing domains are all lower than the activation level corresponding to the task priority of the third computing task, activate the fifth computing domain in the computing domain cluster whose activation level corresponds to the task priority of the third computing task, and allocate the third computing task to the fifth computing domain.

[0262] In this embodiment, tasks can be allocated according to the corresponding relationship between the task priority and the activation level of the computing domain. Here, each computing task can be configured with a specific task priority for the urgency and importance of the task. For example, real-time response, critical business processing, or user interaction, etc. can belong to high-priority scenarios, while background updates, data preprocessing, etc. can belong to low priorities.

[0263] Similar to the previous embodiment, the activation level of the computing domain is proportional to the computing power it can provide. The higher the activation level, the higher the voltage and frequency settings of the computing domain, thus providing stronger computing power, but also consuming more energy at the same time.

[0264] Optionally, the corresponding relationship between the task priority and the computing domain activation level can be predefined. Furthermore, when a high-priority task arrives, it can be compared with the activation levels of the currently activated computing domains to ensure that the task is executed on a computing domain with a sufficient computing power level.

[0265] For example, there is a third computing task to be allocated, and its task priority requires a computing domain with a higher level of computing power than the currently K activated computing domains (whose activation levels are all lower than the activation level corresponding to the task priority of the third computing task) to execute. In this case, the computing domain with the fifth activation level in the computing domain cluster can be automatically activated, and this level exactly matches the task priority of the third computing task. That is to say, the fifth computing domain can provide sufficient computing resources to meet the requirements of the third computing task. Subsequently, the third computing task will be directly allocated to the fifth computing domain for execution, skipping the steps of layer-by-layer activation level matching to ensure that the task can respond quickly and be completed efficiently.

[0266] In this embodiment, for high-priority tasks, the regular activation level order can be skipped, and the highest-level computing domain can be directly activated, which can ensure that critical tasks can be immediately responded to, avoiding delays caused by untimely preparation of computing resources.

[0267] Optionally, when multiple tasks simultaneously request to use the computing resources of the same computing domain, the priority of the tasks can be determined according to the ratio of the performance requirements of the tasks to the power consumption of the current domain.

[0268] Optionally, in a complex computing system, dynamic allocation and activation of resources may encounter delay problems. Such as activation signal transmission, power management unit (PMU) response time, clock tree stabilization, etc., may all increase the activation delay. A long activation delay may lead to resource allocation deadlocks, that is, multiple tasks or modules wait for each other to release resources, and thus cannot execute normally. To prevent the occurrence of deadlock situations, a maximum activation delay constraint (such as 100 μs) can be preset, and all domains are forced to be activated and an alarm is given when the time limit is exceeded.

[0269] Through this embodiment, by activating the computing domain based on the corresponding mechanism between the task priority and the computing domain activation level, timely scheduling and response of computing resources can be achieved, improving the stability and reliability of the system when dealing with sudden computing demands.

[0270] The embodiment of the present application also provides a method for controlling the operation of an electronic component. This method can be implemented by the hardware in the above embodiment and the preferred implementation manner, and those that have been described will not be repeated. The electronic component includes a plurality of computing units, and the plurality of computing units are divided into a computing domain cluster. The computing domains in the computing domain cluster are sets of computing units that independently adjust operation parameters; in the computing domain cluster, one computing domain includes some of the plurality of computing units, and different computing domains include different computing units; Figure 4 is a schematic flowchart of an optional method for controlling the operation of an electronic component according to an embodiment of the present application, as Figure 4 shown, this method includes:

[0271] Step S402, when there are K activated computing domains in the computing domain cluster, monitor a set of state parameters of the activated computing domains, where K is a positive integer greater than or equal to 1;

[0272] Step S404, when there is a first computing domain in the K activated computing domains where the parameter value of a specified state parameter in a set of state parameters meets the corresponding threshold condition, adjust the operation parameters of the first computing domain according to the adjustment method corresponding to the specified state parameter.

[0273] It can be understood that the above-mentioned step S402 and step S404 may be executed by the control component 1011 in the foregoing embodiment, and the specific execution process has been described in the foregoing embodiment and will not be elaborated here.

[0274] Through the above method, when there are K activated computing domains in the computing domain cluster, a set of state parameters of the activated computing domains is monitored, where K is a positive integer greater than or equal to 1; among the K activated computing domains, when there is a first computing domain in which the parameter value of the specified state parameter in a set of state parameters meets the corresponding threshold condition, the operating parameters of the first computing domain are adjusted according to the adjustment method corresponding to the specified state parameter, which can solve the problem of poor timeliness of operation control of electronic components in the related art, and achieve the technical effects of improving the timeliness and flexibility of operation control of electronic components.

[0275] In an exemplary embodiment, a set of state parameters includes the number of instructions per cycle, and the threshold condition corresponding to the number of instructions per cycle is that the value of the number of instructions per cycle is less than or equal to the first instruction number threshold; adjusting the operating parameters of the first computing domain according to the adjustment method corresponding to the specified state parameter includes: when the specified state parameter includes the number of instructions per cycle, performing a first adjustment operation on the first computing domain to adjust the operating parameters of the first computing domain, where the first adjustment operation includes: lowering the frequency of the first computing domain and lowering the voltage of the first computing domain.

[0276] In an exemplary embodiment, the electronic component further includes a hardware performance counter; the above method further includes: monitoring the number of instructions per cycle of the K activated computing domains in real time through the hardware performance counter.

[0277] In an exemplary embodiment, a set of state parameters includes the computing domain temperature, and the threshold condition corresponding to the computing domain temperature is that the temperature value of the computing domain temperature is greater than or equal to the first temperature threshold; adjusting the operating parameters of the first computing domain according to the adjustment method corresponding to the specified state parameter includes: when the specified state parameter includes the computing domain temperature, performing a second adjustment operation on the first computing domain to adjust the operating parameters of the first computing domain, where the second adjustment operation includes at least one of the following: lowering the frequency of the first computing domain and performing voltage compensation on the second computing domain, where the second computing domain is an activated computing domain adjacent to the first computing domain and does not meet the threshold condition corresponding to the computing domain temperature; migrating the task executed by the first computing domain to an activated computing domain that does not meet the threshold condition corresponding to the computing domain temperature.

[0278] In an exemplary embodiment, performing a second adjustment operation on a first computing domain includes: lowering the frequency of the first computing domain by a specified frequency value, where the specified frequency value is the product of a first coefficient and a specified temperature difference, and the specified temperature difference is the difference between the temperature value of the computing domain temperature of the first computing domain and a first temperature value; raising the voltage of a second computing domain by a specified voltage value, where the specified voltage value is the product of a second coefficient, a specified difference, and a specified voltage compensation value, and the specified difference is the difference obtained by subtracting the ratio of the temperature value of the computing domain temperature of the second computing domain to the temperature value of the computing domain temperature of the first computing domain from 1.

[0279] In an exemplary embodiment, the electronic component further includes a distributed sensor network; monitoring a set of state parameters of an activated computing domain includes: when the set of state parameters includes the computing domain temperature, parsing the sensor data corresponding to the activated computing domain to obtain the actual temperature value of the activated computing domain at the current moment, where the sensor data corresponding to the activated computing domain is the sensor data obtained by the sensor network detecting the activated computing domain respectively; predicting the temperature value of the activated computing domain at the next moment based on the actual temperature value of the activated computing domain at the current moment, the current power value of the activated computing domain, and the ambient temperature value at the current moment, where the current power value of the activated computing domain is the power value of the activated computing domain at the current moment, and the temperature value of the computing domain temperature of the activated computing domain is the predicted temperature value of the activated computing domain at the next moment.

[0280] In an exemplary embodiment, predicting the temperature value of the activated computing domain at the next moment based on the actual temperature value of the activated computing domain at the current moment, the current power value of the activated computing domain, and the ambient temperature value at the current moment includes: determining the sum result of the product of the actual temperature value of the activated computing domain at the current moment and a thermal inertia coefficient, the product of the current power value of the activated computing domain and a power-temperature conversion coefficient, and the product of the ambient temperature value at the current moment and a specified heat exchange rate as the predicted temperature value of the activated computing domain at the next moment, where the specified heat exchange rate is the heat exchange rate between the electronic component and the environment where the electronic component is located.

[0281] In an exemplary embodiment, the electronic component further includes a control component, the sensor network includes a set of sensing nodes arranged in a grid layout, and one computing domain in the computing domain cluster corresponds to a part of the sensing nodes in the set of sensing nodes; the above method further includes: obtaining the sensor data detected by the sensing nodes corresponding to K activated computing domains in an address polling manner through the control component to obtain the sensor data corresponding to the K activated computing domains.

[0282] In an exemplary embodiment, the electronic component further includes a temperature sensor; the method further includes: periodically detecting the ambient temperature through the temperature sensor, and calculating the temperature difference between the detected ambient temperature value and the recorded reference ambient temperature value; when the temperature difference between the detected ambient temperature value and the recorded reference ambient temperature value is greater than or equal to the temperature difference threshold, multiplying the detected ambient temperature value by a specified coefficient and adding the second temperature value to determine the updated first temperature threshold, and updating the recorded reference ambient temperature value to the detected ambient temperature value.

[0283] In an exemplary embodiment, the method further includes: when there is a first computing task to be allocated, selecting a third computing domain for executing the first computing task from the K activated computing domains according to the task label of the first computing task and the operation characteristic data of the K activated computing domains, where the task label of the first computing task is used to indicate the task attribute of the first computing task and the execution constraint conditions of the first computing task; allocating the first computing task to the third computing domain for the first computing unit of the third computing domain to execute the first computing task.

[0284] In an exemplary embodiment, the operation state data of the activated computing domain includes the historical performance data of the activated computing domain; selecting a third computing domain for executing the first computing task from the K activated computing domains according to the task label of the first computing task and the operation characteristic data of the K activated computing domains includes: inputting the task label of the first computing task and the historical performance data of the activated computing domain into the task level prediction engine to obtain the task level of the first computing task predicted by the task level prediction engine, where the task level of the first computing task is used to characterize the computing power requirement for executing the first computing task; selecting one activated computing domain that matches the task level of the first computing task from the K activated computing domains to obtain the third computing domain.

[0285] In an exemplary embodiment, one of the multiple computing units allows a computing task to be executed in at least one precision mode. The higher the computing precision corresponding to the precision mode, the more floating-point bits are used to execute the computing task; the task attribute of the first computing task includes the precision requirement of the first computing task, and the precision requirement of the first computing task is used to indicate the lowest precision mode required to execute the first computing task; selecting an activated computing domain that matches the task level of the first computing task from the K activated computing domains to obtain a third computing domain, including: when there are at least two candidate computing domains whose allowed precision modes match the precision requirement of the first computing task among the K activated computing domains, calculating the computing domain score of the candidate computing domains according to the performance of the candidate computing domains, the power consumption of the candidate computing domains, and the temperature of the candidate computing domains, where the computing domain score of the candidate computing domains is positively correlated with the performance of the candidate computing domains and negatively correlated with the power consumption and the temperature of the candidate computing domains, and the performance of the candidate computing domains is the number of instructions per cycle of the candidate computing domains; determining the candidate computing domain with the highest computing domain score among the at least two candidate computing domains as the third computing domain.

[0286] In an exemplary embodiment, one of a plurality of computing units allows a computing task to be executed in at least one precision mode. The higher the computing precision corresponding to the precision mode, the more floating-point bits are used to execute the computing task. The task attribute of the first computing task includes the precision requirement of the first computing task, and the precision requirement of the first computing task is used to indicate the lowest precision mode required to execute the first computing task. The third computing domain includes at least two types of computing units. The method further includes: screening computing units from the computing units in the third computing domain according to a preset screening condition, where the preset screening condition includes one of the following: when the temperature of the electronic component is greater than or equal to the second temperature threshold, screening computing units with a temperature lower than the third temperature threshold; screening computing units that meet the execution constraint conditions; screening computing units whose allowed precision mode matches the precision requirement of the first computing task; when at least two candidate computing units are screened out, calculating a computing unit score of the candidate computing units according to the performance of the candidate computing units, the power consumption of the candidate computing units, and the temperature of the candidate computing units, where the computing unit score of the candidate computing units is positively correlated with the performance of the candidate computing units and negatively correlated with the power consumption of the candidate computing units and the temperature of the candidate computing units, and the performance of the candidate computing units is the number of instructions per cycle of the candidate computing units; when the precision mode allowed by the candidate computing unit with the highest computing unit score among at least two candidate computing units matches the precision requirement of the first computing task, determining the candidate computing unit with the highest computing unit score as the first computing unit; when the precision mode allowed by the candidate computing unit with the highest computing unit score among at least two candidate computing units does not match the precision requirement of the first computing task, determining the specified computing unit in the third computing domain as the first computing unit.

[0287] In an exemplary embodiment, one of a plurality of computing units allows a computing task to be executed in at least one precision mode. The higher the computing precision corresponding to the precision mode, the more floating-point bits are used to execute the computing task. The method further includes: extracting a to-be-executed second computing task from the computing task queue of the fourth computing domain among K activated computing domains, where the second computing task has at least two computing stages, and the precision requirements of the second computing task are different in different computing stages among the at least two computing stages, and the precision requirement of the second computing task is used to indicate the lowest precision mode required to execute the second computing task; in different computing stages of the second computing task, executing the second computing task in different precision modes through at least one computing unit in the fourth computing domain, so that the precision mode adopted by the computing unit in the fourth computing domain matches the precision requirement of the computing stage of the second computing task.

[0288] In an exemplary embodiment, the activated computing domain is configured with multiple load levels; the method further includes: obtaining prediction reference data of the activated computing domain, where the prediction reference data includes the instruction mix, cache miss rate, and task queue depth of the task queue of the activated computing domain in N historical periods, and N is a positive integer greater than or equal to 1; inputting the prediction reference data into a load level prediction engine to obtain the load levels predicted by the load level prediction engine for the activated computing domain in M future periods, where M is a positive integer greater than or equal to 1; and pre-adjusting the voltage and frequency of the activated computing domain based on the predicted load levels of the activated computing domain in M future periods.

[0289] In an exemplary embodiment, the method further includes: when the instructions per cycle of the activated computing domain are less than or equal to a second instruction number threshold and the temperature of the activated computing domain is less than or equal to a fourth temperature threshold, controlling the activated computing domain to enter a near-threshold computing mode; in the near-threshold computing mode, adjusting the voltage offset of the activated computing domain based on the operating state data of the activated computing domain, where the operating state data of the activated computing domain is used to indicate the temperature, voltage, and process deviation coefficient of the activated computing domain, and the voltage offset of the activated computing domain is obtained by summing the product of the difference between the temperature of the activated computing domain and a specified temperature and a third coefficient, the product of the difference between the voltage of the activated computing domain and a specified voltage and a fourth coefficient, and the product of the process deviation coefficient and a fifth coefficient.

[0290] In an exemplary embodiment, after controlling the activated computing domain to enter the near-threshold computing mode, the method further includes: in the near-threshold computing mode, sampling data of the activated computing domain through a master latch and a shadow latch, and detecting timing errors of the activated computing domain, where the trigger edges of the master latch and the shadow latch are different; when a timing error of the activated computing domain is detected, continuously increasing the voltage of the activated computing domain until the error instruction is successfully executed, where the error instruction is the instruction executed by the activated computing domain corresponding to the timing error of the activated computing domain; or controlling the activated computing domain to exit the near-threshold computing mode.

[0291] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.

[0292] The above has introduced in detail an electronic component and a method for controlling the operation of the electronic component provided in this application. Specific examples are used in this text to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. It should be noted that for those of ordinary skill in the art, without departing from the principle of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the protection scope of the claims of this application.

Claims

1. An electronic component, characterized in that, It includes a control component and multiple computing units, and the multiple computing units are divided into computing domain clusters, where a computing domain in the computing domain cluster is a set of computing units that independently adjust operating parameters; In the computing domain cluster, one computing domain includes some of the multiple computing units, and different computing domains include different computing units; where, The control component is configured to monitor a set of state parameters of the activated computing domains when there are K activated computing domains in the computing domain cluster, where K is a positive integer greater than or equal to 1; when there is a first computing domain in the K activated computing domains where the parameter value of a specified state parameter in the set of state parameters meets the corresponding threshold condition, the operating parameters of the first computing domain are adjusted according to the adjustment method corresponding to the specified state parameter.

2. The electronic component according to claim 1, characterized in that, The electronic component further includes at least one of the following: A voltage regulation unit respectively configured for the computing domains in the computing domain cluster, which is used to regulate the voltage of the corresponding computing domain. Wherein, the operating parameters of the computing domains in the computing domain cluster include the voltage of the computing domains in the computing domain cluster, and the voltage regulation unit includes at least one of the following: a low dropout regulator, a switched capacitor converter; A local clock tree synthesis element, an asynchronous bridge, and a digital phase-locked loop respectively configured for the computing domains in the computing domain cluster. Wherein, the local clock tree synthesis element is used to distribute the global clock signal of the electronic component to different computing domains in the computing domain cluster to form clock domains corresponding to different computing domains in the computing domain cluster, the asynchronous bridge establishes a data transmission path between different clock domains, and the digital phase-locked loop is used to regulate the frequency of the corresponding computing domain. The operating parameters of the computing domains in the computing domain cluster include the frequency of the computing domains in the computing domain cluster.

3. The electronic component according to claim 1, characterized in that, The control component is connected to the computing units in the computing domains in the computing domain cluster through an on-chip interconnect network.

4. The electronic component according to claim 1, characterized in that, The set of state parameters includes the instructions per cycle, and the threshold condition corresponding to the instructions per cycle is that the value of the instructions per cycle is less than or equal to a first instruction number threshold; The control component is further configured to perform a first adjustment operation on the first computing domain to adjust the operating parameters of the first computing domain when the specified state parameter includes the instructions per cycle. Wherein, the first adjustment operation includes: lowering the frequency of the first computing domain and lowering the voltage of the first computing domain.

5. The electronic component according to claim 4, wherein The electronic component further includes: A hardware performance counter, which is used to monitor the instructions per cycle of the K activated computing domains in real time.

6. The electronic component according to claim 1, characterized in that, The set of state parameters includes the computing domain temperature, and the threshold condition corresponding to the computing domain temperature is that the temperature value of the computing domain temperature is greater than or equal to a first temperature threshold; The control component is further configured to, when the specified state parameter includes the computing domain temperature, perform a second adjustment operation on the first computing domain to adjust the operating parameters of the first computing domain, where the second adjustment operation includes at least one of the following: lowering the frequency of the first computing domain and performing voltage compensation on a second computing domain, where the second computing domain is an activated computing domain that is adjacent to the first computing domain and does not meet the threshold condition corresponding to the computing domain temperature; migrating the tasks executed by the first computing domain to the activated computing domain that does not meet the threshold condition corresponding to the computing domain temperature.

7. The electronic component according to claim 6, characterized in that, The control component is further configured to lower the frequency of the first computing domain by a specified frequency value, where the specified frequency value is the product of a first coefficient and a specified temperature difference, and the specified temperature difference is the difference between the temperature value of the computing domain temperature of the first computing domain and a first temperature value; raise the voltage of the second computing domain by a specified voltage value, where the specified voltage value is the product of a second coefficient, a specified difference, and a specified voltage compensation value, and the specified difference is the difference obtained by subtracting the ratio of the temperature value of the computing domain temperature of the second computing domain to the temperature value of the computing domain temperature of the first computing domain from 1.

8. The electronic component according to claim 6, wherein The electronic component further includes a distributed sensor network; where the sensor network is configured to, when the set of state parameters includes the computing domain temperature, respectively detect the temperatures of the activated computing domains to obtain the sensor data corresponding to the activated computing domains; The control component is further configured to analyze the sensor data corresponding to the activated computing domains to obtain the actual temperature value of the activated computing domains at the current moment; predict the temperature value of the activated computing domains at the next moment based on the actual temperature value of the activated computing domains at the current moment, the current power value of the activated computing domains, and the ambient temperature value at the current moment, where the current power value of the activated computing domains is the power value of the activated computing domains at the current moment, and the temperature value of the computing domain temperature of the activated computing domains is the predicted temperature value of the activated computing domains at the next moment.

9. The electronic component according to claim 8, wherein The control component is further configured to determine the sum result of the product of the actual temperature value of the activated computing domains at the current moment and the thermal inertia coefficient, the product of the current power value of the activated computing domains and the power-temperature conversion coefficient, and the product of the ambient temperature value at the current moment and the specified heat exchange rate as the predicted temperature value of the activated computing domains at the next moment, where the specified heat exchange rate is the heat exchange rate between the electronic component and the environment where the electronic component is located.

10. The electronic component according to claim 8, wherein The sensor network includes a set of sensing nodes arranged in a grid layout, and one computing domain in the computing domain cluster corresponds to a part of the sensing nodes in the set of sensing nodes; The control component is further configured to obtain the sensor data detected by the sensing nodes corresponding to the K activated computing domains in an address polling manner, so as to obtain the sensor data corresponding to the K activated computing domains.

11. The electronic component according to claim 7, characterized in that, The electronic component further includes a temperature sensor; wherein, The temperature sensor is configured to periodically detect the ambient temperature; The temperature sensor or the control component is further configured to calculate the temperature difference between the detected ambient temperature value and the recorded reference ambient temperature value; when the temperature difference between the detected ambient temperature value and the recorded reference ambient temperature value is greater than or equal to the temperature difference threshold, the product of the detected ambient temperature value and a specified coefficient is added to the second temperature value to determine the updated first temperature threshold, and the recorded reference ambient temperature value is updated to the detected ambient temperature value.

12. The electronic component according to claim 1, characterized in that, The control component is further configured to, when there is a first computing task to be allocated, select a third computing domain for executing the first computing task from the K activated computing domains according to the task label of the first computing task and the operation characteristic data of the K activated computing domains, wherein the task label of the first computing task is used to indicate the task attribute of the first computing task and the execution constraint conditions of the first computing task; and allocate the first computing task to the third computing domain, so that the first computing unit of the third computing domain executes the first computing task.

13. The electronic component according to claim 12, wherein A task level prediction engine runs on the control component, and the task level prediction engine is configured to predict the task level of a computing task corresponding to an input task label based on the input task label and the input historical performance data; wherein, The control component is further configured to input the task label of the first computing task and the historical performance data of the activated computing domains into the task level prediction engine to obtain the task level of the first computing task predicted by the task level prediction engine, wherein the task level of the first computing task is used to characterize the computing power requirement for executing the first computing task; and select one of the K activated computing domains that matches the task level of the first computing task from the K activated computing domains to obtain the third computing domain.

14. The electronic component according to claim 13, characterized in that, One of the multiple computing units is allowed to execute a computing task in at least one precision mode, and the higher the computing precision corresponding to the precision mode, the more floating-point bits are used for executing the computing task; the task attribute of the first computing task includes the precision requirement of the first computing task, and the precision requirement of the first computing task is used to indicate the lowest precision mode required for executing the first computing task. The control component is further configured to, when there are at least two candidate computing domains among the K activated computing domains whose allowed precision modes match the precision requirement of the first computing task, calculate a computing domain score for the candidate computing domains according to the performance, power consumption, and temperature of the candidate computing domains, where the computing domain score of the candidate computing domain is positively correlated with the performance of the candidate computing domain and negatively correlated with the power consumption and temperature of the candidate computing domain, and the performance of the candidate computing domain is the number of instructions per cycle of the candidate computing domain; Among at least two of the candidate computing domains, the candidate computing domain with the highest computing domain score is determined as the third computing domain.

15. The electronic component according to claim 13, characterized in that, One of the multiple computing units is allowed to execute a computing task in at least one precision mode. The higher the computing precision corresponding to the precision mode, the more floating-point bits are used to execute the computing task; the task attribute of the first computing task includes the precision requirement of the first computing task, and the precision requirement of the first computing task is used to indicate the lowest precision mode required to execute the first computing task; the third computing domain includes at least two types of computing units; The third computing domain is configured to screen computing units from the computing units of the third computing domain according to a preset screening condition; when at least two candidate computing units are screened out, calculate a computing unit score for the candidate computing units according to the performance, power consumption, and temperature of the candidate computing units; When the precision mode allowed by the candidate computing unit with the highest computing unit score among at least two candidate computing units matches the precision requirement of the first computing task, the candidate computing unit with the highest computing unit score is determined as the first computing unit; when the precision mode allowed by the candidate computing unit with the highest computing unit score among at least two candidate computing units does not match the precision requirement of the first computing task, the specified computing unit in the third computing domain is determined as the first computing unit; Wherein, the preset screening condition includes one of the following: when the temperature of the electronic component is greater than or equal to a second temperature threshold, screening the computing units with a temperature lower than a third temperature threshold; screening the computing units that meet the execution constraint conditions; screening the computing units whose allowed precision modes match the precision requirement of the first computing task; Wherein, the computing unit score of the candidate computing unit is positively correlated with the performance of the candidate computing unit and negatively correlated with the power consumption and temperature of the candidate computing unit, and the performance of the candidate computing unit is the number of instructions per cycle of the candidate computing unit.

16. The electronic component according to claim 1, wherein One of the multiple computing units is allowed to execute a computing task in at least one precision mode. The higher the computing precision corresponding to the precision mode, the more floating-point bits are used to execute the computing task; The fourth computational domain among the K activated computational domains is configured to extract a second computational task to be executed from the computational task queue of the fourth computational domain. The second computational task has at least two computational stages, and the precision requirements of the second computational task are different in different computational stages among the at least two computational stages. The precision requirements of the second computational task are used to indicate the lowest precision mode required to execute the second computational task. In different computational stages of the second computational task, at least one computational unit in the fourth computational domain executes the second computational task in different precision modes, so that the precision mode adopted by the computational unit in the fourth computational domain matches the precision requirements of the computational stage of the second computational task.

17. The electronic component according to claim 1, characterized in that, The activated computational domain is configured with multiple load levels, and a load level prediction engine runs on the control component. The load level prediction engine is configured to predict the load level of the computational domain corresponding to the input prediction reference data in M future cycles based on the input prediction reference data. The prediction reference data includes the following data of the computational domains in the computational domain cluster in N historical cycles: instruction mix, cache miss rate, and task queue depth of the computational task queue. M is a positive integer greater than or equal to 1, and N is a positive integer greater than or equal to 1. Among them, The control component is further configured to obtain the prediction reference data of the activated computational domain; input the prediction reference data of the activated computational domain into the load level prediction engine to obtain the load level of the activated computational domain predicted by the load level prediction engine in the M future cycles; pre-adjust the operating parameters of the activated computational domain based on the predicted load level of the activated computational domain in the M future cycles.

18. The electronic component according to claim 1, characterized in that, The control component is further configured to control the activated computational domain to enter the near-threshold computing mode when the instructions per cycle of the activated computational domain are less than or equal to a second instruction number threshold and the temperature of the activated computational domain is less than or equal to a fourth temperature threshold. In the near-threshold computing mode, the voltage offset of the activated computational domain is adjusted based on the operating state data of the activated computational domain. The operating state data of the activated computational domain is used to indicate the temperature of the activated computational domain, the voltage of the activated computational domain, and the process deviation coefficient of the activated computational domain. The voltage offset of the activated computational domain is obtained by summing the product of the difference between the temperature of the activated computational domain and the specified temperature and the third coefficient, the product of the difference between the voltage of the activated computational domain and the specified voltage and the fourth coefficient, and the product of the process deviation coefficient and the fifth coefficient.

19. The electronic component according to claim 18, characterized in that, The electronic component further includes: a master latch and a shadow latch, where the triggering edges of the master latch and the shadow latch are different. Among them, The master latch and the shadow latch are configured to sample data of the activated computational domain when the activated computational domain enters the near-threshold computing mode. The control component is further configured to sample data of the activated computing domain through the master latch and the shadow latch, and detect timing errors of the activated computing domain; in the case that a timing error of the activated computing domain is detected, continuously increase the voltage of the activated computing domain until the error instruction is successfully executed, where the error instruction is an instruction executed by the activated computing domain and corresponding to the timing error of the activated computing domain; or control the activated computing domain to exit the near-threshold computing mode.

20. A method for operating control of an electronic component, characterized in that, The electronic component includes a plurality of computing units, and the plurality of computing units are divided into computing domain clusters, and the computing domains in the computing domain clusters are sets of computing units that independently adjust operating parameters; In the computing domain cluster, one computing domain includes some of the plurality of computing units, and different computing domains include different computing units; the method includes: When there are K activated computing domains in the computing domain cluster, monitor a set of state parameters of the activated computing domains, where K is a positive integer greater than or equal to 1; In the case that there is a first computing domain among the K activated computing domains where the parameter value of a specified state parameter in the set of state parameters satisfies the corresponding threshold condition, adjust the operating parameters of the first computing domain according to the adjustment method corresponding to the specified state parameter.

Citation Information

Patent Citations

  • High energy efficiency resource allocating method for isomorphic cluster system of computer

    CN102749987A

  • Adaptive adjustment method for calculation resources in multiple working areas

    CN106708624A

  • Node state adjusting method and device, computer equipment and storage medium

    CN116841698A

  • Resource scheduling method and device, electronic equipment and storage medium

    CN116880984A

  • Cluster management and dynamic scheduling system and method for computing host

    TW202115585A