Apparatus, system and method for detecting violations of physical infrastructure constraints - Patents.com

The system detects and prevents violations of physical infrastructure constraints in computing systems by using a microcontroller to trigger machine check exceptions for timely corrective actions, addressing the limitations of existing architectures in detecting these violations.

JP2025532828APending Publication Date: 2025-10-03ADVANCED MICRO DEVICES INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025517588
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-09-30
Filing Date
2023-09-29
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Modern computing systems lack the ability to detect violations of physical infrastructure constraints such as voltage, power, sustained current, peak current, and temperature before they cause hardware failure, as existing machine check architectures and exception functions do not provide visibility into these lower-level violations.

Method used

A system and method that utilizes a microcontroller to report violations of physical infrastructure constraints to a machine check architecture, triggering a machine check exception for recording and initiating corrective actions, including logging, reporting, and profiling telemetry events, thereby preventing potential hardware failures.

Benefits of technology

The solution enables early detection and prevention of physical infrastructure constraint violations, reducing the risk of hardware failure by allowing for proactive corrective actions before the constraints are breached, thus enhancing system reliability and stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025532828000001_ABST
    Figure 2025532828000001_ABST
Patent Text Reader

Abstract

The disclosed method may include (i) the microcontroller reporting the detection of a violation of a physical infrastructure constraint to a machine check architecture, (ii) the machine check architecture, in response to the reporting, triggering a machine check exception such that the violation of the physical infrastructure constraint is logged, and (iii) performing a corrective action based on the triggering of the machine check exception. Various other apparatus, systems, and methods are also disclosed.
Need to check novelty before this filing date? Find Prior Art

Description

[Background technology]

[0001] Modern computer chips may implement several different features to detect errors and facilitate debugging. In this regard, this application discloses improved apparatus, systems, and methods for detecting physical infrastructure constraint violations.

[0002] The accompanying drawings illustrate several exemplary versions and are a part of this specification, and together with the following description, these drawings demonstrate and explain various principles of the present disclosure. [Brief explanation of the drawings]

[0003] [Figure 1] FIG. 1 is a flow diagram of an example method for detecting physical infrastructure constraint violations. [Figure 2] FIG. 1 is a block diagram of an exemplary system-on-chip. [Figure 3] FIG. 1 is a block diagram of a taxonomy of detected and undetected errors in a computing environment. [Figure 4] FIG. 10 is an exemplary block diagram of a machine check architecture register. [Figure 5] FIG. 2 is a flow diagram of an exemplary method for detecting physical infrastructure constraint violations in one version. DETAILED DESCRIPTION OF THE INVENTION

[0004] Throughout the drawings, identical reference numerals and descriptions indicate similar, but not necessarily identical, elements. While the exemplary versions described herein are susceptible to various modifications and alternative forms, specific versions have been shown by way of example in the drawings and are described in detail herein. However, the exemplary versions described herein are not limited to the particular forms disclosed. Rather, the present disclosure covers all modifications, equivalents, and alternatives falling within the scope of the appended claims.

[0005] This disclosure describes various apparatus, systems, and methods for detecting violations of physical infrastructure constraints. Modern computing microprocessors may feature machine check architectures and corresponding machine check exception functions that report errors to operating systems or other software components. These mechanisms may detect and report hardware or machine errors (e.g., system bus errors, errors caused by communication over noisy channels, parity errors, cache errors, translation lookaside buffer errors, etc.). Nevertheless, these related methods may not necessarily detect lower-level violations of physical infrastructure constraints. In other words, physical or hardware components of modern computing systems may feature physical infrastructure constraints related to voltage, power, sustained current, peak current, temperature, etc.

[0006] More generally, as used herein, the term "physical infrastructure constraint" generally refers to constraints defined on physical or performance-related properties of corresponding physical or hardware components of a computing device, where, as explained further below, the properties should not meet a threshold (e.g., a maximum or minimum threshold) or should not meet the threshold for more than a certain amount of time. Thus, these components should not achieve values ​​that violate these constraints (e.g., violate the constraint for more than a predetermined amount of time). In some examples, a physical infrastructure constraint may be violated even if the violation has not yet caused the corresponding physical or hardware component to fail.

[0007] As one illustrative example, a particular conductor or lead may have an electrical design current constraint that specifies a maximum value for the current passing through it. Thus, the conductor or lead should not carry a current greater than that maximum value, or for a predetermined amount of time, yet the conductor or lead may continue to carry such a current without causing the conductor or lead to fail, violating the constraint. Nevertheless, for a sufficient amount of time, physical components are likely to fail, which is one purpose behind the imposed physical infrastructure constraint. Again, this presents another reason why monitoring for such constraint violations is important, yet modern machine check architectures and corresponding machine check exception functions may not have visibility into such violations.

[0008] As described in more detail below, the present disclosure generally relates to apparatus, systems, and methods for detecting violations of physical infrastructure constraints. In one example, the method may include: (i) a microcontroller reporting a detection of a violation of a physical infrastructure constraint to a machine check architecture; (ii) in response to the reporting, the machine check architecture triggering a machine check exception such that the violation of the physical infrastructure constraint is recorded; and (iii) performing a corrective action based on the triggering of the machine check exception.

[0009] In some examples, the microcontroller includes a system management unit.

[0010] In some examples, violations of physical infrastructure constraints are defined in terms of at least one of voltage, power, sustained current, peak current, or temperature.

[0011] In some examples, violations of physical infrastructure constraints include violations of electrical design current constraints.

[0012] In some examples, the corrective action includes at least one of logging, reporting, or profiling telemetry events.

[0013] In some examples, the corrective action is performed by a prediction and prevention engine.

[0014] In some instances, corrective action is performed as part of a debugging diagnosis.

[0015] In some examples, the corrective action includes entering a debugging mode or resetting the processor.

[0016] In some instances, violations of physical infrastructure constraints are detected before the violations cause any physical components to fail.

[0017] In some examples, the microcontroller and machine check architecture are located in a system-on-chip.

[0018] In a further example, the machine check architecture includes (i) a receptor that receives a report of a violation of a physical infrastructure constraint from a microcontroller; (ii) a trigger that, in response to receiving the report, triggers a machine check exception such that the violation of the physical infrastructure constraint is recorded; and (iii) an initiator that initiates corrective action based on the recording of the violation of the physical infrastructure constraint.

[0019] In a further example, the system may include a microcontroller that issues a report of a violation of a physical infrastructure constraint, and a machine check architecture that, in response to receiving the report, triggers a machine check exception so that the violation of the physical infrastructure constraint is recorded and initiates corrective action based on the recording of the violation of the physical infrastructure constraint.

[0020] 1 shows a flow diagram of a method 100 for detecting a violation of a physical infrastructure constraint. The steps of method 100 may be performed by various devices, microcontrollers, machine check architectures, and / or systems described herein. For example, in step 102, a microcontroller may report the detection of a violation of a physical infrastructure constraint to a machine check architecture. As used herein, the term "microcontroller" generally refers to a small integrated circuit designed to manage operation, for example, within an embedded system.

[0021] To help provide context for the environment in which method 100 may be performed, Figure 2 illustrates an example block diagram of a system-on-chip that may include both a microcontroller and a machine check architecture, which are described further below.

[0022] In an exemplary version, SOC 200 includes multiple processing cores 220, 222, 224, 226, multiple trace data storage elements 240, 242, 244, 246, 254, 264, 274 (e.g., L2 cache and Trace Capture Buffers (TCBs)), a northbridge 250 (or memory controller), a southbridge 260, a GPU 270, and a cross-trigger bus 280. Cross-triggering to debugging state machines (not shown) on another die within the same package (e.g., within an MCM) and / or other packages may be achieved via an off-chip debugging state machine interface 212.

[0023] While SOC 200 is shown as including four cores 220, 222, 224, and 226, the SOC may include more or fewer cores in other versions (including as few as one single core). Note that while SOC 200 is shown as including a single northbridge 250, southbridge 260, and GPU 270, in other versions, some or all of these electronic modules may be omitted from SOC 200 (e.g., they may be located off-chip). Furthermore, while SOC 200 is shown as including only a single northbridge 250, in other versions, the SOC may include two or more northbridges. In other versions, in addition to the processing components and buses shown in FIG. 2 , the SOC may include additional or different processing components, buses, and other electronic devices and circuits.

[0024] Processing cores 220, 222, 224, 226 generally represent the primary processing hardware, logic, and / or circuitry for SOC 200, and each processing core 220, 222, 224, 226 may be implemented using one or more arithmetic logic units (ALUs), one or more floating point units (FPUs), one or more memory elements (e.g., one or more caches), discrete gate or transistor logic, discrete hardware components, or any combination thereof. Although not shown in FIG. 2, each processing core 220, 222, 224, 226 may implement its own associated cache memory element (e.g., a level 1 or L1 cache) in close proximity to the respective processing circuitry to reduce latency.

[0025] Northbridge 250, which may also be referred to as a "memory controller" in some systems, is configured to interface with I / O peripherals (e.g., I / O peripherals 140 of FIG. 1) and memory (e.g., memory 150 of FIG. 1). Northbridge 250 controls communication between the components of SOC 200 and the I / O peripherals and / or external memory. Southbridge 260, which may also be referred to as an "I / O controller hub" in some systems, is configured to connect and control peripheral devices (e.g., relatively slower peripheral devices). GPU 270 is a dedicated microprocessor that offloads and accelerates graphics rendering from cores 220, 222, 224, and 226.

[0026] In the illustrated version, caches 240, 242, 244, 246 and TCBs 254, 264, 274 provide intermediate memory elements having reduced sizes relative to external memory for temporarily storing data and / or instructions retrieved from external memory or elsewhere and / or data generated by processing cores 220, 222, 224, 226, northbridge 250, southbridge 260, and GPU 270. For example, in one version, caches 240, 242, 244, 246 and TCBs 254, 264, 274 provide memory elements for storing debug information (or “trace data”) collected and / or generated by debugging state machines during debug operations associated with the respective electronic modules into which they are integrated. In the illustrated version, caches 240, 242, 244, 246 are closely coupled between their respective processing cores 220, 222, 224, 226 and northbridge 250. In this regard, caches 240, 242, 244, 246 may alternatively be referred to as core-coupled caches, with each core-coupled cache 240, 242, 244, 246 maintaining data and / or program instructions previously fetched from external memory that have either previously been used and / or are likely to be used by its associated processing core 220, 222, 224, 226. Caches 240, 242, 244, 246 are preferably larger than the L1 caches implemented by processing cores 220, 222, 224, 226 and function as level 2 caches (or L2 caches) in the memory hierarchy. SOC 200 may also include another higher level cache (e.g., a level 3 or L3 cache, not shown), which is preferably larger than L2 cache 240, 242, 244, 246.

[0027] In an exemplary version, SOC 200 includes test interface 210 that includes a number of pins dedicated for use in testing and / or configuring the functionality of SOC 200. In one version, test interface 210 conforms to the IEEE 1149.1 Standard Test Access Port and Boundary Scan Architecture, i.e., the Joint Test Action Group (JTAG) standard.

[0028] 2 also illustrates how SOC 200 may further include a subsystem 288 for performing method 100. Subsystem 288 may further include a machine check architecture 290 (MCA). Machine check architecture 290 may include a receptor 292 that may perform step 102 of method 100 and a trigger 294 that may perform step 104 of method 100, as described further below. Subsystem 288 may further include an initiator 296 for performing step 106, as described further below. Subcomponents 292-296 may be implemented in terms of hardware or logic within SOC 200, and / or one or more of them may be implemented as firmware or software components in conjunction with a hardware-based MCA.

[0029] As mentioned above, the present technology may modify or improve the machine check architecture and corresponding machine check exception functionality. Accordingly, the following provides an overview of the machine check architecture and its corresponding machine check exception functionality.

[0030] An MCA, such as MCA 290, may refer to a processor-centric error detection and reporting mechanism. Figure 3 shows a useful taxonomy 300 that breaks down system errors 302 into sets. The two main sets are detected errors 304 and undetected errors 306. Undetected errors may be important but are not handled by the MCA. Undetected errors may include benign errors 312 and critical errors 314.

[0031] Detected errors can be divided into corrected errors 308 and uncorrected errors 310. Corrected errors are harmless after being corrected. Despite being harmless, these errors can still be reported to software for tracking purposes. Some errors should indicate that an error is expected, and therefore it is useful to monitor them. Uncorrected errors can include three different sets: catastrophic errors 316, fatal errors 318, and recoverable errors 320. Both catastrophic and fatal errors trigger a system reset. Uncorrected recoverable errors (UCR) may allow the system to still function, but they can be categorized into uncorrected no action errors 322, software recoverable action optional errors 324, and software recoverable action required errors 326.

[0032] Assuming these error sets are detected by hardware components, functions are created to alert software (such as firmware, BIOS, operating system, or hypervisor) to the errors. Corrective action can include two versions: recording and reporting. Recording can be performed by recording data in Machine Check (MC) registers, as shown in FIG. 4, which can be reviewed by software and saved to memory. As further shown in FIG. 4, MC registers 400 include a set of global registers 402 and a set of bank registers 404-410. Reporting can use CMCI reporting and MCE reporting. The CMCI (Corrected Machine Check Interrupt) function can report corrected errors (CE) and uncorrected recoverable errors (UCR), such as uncorrected no action (UCNA) recoverable errors. MCE (machine-check exception) reporting may address uncorrected recoverable errors such as the software recoverable action optional (SRAO) and software recoverable action required (SRAR) errors listed above. This function may also report catastrophic errors that result in a system reboot. In the scenarios outlined above, data about the event may be saved to the MC bank before reporting.

[0033] 1, step 102 may be performed by a microcontroller that detects violations of physical infrastructure constraints. As further described above, physical infrastructure constraints may refer to maximum values ​​or thresholds that define properties of physical or hardware components that should be avoided. Thus, examples of such physical infrastructure constraints may include constraints on voltage, power, sustained current, peak current, or temperature.

[0034] In one illustrative example, the microcontroller may correspond to a System Management Unit (SMU). The SMU may be a subcomponent of a larger processor (e.g., SOC), which may be responsible for various system and power management tasks during boot and runtime. The SMU may include a thermal block, which may further include features related to temperature sensing, control, and reporting. The thermal block may include temperature collection and calculation logic, fan speed control for off-chip fans, and temperature reporting capabilities. The SMU may also include several registers, which may provide current-controlled temperature, among other outputs.

[0035] 1 , in step 104, one or more of the devices or systems described herein may trigger a machine check exception such that a violation of a physical infrastructure constraint is logged. For example, a machine check architecture may perform step 104 as further described above. As used herein, "corrective action" may refer to any action taken to correct an error.

[0036] 5 shows an example flow diagram 500 of a more detailed method for detecting and reporting physical infrastructure constraint violations. This example method focuses on physical infrastructure constraints related to electrical design current constraints. Nevertheless, those skilled in the art will understand that the overall concepts and functionality of FIG. 5 may be applied in parallel to other types of physical infrastructure constraints, such as constraints defined in terms of voltage, power, temperature, etc.

[0037] The concept reflected in Figure 5 stems in part from an electrical design current constraint debugging procedure on a hardware accelerator. A corresponding system trips a thermal constraint due to exceeding the electrical design current limit when executing a workload through the hardware accelerator. In attempting to determine the root cause of this problem, it is useful to gain an understanding of the specific software components that cause the problem and correlate this problem to the electrical design current parameters. These investigations result in the system management unit communicating with the machine check architecture as an indication that the system is in a danger zone. Thus, when the machine check architecture receives an indication of a physical infrastructure constraint violation from the microcontroller, it can take appropriate predetermined actions. This method allows for greater control in response to such constraint violations and allows for triggering of investigations into the corresponding software through kernel logs or through the use of hardware debugging tools.

[0038] As one illustrative example, a system management unit may work with machine check architecture hardware to assist in debugging telemetry events caused by exceeding electrical design current limits or other physical infrastructure constraints. When the present technology is enabled, the microcontroller may provide signaling to the machine check architecture as a trigger for corresponding hardware components to take appropriate action. These actions may be simply logging that the event occurred and / or performing more active debugging steps in terms of writing specific data to memory, or performing debugging through the use of machine check breakpoints.

[0039] One illustrative example for applying the method corresponds to method 500 of FIG. 5. When this feature is enabled in step 502, the system management unit begins measuring an Electrical Design Current (EDC), which may correspond to a specific value representing the peak current that a physical component, such as a voltage regulator module, can handle in the short term. According to the methodology of method 500, if the value reaches a predetermined threshold (i.e., 254 in the example of FIG. 5) in comparison step 504, an internal counter may be incremented in step 506. The system management unit may loop these steps (see FIG. 5) until the internal counter reaches the predetermined threshold in comparison step 508. Once the threshold is reached, the system management unit generates a correctable machine check exception in step 510, which may be triggered in machine check bank 24 (see FIG. 4). Step 510 may correspond to a specific embodiment of step 104 of method 100 shown in FIG. 1, as further described above. Once a machine check exception is logged or recorded, all other counters may be reset in step 512 and the process may be repeated.

[0040] 5 focuses on violations of constraints related to electrical design current. Nevertheless, the present techniques may also be applied to other physical infrastructure constraints, such as constraints related to voltage (setpoint and load voltage), power, clock frequency, temperature, etc. In some examples, the computer code may use 4 bits to represent a clock frequency limit, a Boolean value to represent the current power state, 3 bits to represent the load voltage, and / or 5 bits for the setpoint voltage range.

[0041] 1, one or more of the systems described herein may perform corrective actions based on the triggering of a machine check exception in step 106. For example, particular software components (such as BIOS, firmware, operating system, and / or hypervisor) may perform corrective actions in response to the triggering of a machine check exception.

[0042] As outlined above, the machine check architecture provides a mechanism for hardware or low-level logic to detect violations of physical infrastructure constraints and report this detection to corresponding software components. In response to a machine check exception, the corresponding software component, such as an operating system, may perform one or more corrective actions. These corrective actions may include logging, reporting, or profiling the telemetry event. Logging a telemetry event may simply involve storing data and metadata describing the telemetry event in memory. Reporting a telemetry event may involve reporting the telemetry event to another software component or to a user or administrator, which may take further action to address the violation of a physical infrastructure constraint. Profiling a telemetry event, in some examples, may involve categorizing or classifying the telemetry event and identifying one or more causes or reasons for the telemetry event. Generally speaking, a telemetry event may refer to the reporting of a violation of a physical infrastructure constraint, as outlined above.

[0043] In some examples, the corrective action is performed by a prediction and prevention engine. In these examples, the machine check architecture in cooperation with one or more higher-level software components may collectively form a prediction and prevention engine. The prediction and prevention engine may be useful because it can monitor, identify, and correct physical infrastructure constraint violations even before these violations result in the failure of one or more physical components. Thus, the prediction and prevention engine may improve upon related machine check architecture configurations that are substantially limited to detecting hardware or other physical component failures (i.e., detecting failures of these physical components after they have already occurred). In contrast to such related methodologies, the present technology may identify violations of physical infrastructure constraints that tend to result in the eventual failure of the corresponding physical component, but may identify the violations before the actual failure of the component. As one illustrative example, the present technology may detect violations of electrical design current constraints and, further, detect the violations before the actual failure of the corresponding physical component, such as a conductor, lead, or voltage regulator. In other words, the present technology predicts that a physical or hardware component may eventually fail due to a violation of a physical infrastructure constraint, and, as further explained above, may prevent such a failure from actually occurring because the violation is detected early enough.

[0044] In some examples, corrective actions may be performed as part of debug diagnostics. For example, in a laboratory environment or in the field, the operation or execution of one or more software, firmware, or hardware components may result in an output that deviates from the intended design specifications. For example, during the development of a software, firmware, or hardware component, one or more “bugs” or undesired instances or unintended consequences of functionality may further result in the violation of corresponding physical infrastructure constraints. Accordingly, a developer may enter one or more debugging procedures to determine the root cause of the undesired functionality and / or to develop a remedy to eliminate the corresponding “bug.” As one example, an optimization bug or failure in a software, firmware, or hardware component may result in the violation of an electrical design current constraint. Other corresponding physical infrastructure constraints may include constraints defined for clock speed, voltage, temperature, or the like, as further described above. Accordingly, if a “bug” causes a violation of such a physical infrastructure constraint, the violation is detected by the present technology through a machine check architecture, thereby enabling one or more software components, an administrator, or a developer to perform corrective action. In some simple examples, the corrective action may correspond to entering a debugging mode or resetting the processor.

[0045] In summary, this application relates to a technique that may use an on-chip microcontroller, such as a system management unit, in conjunction with machine check architecture registers to support logging, reporting, and profiling of telemetry events within a computing environment. As further described above, examples of such telemetry events may include violations of electrical design current limits or constraints. This technique may be used as a fault prediction and prevention engine, which may be used to perform laboratory debugging or field diagnostics. This technique may also be used to take corrective action for critical events, such as entering a debugging mode, resetting the processor, or any other suitable and programmable action. When this technique's features are enabled, the microcontroller may provide a signal to the machine check architecture as a trigger for the corresponding hardware component to take appropriate fault prevention action or to record debugging data for diagnostics.

[0046] The present technology may improve upon related methodologies in various ways. As one example, a control loop in the context of thermal throttling may perform similar functions, but the control loop may lack any error reporting or any user-configurable corrective actions. Similarly, in some related methods, a microcontroller (e.g., a system management unit or platform security processor) may log a machine check architectural event due to a malfunction of the microcontroller itself (e.g., poisoned data consumption). Nevertheless, these additional cases differ from the improved technology described herein because they may involve logging and reporting physical defects or cosmic background radiation (e.g., soft error rate) events rather than reporting physical infrastructure constraint violations that are still not necessarily malfunctions and are not necessarily directly related to the operation of the corresponding microcontroller.

[0047] Embodiments described herein may relate to improving or modifying the configuration of machine check architectures to enable monitoring, reporting, detection, and prevention of physical infrastructure constraint violations (e.g., voltage, current, temperature, etc.) rather than simply detecting higher level errors such as data corruption. The associated machine check architectures do not necessarily check for lower level or physical level violations of constraints on values ​​such as current and temperature.

[0048] The examples described herein directly address the problem of correlating physical infrastructure values ​​with the micro-architectural or software context of modern SOCs and provide a solution that enables SOCs to avoid potentially catastrophic events that are the result of violations of physical infrastructure constraints. The present technology also supports scale debugging of such events by providing enhanced logging via existing scalable machine check architectures without system software intervention.

[0049] Various processors described herein may include and / or represent any type or form of hardware-implemented device capable of interpreting and / or executing computer-readable instructions. In one example, a processor may include and / or represent one or more semiconductor devices implemented and / or deployed as part of a computing system. Examples of processors include a central processing unit (CPU) and a microprocessor. Other examples may include a microcontroller, a field-programmable gate array (FPGA) implementing a soft-core processor, an application-specific integrated circuit (ASIC), a system on a chip (SoC), one or more portions thereof, one or more variations or combinations thereof, and / or any other suitable processor, depending on the context.

[0050] A processor may implement and / or be configured using any of a variety of different architectures and / or microarchitectures. For example, a processor may be implemented and / or configured as a Reduced Instruction Set Computer (RISC) architecture, or a processor may be implemented and / or configured as a Complex Instruction Set Computer (CISC) architecture. Additional examples of such architectures and / or microarchitectures include, without limitation, 16-bit computer architecture, 32-bit computer architecture, 64-bit computer architecture, x86 computer architecture, advanced RISC machine (ARM) architecture, microprocessor without interlocked pipelined stage (MIPS) architecture, scalable processor architecture (SPARC), load-store architecture, portions of one or more of these, combinations or variations of one or more of these, and / or any other suitable architecture or microarchitecture.

[0051] In some examples, a processor may include and / or incorporate one or more additional components not explicitly depicted and / or shown in the figures. Examples of such additional components include, but are not limited to, registers, memory devices, circuits, transistors, resistors, capacitors, diodes, connections, traces, buses, semiconductor (e.g., silicon) devices and / or structures, combinations or variations of one or more of these, and / or any other suitable components.

[0052] While the above disclosure describes various variations using specific block diagrams, flow diagrams, and examples, each block diagram component, flow diagram step, operation, and / or component described and / or shown herein may be individually and / or collectively implemented using a wide variety of hardware, software, or firmware (or any combination thereof) configurations. However, since many other architectures can be implemented to achieve the same functionality, any disclosure of components included within other components should be considered exemplary in nature.

[0053] The apparatuses, systems, and methods described herein may employ any number of software, firmware, and / or hardware configurations. For example, one or more of the example variations and / or embodiments disclosed herein may be encoded as a computer program (also referred to as computer software, a software application, computer-readable instructions, and / or computer control logic) on a computer-readable medium. The term "computer-readable medium" generally refers to any form of device, carrier, or medium capable of storing or carrying computer-readable instructions. Examples of computer-readable media include, but are not limited to, transmission-type media such as carrier waves, magnetic storage media (e.g., hard disk drives and floppy disks), optical storage media (e.g., compact disks (CDs) and digital video disks (DVDs)), electronic storage media (e.g., solid-state drives and flash media), and / or non-transitory media such as other distribution systems.

[0054] It should be noted that one or more of the modules, instructions, and / or microoperations described herein may transform data, physical devices, and / or representations of physical devices from one form to another. Additionally or alternatively, one or more of the modules, instructions, and / or microoperations described herein may transform a processor, volatile memory, non-volatile memory, and / or any other portion of a physical computing device from one form to another by executing on a computing device, storing data on a computing device, and / or otherwise interacting with a computing device.

[0055] The process parameters and order of steps described and / or illustrated herein are given by way of example only and can be changed as desired. For example, although the steps illustrated and / or described herein may be illustrated or described in a particular order, these steps do not necessarily have to be performed in the order illustrated or described. The various exemplary methods described and / or illustrated herein may omit one or more of the steps described or illustrated herein or may include additional steps in addition to those disclosed.

[0056] The foregoing description has been provided to enable those skilled in the art to best utilize various aspects of the exemplary variations disclosed herein. This exemplary description is not intended to be exhaustive or to be limited to any precise form disclosed. Many changes and modifications are possible without departing from the spirit and scope of the present disclosure. The variations disclosed herein should be considered in all respects as illustrative and not restrictive. Reference should be made to the appended claims and their equivalents in determining the scope of the present disclosure.

[0057] Unless otherwise specified, the terms "connected to" and "coupled to" (and their derivatives) as used in this specification and claims should be interpreted as allowing both direct and indirect connections (i.e., via other elements or components). Additionally, the terms "a" or "an" as used in this specification and claims should be interpreted as meaning "at least one of." Finally, for ease of use, the terms "including" and "having" (and their derivatives) as used in this specification and claims are interchangeable with the term "comprising," and have the same meaning.

Claims

1. the microcontroller reporting the detection of a violation of a physical infrastructure constraint to a machine check architecture; the machine check architecture, in response to the report, triggering a machine check exception such that a violation of the physical infrastructure constraint is logged; and performing a corrective action upon triggering the machine check exception. method.

2. the microcontroller includes a system management unit; 10. The method of claim 1.

3. The violation of the physical infrastructure constraint is defined in terms of at least one of voltage, power, sustained current, peak current, or temperature. The method of claim 2.

4. the violation of the physical infrastructure constraints includes a violation of an electrical design current constraint; 10. The method of claim 1.

5. the corrective action includes at least one of telemetry event logging, reporting, or profiling; 10. The method of claim 1.

6. The corrective action is performed by a prediction and prevention engine.

10. The method of claim 1.

7. The corrective action is performed as part of a debugging diagnostic.

10. The method of claim 1.

8. The corrective action includes entering a debugging mode or resetting the processor.

10. The method of claim 1.

9. Violations of physical infrastructure constraints are detected before said violations cause any physical component to fail; 10. The method of claim 1.

10. the microcontroller and the machine check architecture are disposed in a system-on-chip; 10. The method of claim 1.

11. a receptor for receiving reports of violations of physical infrastructure constraints from the microcontroller; a trigger, in response to receiving the report, for triggering a machine check exception such that a violation of the physical infrastructure constraint is logged; an initiator that initiates corrective action based on a record of a violation of the physical infrastructure constraint; Equipped with Machine check architecture.

12. the microcontroller includes a system management unit; 12. The machine check architecture of claim 11.

13. The violation of the physical infrastructure constraint is defined in terms of at least one of voltage, power, sustained current, peak current, or temperature.

13. The machine check architecture of claim 12.

14. the violation of the physical infrastructure constraints includes a violation of an electrical design current constraint; 12. The machine check architecture of claim 11.

15. the corrective action includes at least one of telemetry event logging, reporting, or profiling; 12. The machine check architecture of claim 11.

16. the machine check architecture is configured to initiate the corrective action as part of a prediction and prevention engine.

12. The machine check architecture of claim 11.

17. the machine check architecture is configured to initiate the corrective action as part of debugging diagnostics.

12. The machine check architecture of claim 11.

18. The corrective action includes entering a debugging mode or resetting the processor.

12. The machine check architecture of claim 11.

19. the microcontroller is configured to detect a violation of the physical infrastructure constraint before the violation causes any physical component to fail.

12. The machine check architecture of claim 11.

20. a microcontroller that issues a report of a violation of a physical infrastructure constraint; a machine check architecture; The machine check architecture comprises: In response to receiving the report, triggering a machine check exception such that a violation of the physical infrastructure constraint is logged; initiating corrective action based on the recorded violation of the physical infrastructure constraint; configured to: system.