Apparatus, system and method for detecting physical infrastructure constraint violations

By introducing a communication mechanism of microcontroller and machine inspection architecture in the computing system, the problem of inability to effectively detect physical infrastructure constraint violations in the prior art is solved, and early detection and recording of these violations is achieved, and the reliability of the system is improved.

CN119948463APending Publication Date: 2025-05-06ADVANCED MICRO DEVICES INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380068888.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-09-30
Filing Date
2023-09-29
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Modern computing systems have shortcomings in detecting lower-level violations of physical infrastructure constraints, and existing machine inspection architectures and exceptional functions may not be able to effectively monitor and report these violations.

Method used

By establishing a communication mechanism between the microcontroller and the machine inspection architecture, the microcontroller reports detection results that violate physical infrastructure constraints, the machine inspection architecture triggers the machine inspection exception, and records and performs correction actions.

Benefits of technology

Early detection and recording of physical infrastructure constraint violations is achieved, potentially catastrophic events before hardware failure occurs due to constraint violations, and system reliability and stability are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119948463A_ABST
    Figure CN119948463A_ABST
Patent Text Reader

Abstract

A disclosed method may include (i) reporting, by a microcontroller, a detection of violation of a physical infrastructure constraint to a machine inspection architecture, (ii) in response to the reporting, triggering, by the machine inspection architecture, a machine inspection anomaly such that the violation of the physical infrastructure constraint is recorded, and (iii) performing a corrective action based on the triggering of the machine inspection anomaly. Various other apparatuses, systems, and methods are also disclosed.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Modern computer chips may implement many different features to detect errors and facilitate debugging.Within this context, the present application discloses improved apparatus, systems, and methods for detecting physical infrastructure constraint violations. BRIEF DESCRIPTION OF THE DRAWINGS

[0002] The accompanying drawings illustrate several exemplary variations and are a part of the specification. Together with the following description, these drawings illustrate and explain the various principles of the present disclosure.

[0003] Figure 1 A flow chart of an exemplary method for detecting physical infrastructure constraint violations is shown.

[0004] Figure 2 A block diagram of an exemplary system-on-chip is shown.

[0005] Figure 3 A block diagram illustrating classification of detected and undetected errors within a computing environment.

[0006] Figure 4 A block diagram of an illustrative example of a machine check architecture register is shown.

[0007] Figure 5 A flow chart of an exemplary method for detecting physical infrastructure constraint violations in one variation is shown.

[0008] In all drawings, the same reference numerals and descriptions indicate similar but not necessarily identical elements. Although the exemplary variations described herein are susceptible to various modifications and alternative forms, specific variations have been shown by way of example in the drawings and will be described in detail herein. However, the exemplary variations described herein are not intended to be limited to the specific forms disclosed. More specifically, the present disclosure covers all modifications, equivalents, and alternatives that fall within the scope of the appended claims. DETAILED DESCRIPTION

[0009] The present disclosure describes various devices, systems and methods for detecting physical infrastructure constraint violations. Modern computing microprocessors can feature a machine check architecture and corresponding machine check exception functionality that can report errors to an operating system or other software component. These mechanisms can detect and report hardware or machine errors, such as system bus errors, errors caused by communications over noisy channels, parity errors, cache errors, and translation lookaside buffer errors. However, these related methods may not necessarily detect lower-level violations of physical infrastructure constraints. In other words, physical components or hardware components of modern computing systems can be characterized by physical infrastructure constraints based on voltage, power, continuous current, peak current, or temperature, etc.

[0010] More generally, as used herein, the term "physical infrastructure constraint" generally refers to a constraint defined according to a physical or performance-related property of a corresponding physical component or hardware component of a computing device such that the property should not satisfy a threshold (e.g., a maximum or minimum threshold) or should not satisfy a threshold for more than a certain amount of time, as discussed further below. Thus, these components should not achieve values ​​that violate these constraints (e.g., violate the constraints for more than a predetermined amount of time). In some examples, a physical infrastructure constraint may be violated even if the violation has not yet caused the corresponding physical component or hardware component to fail.

[0011] As an illustrative example, a particular wire or lead may have an electrical design current constraint that specifies a maximum value for the current that traverses the wire or lead. Thus, the wire or lead should not carry a current in excess of that maximum value or should not carry a current in excess of that maximum value for a predetermined amount of time, yet the wire or lead may continue to carry such current while violating the constraint without the wire or lead failing. However, within a sufficient amount of time, physical components will tend to fail, which is one purpose behind the physical infrastructure constraints that are imposed. This also presents another reason why monitoring for violations of such constraints is important, yet modern machine check architectures and corresponding machine check exception functionality may not have visibility into such violations.

[0012] As will be described in more detail below, the present disclosure generally relates to apparatus, systems, and methods for detecting physical infrastructure constraint violations. In one example, the method may include (i) reporting, by a microcontroller, detection of a violation of a physical infrastructure constraint to a machine check architecture, (ii) in response to the reporting, triggering, by the machine check architecture, a machine check exception such that the violation of the physical infrastructure constraint is recorded, and (iii) performing a corrective action based on triggering the machine check exception.

[0013] In some examples, the microcontroller includes a system management unit.

[0014] In some examples, violating a physical infrastructure constraint is defined in terms of at least one of voltage, power, continuous current, peak current, or temperature.

[0015] In some examples, violating a physical infrastructure constraint includes violating an electrical design current constraint.

[0016] In some examples, the corrective action includes at least one of logging, reporting, or analyzing the telemetry event.

[0017] In some examples, corrective actions are performed by a prediction and prevention engine.

[0018] In some examples, the corrective action is performed as part of a debug diagnostic.

[0019] In some examples, the corrective action includes entering a debug mode or resetting the processor.

[0020] In some examples, violations of physical infrastructure constraints are detected before the violations cause any physical components to fail.

[0021] In some examples, the microcontroller and the machine check architecture are provided on a system on a chip.

[0022] In another example, a machine check architecture includes (i) a receiver that receives a report of a violation of a physical infrastructure constraint from a microcontroller, (ii) a trigger that triggers a machine check exception in response to receiving the report, such that a violation of the physical infrastructure constraint is recorded, and (iii) an initiator that initiates a corrective action based on the recording of the violation of the physical infrastructure constraint.

[0023] In another example, a system may include: a microcontroller that issues a report of a violation of a physical infrastructure constraint and triggers a machine check exception in response to receiving the report, such that a machine check architecture records the violation of the physical infrastructure constraint and initiates corrective action based on the recording of the violation of the physical infrastructure constraint.

[0024] Figure 1 A flow chart of a method 100 for detecting physical infrastructure constraint violations is shown. The steps of the method 100 may be performed by various devices, microcontrollers, machine check architectures, and / or systems described herein. For example, at step 102, the microcontroller may report the detection of a violation of a physical infrastructure constraint to the machine check architecture. As used herein, the term "microcontroller" generally refers to a compact integrated circuit designed to manage operations within, for example, an embedded system.

[0025] To help provide context for the context in which method 100 may be performed, Figure 2 An exemplary block diagram of a system on a chip is shown. The system on a chip may include both a microcontroller and a machine inspection architecture discussed further below.

[0026] In an exemplary variation, the SOC 200 includes a plurality of processing cores 220, 222, 224, 226, a plurality of trace data storage elements 240, 242, 244, 246, 254, 264, 274 (e.g., L2 cache and trace capture buffers (TCBs)), a north bridge 250 (or memory controller), a south bridge 260, a GPU 270, and a cross trigger bus 280. Cross triggering of a debug state machine (not illustrated) on another die within the same package (e.g., in an MCM) and / or within other packages may be implemented via an off-chip debug state machine interface 212.

[0027] Although SOC 200 is illustrated as including four cores 220, 222, 224, 226, in other variations (including as few as a single core), the SOC may include more or fewer cores. In addition, although SOC 200 is illustrated as including a single north bridge 250, south bridge 260, and GPU 270, in other variations, some or all of these electronic modules may be excluded from SOC 200 (e.g., they may be located off-chip). Furthermore, although SOC 200 is illustrated as including only one north bridge 250, in other variations, the SOC may include more than one north bridge. In other variations, in addition to the north bridge 250, the south bridge 260 may include more than one north bridge. Figure 2 In addition to the processing components and buses illustrated in the figure, the SOC may include additional or different processing components, buses, and other electronic devices and circuits.

[0028] Processing cores 220, 222, 224, 226 generally represent the main processing hardware, logic and / or circuitry for SOC 200, and each processing core 220, 222, 224, 226 may be implemented using one or more arithmetic logic units (ALUs), one or more floating point units (FPUs), one or more memory elements (e.g., one or more caches), discrete gate or transistor logic, discrete hardware components, or any combination thereof. Figure 2 Not illustrated in , but each processing core 220 , 222 , 224 , 226 may implement its own associated cache memory element (eg, a level one or L1 cache) proximate to its respective processing circuitry for reduced latency.

[0029] Northbridge 250 (which may also be referred to as a "memory controller" in some systems) is configured to communicate with I / O peripherals (e.g., I / O peripherals 140, Figure 1 ) and memory (e.g., memory 150, Figure 1 ) interface. Northbridge 250 controls communication between components of SOC 200 and I / O peripherals and / or external memory. In some systems, southbridge 260, which may also be referred to as an "I / O controller hub", is configured to connect and control peripherals (e.g., relatively slow peripherals). GPU 270 is a dedicated microprocessor that offloads and accelerates graphics rendering from cores 220, 222, 224, 226.

[0030] In the illustrated variation, caches 240, 242, 244, 246 and TCBs 254, 264, 274 provide intermediate memory elements of reduced size relative to external memory for temporarily storing data and / or instructions retrieved from external memory or elsewhere, and / or data generated by processing cores 220, 222, 224, 226, north bridge 250, south bridge 260, and GPU 270. For example, in one variation, caches 240, 242, 244, 246 and TCBs 254, 264, 274 provide memory elements for storing debug information (or "trace data") collected and / or generated by a debug state machine during debug operations associated with the respective electronic modules with which they are integrated. In the illustrated variation, caches 240, 242, 244, 246 are in close proximity and coupled between respective processing cores 220, 222, 224, 226 and north bridge 250. In this regard, caches 240, 242, 244, 246 may alternatively be referred to as core-coupled caches, and each core-coupled cache 240, 242, 244, 246 maintains data and / or program instructions previously fetched from external memory that were previously used and / or may be used by its associated processing core 220, 222, 224, 226. Caches 240, 242, 244, 246 are preferably larger than the L1 caches implemented by processing cores 220, 222, 224, 226 and serve as secondary caches (or L2 caches) in the memory hierarchy. The SOC 200 may also include another higher level cache (eg, a level three or L3 cache, not illustrated) that is preferably larger than the L2 caches 240 , 242 , 244 , 246 .

[0031] In an exemplary variation, SOC 200 includes a test interface 210 that includes a plurality of pins dedicated to testing and / or configuring functionality of SOC 200. In one variation, test interface 210 complies with the IEEE 1149.1 Standard Test Access Port and Boundary Scan Architecture, the Joint Test Action Group (JTAG) standard.

[0032] Figure 2Also illustrated is how the SOC 200 may include a subsystem 288 for performing the method 100. The subsystem 288 may also include a machine check architecture 290 (MCA). The machine check architecture 290 may include a receiver 292 that may perform step 102 of the method 100 and a trigger 294 that may perform step 104 of the method 100, as discussed further below. In addition, the subsystem 288 may also include an initiator 296 for performing step 106, as discussed further below. The subcomponents 292-296 may be implemented according to hardware or logic within the SOC 200, and / or one or more of these may be implemented as firmware components or software components coordinated with a hardware-based MCA.

[0033] As discussed above, the techniques of the present application may modify or improve the machine check architecture and the corresponding machine check exception functionality. Accordingly, the following provides an overview of the machine check architecture and its corresponding machine check exception functionality.

[0034] An MCA, such as MCA 290, may refer to a processor-centric error detection and reporting mechanism. Figure 3 A useful classification 300 is shown that breaks down system errors 302 into sets. The two main sets are detected errors 304 and undetected errors 306. Undetected errors, although they may be important, are not handled by the MCA. Undetected errors may include benign errors 312 and serious errors 314.

[0035] Detected errors can be divided into corrected errors 308 and uncorrected errors 310. Corrected errors are benign after being corrected. Although benign, these errors can still be reported to the software for tracking purposes. Some errors should indicate that the errors are expected, and therefore monitoring them is helpful. Uncorrected errors can include three different sets: catastrophic errors 316, fatal errors 318, and recoverable errors 320. Catastrophic errors and fatal errors all trigger system resets. Uncorrected recoverable errors (UCR) may allow the system to still function, and these can be classified into uncorrected no-action errors 322, software recoverable action optional errors 324, and software recoverable action required errors 326.

[0036] Given these error sets, functionality has been created to alert software (such as firmware, BIOS, operating system, or hypervisor) of the errors when detected by hardware components. Remedial actions may include two versions, logging and reporting. Logging may be performed by logging data into a machine check (MC) register, such as Figure 4 As shown, the data can be reviewed by software and saved to memory. Figure 4As further shown in , MC registers 400 include a set of global registers 402 and a set of grouped registers 404-410. Reports can use CMCI and MCE reports. The CMCI (corrected machine check interrupt) function can report corrected errors (CE) and uncorrected recoverable errors (UCR), such as uncorrected no action (UCNA) recoverable errors. MCE (machine check exception) reporting can address uncorrected recoverable errors, such as software recoverable action optional (SRAO) errors and software recoverable action required (SRAR) errors listed above. This function can also report catastrophic errors involving restarting the system. In the scenario outlined above, data about the event can be saved to the MC library before reporting.

[0037] Back to Figure 1 In the method 100, step 102 may be performed by a microcontroller that detects a violation of a physical infrastructure constraint. As discussed further above, a physical infrastructure constraint may refer to a maximum value or threshold value that defines a property of a physical component or hardware component that should be avoided. Thus, illustrative examples of such physical infrastructure constraints may include constraints based on voltage, power, continuous current, peak current, or temperature.

[0038] In one illustrative example, the microcontroller may correspond to a system management unit (SMU). The system management unit may be a subcomponent of a larger processor (e.g., a SOC), and this subcomponent may be responsible for a variety of system and power management tasks during boot and run time. The system management unit may include a thermal block, which may also include features related to temperature sensing, control, and reporting. The thermal block may include temperature collection and calculation logic, fan speed control for off-chip fans, and temperature reporting functions. The system management unit may also include a plurality of registers, which may provide the current control temperature among other outputs.

[0039] Back to Figure 1 At step 104, one or more of the devices or systems described herein may trigger a machine check exception such that a violation of a physical infrastructure constraint is recorded. For example, as further described above, a machine check architecture may perform step 104. As used herein, a "corrective action" may refer to any action performed to correct an error.

[0040] Figure 5 An exemplary flow chart 500 of a more detailed method for detecting and reporting physical infrastructure constraint violations is shown. The example of the method focuses on physical infrastructure constraints in terms of electrical design current constraints. However, those skilled in the art will appreciate that Figure 5 The overall concepts and capabilities of can be applied in parallel to other types of physical infrastructure constraints, such as constraints defined in terms of voltage, power, temperature, etc.

[0041] Figure 5 The concepts reflected in are generated in part by an electrical design current constraint debugger on a hardware accelerator. When a workload is run through a hardware accelerator, the corresponding system breaks the temperature constraint due to crossing the electrical design current limit. In an attempt to identify the root cause of the problem, it is helpful to gain an understanding of the specific software components that cause the problem and associate the problem with the electrical design current parameters. The results of these investigations lead to the system management unit communicating with the machine check architecture as an indication that the system is in a dangerous area. Therefore, once the machine check architecture obtains an indication of a physical infrastructure constraint violation from the microcontroller, appropriate predefined actions can be taken. The method enables greater control in response to such constraint violations, as well as allowing triggers to be studied through kernel logs or through the use of hardware debugging tools.

[0042] As an illustrative example, a system management unit may work in conjunction with machine check architecture hardware to help debug telemetry events due to crossing electrical design current limits or other physical infrastructure constraints. When the techniques of the present application are enabled, the microcontroller may provide signaling to the machine check architecture as a trigger for the corresponding hardware component to take appropriate actions. These actions may simply log the occurrence of the event and / or perform more active debugging steps in terms of writing specific data to memory or performing debugging by using machine check breakpoints.

[0043] An illustrative example of applying this method corresponds to Figure 5 When this feature is enabled at step 502, the system management unit begins measuring the electrical design current (EDC), which may correspond to a specific value representing the peak current that a physical component (such as a voltage regulator module) can handle on a short-term basis. According to the method 500, if this value reaches a predetermined threshold (i.e., Figure 5 254 in the example of ), the internal counter may be incremented at step 506. The system management unit may loop through these steps (see Figure 5 ) until the internal counter reaches a predetermined threshold at comparison step 508. Once the threshold is reached, the system management unit generates a correctable machine check exception at step 510, which can be triggered in the machine check library 24 (see Figure 4 ). Step 510 may correspond to Figure 1 A specific implementation of step 104 of method 100 is shown in , as discussed further above. Once a machine check exception is logged, all other counters may be reset at step 512 and the process may be repeated.

[0044] As discussed further above, Figure 5The flowchart of the present invention focuses on the violation of the constraints in the electrical design current. However, the technology of the present application can also be applied to other physical infrastructure constraints, such as voltage (set voltage and load voltage), power, clock frequency and temperature. In some exemplary examples, the computer code can use four bits to represent the clock frequency limit value, use a Boolean value to represent the current power state, use three bits to represent the load voltage, and / or use five bits as the range of the set voltage.

[0045] Back to Figure 1 At step 106, one or more of the systems described herein may perform corrective actions based on the triggering of a machine check exception. For example, certain software components (BIOS, firmware, operating system, and / or hypervisor) may perform corrective actions in response to the triggering of a machine check exception.

[0046] As outlined above, the machine check architecture provides a mechanism for hardware or low-level logic to detect violations of physical infrastructure constraints and report the detection to corresponding software components. In response to a machine check exception, a corresponding software component such as an operating system may perform one or more corrective actions. These corrective actions may include recording, reporting, or analyzing telemetry events. Recording telemetry events may simply involve storing data and metadata describing the telemetry events in a memory. Reporting telemetry events may involve reporting the telemetry events to another software component or to a user or administrator, which may take further action to resolve the violation of the physical infrastructure constraints. In some examples, analyzing telemetry events may involve categorizing or classifying telemetry events, and ascertaining one or more causes of the telemetry events. In general, a telemetry event may refer to reporting a violation of a physical infrastructure constraint, as further outlined above.

[0047] In some examples, the corrective action is performed by a prediction and prevention engine. In these examples, a machine check architecture coordinated with one or more higher-level software components can collectively form a prediction and prevention engine. The prediction and prevention engine can be useful because the engine can monitor, identify and remedy these violations even before the violation of physical infrastructure constraints causes one or more physical components to fail. Therefore, the prediction and prevention engine can improve the relevant machine check architecture configuration, which is effectively limited to detecting hardware or other physical component failures (i.e., detecting these physical component failures after these physical component failures have occurred). In contrast to such related methods, the technology of the present application can identify the violation of physical infrastructure constraints that will tend to cause the final failure of the corresponding physical component, but the violation can be identified before the actual failure of the component. As an illustrative example, the technology of the present application can detect the violation of electrical design current constraints, but the violation can be detected before the actual failure of the corresponding physical component such as a wire, a lead or a voltage regulator. In other words, the technology of the present application can predict that a physical or hardware component may eventually fail due to a violation of a physical infrastructure constraint, and can prevent such failures from actually occurring because the violation is detected sufficiently early, as further discussed above.

[0048] In some examples, the corrective action may be performed as part of the debug diagnostics. For example, in a laboratory environment or in the field, the operation or execution of one or more software components, firmware components, or hardware components may result in an output that deviates from the expected design specification. For example, during the development of a software component, firmware component, or hardware component, one or more "bugs" or unexpected instances or unexpected results of a function may further result in a violation of the corresponding physical infrastructure constraints. Therefore, the developer may enter one or more debuggers in an attempt to find out the root cause of this unexpected functionality and / or develop a remedy to eliminate the corresponding "bug". As an illustrative example, an optimized bug or failure in a software component, firmware component, or hardware component may result in a violation of an electrical design current constraint. Other corresponding physical infrastructure constraints may include constraints defined according to clock speed, voltage, or temperature, etc., as further discussed above. Therefore, when a "bug" results in a violation of such physical infrastructure constraints, such a violation may be detected by the technology of the present application through a machine check architecture, thereby enabling one or more software components, administrators, or developers to perform corrective actions. In some simple examples, the corrective action may correspond to entering a debug mode or resetting a processor.

[0049] In summary, the present application relates to techniques that can use an on-chip microcontroller (such as a system management unit) in conjunction with a machine check architecture register to facilitate the recording, reporting, and analysis of telemetry events within a computing environment. As further discussed above, illustrative examples of such telemetry events may include violations of electrical design current limits or constraints. The technology can be used as a failure prediction and prevention engine that can be used to perform laboratory debugging or field diagnostics. The technology can also be used to take corrective actions for critical events, such as entering a debug mode, resetting a processor, or any other suitable and programmable action. When features of this technology are enabled, the microcontroller can provide signaling to the machine check architecture as a trigger for the corresponding hardware component to take appropriate failure prevention actions or record debug data for diagnosis.

[0050] The techniques of the present application may improve related methods in various ways. As an example, a control loop in the context of thermal suppression may perform similar functions, but the control loop may lack any error reporting or any user-configurable corrective actions. Similarly, in some related methods, a microcontroller (e.g., a system management unit or a platform security processor) may record machine check architecture events caused by failures of the microcontroller itself (e.g., poisoned data consumption). However, these additional situations are different from the improved techniques described herein because these situations may involve recording and reporting physical defects or cosmic background radiation (e.g., soft error rate) events rather than reporting violations of physical infrastructure constraints, which are not necessarily failures, but which are not necessarily directly related to the operation of the corresponding microcontroller.

[0051] Specific implementations described within this application may involve improving or modifying the configuration of a machine check architecture to enable monitoring, reporting, detection, and prevention of violations of physical infrastructure constraints (e.g., voltage, current, temperature, etc.), rather than just detecting high-level errors such as data corruption. The associated machine check architecture does not necessarily check for lower or physical level violations of constraints in terms of values ​​such as current and temperature.

[0052] The examples described in this application directly address the difficulty of relating physical infrastructure values ​​to the microarchitecture or software context of modern SOCs, and provide a solution that will enable the SOC to avoid potentially catastrophic events as a result of physical infrastructure constraint violations. The techniques of this application will also aid in extended debugging of such events by providing enhanced logging via existing scalable machine inspection architectures without any system software intervention.

[0053] The various processors described herein may include and / or represent a device that can interpret and / or implement any type or form of hardware implementation of computer-readable instructions. In one example, a processor may include and / or represent one or more semiconductor devices that are implemented and / or deployed as a part of a computing system. An example of a processor includes a central processing unit (CPU) and a microprocessor. Depending on the context, other examples may include a microcontroller, a field programmable gate array (FPGA) implementing a softcore processor, an application specific integrated circuit (ASIC), a system on a chip (SOC), a portion of one or more of them, a variation or combination of one or more of them, and / or any other suitable processor.

[0054] The processor can be implemented and / or configured with any of a variety of different architectures and / or micro-architectures. For example, the processor can be implemented and / or configured as a reduced instruction set computer (RISC) architecture, or the processor can be implemented and / or configured as a complex instruction set computer (CISC) architecture. Additional examples of such architectures and / or micro-architectures include, but are not limited to, 16-bit computer architectures, 32-bit computer architectures, 64-bit computer architectures, x86 computer architectures, advanced RISC machine (ARM) architectures, microprocessors without interlocked pipeline stages (MIPS) architectures, scalable processor architectures (SPARC), load-store architectures, parts of one or more of them, combinations or variations of one or more of them, and / or any other suitable architecture or micro-architecture.

[0055] In some examples, the processor may include and / or incorporate one or more additional components not explicitly represented and / or illustrated in the figures. Examples of such additional components include, but are not limited to, registers, memory devices, circuits, transistors, resistors, capacitors, diodes, connections, traces, buses, semiconductor (e.g., silicon) devices and / or structures, combinations or variations of one or more of them, and / or any other suitable components.

[0056] Although the foregoing disclosure uses specific block diagrams, flow charts, and examples to illustrate various variations, each block diagram component, flow chart step, operation, and / or component described and / or illustrated herein may be implemented individually and / or collectively using a wide range of hardware, software, or firmware (or any combination thereof) configurations. In addition, any disclosure of components contained within other components should be considered exemplary in nature, as many other architectures may be implemented to achieve the same functionality.

[0057] The devices, systems, and methods described herein may employ any number of software, firmware, and / or hardware configurations. For example, one or more exemplary variations and / or specific implementations disclosed herein may be encoded as a computer program (also referred to as computer software, software applications, computer-readable instructions, and / or computer control logic) on a computer-readable medium. The term "computer-readable medium" generally refers to any form of device, carrier, or medium capable of storing or carrying computer-readable instructions. Examples of computer-readable media include, but are not limited to, transmission-type media such as carrier waves and non-transient media such as magnetic storage media (e.g., hard drives and floppy disks), optical storage media (e.g., compact disks (CDs) and digital video disks (DVDs)), electronic storage media (e.g., solid-state drives and flash memory media), and / or other distribution systems.

[0058] In addition, one or more of the modules, instructions, and / or micro-operations described herein may transform data, physical devices, and / or representations of physical devices from one form to another. Additionally or alternatively, one or more of the modules, instructions, and / or micro-operations described herein may transform a processor, volatile memory, non-volatile memory, and / or any other portion of a physical computing device from one form to another by executing on a computing device, storing data on a computing device, and / or otherwise interacting with a computing device.

[0059] The order of process parameters and steps described and / or illustrated herein is given by way of example only and may vary as desired. For example, although the steps shown and / or illustrated herein can be shown or discussed in a particular order, these steps do not necessarily need to be performed in the order shown or discussed. The various exemplary methods described and / or illustrated herein may also omit one or more steps described or illustrated herein, or include additional steps in addition to those disclosed.

[0060] The previous description has been provided to enable others skilled in the art to best utilize various aspects of the exemplary variations disclosed herein. This exemplary description is not intended to be exhaustive or limited to any precise form disclosed. Many modifications and variations are possible without departing from the spirit and scope of the present disclosure. The variations disclosed herein should be considered to be exemplary and non-restrictive in all respects. In determining the scope of the present disclosure, reference should be made to the appended claims and their equivalents.

[0061] Unless otherwise indicated, the terms "connected to" and "coupled to" (and their derivatives) as used in the specification and claims will be deemed to allow both direct and indirect (i.e., via other elements or components) connections. In addition, the terms "a" or "an" as used in the specification and claims will be deemed to mean "at least one". Finally, for ease of use, the terms "including" and "having" (and their derivatives) as used in the specification and claims are interchangeable with the word "comprising" and have the same meaning.

Claims

1. A method, comprising: Detection of violations of physical infrastructure constraints is reported by the microcontroller to the machine check architecture; In response to the report, triggering a machine check exception by the machine check architecture such that the violation of the physical infrastructure constraint is recorded; and Corrective action is performed based on the triggering of the machine check exception.

2. The method of claim 1, wherein the microcontroller comprises a system management unit.

3. The method of claim 2, wherein the violation of the physical infrastructure constraint is defined in terms of at least one of voltage, power, continuous current, peak current, or temperature. The method of claim 1 , wherein the violating the physical infrastructure constraint comprises violating an electrical design current constraint.

5. The method of claim 1, wherein the corrective action comprises at least one of logging, reporting, or profiling a telemetry event. The method of claim 1 , wherein the corrective action is performed by a prediction and prevention engine. The method of claim 1 , wherein the corrective action is performed as part of a debug diagnostic.

8. The method of claim 1, wherein the corrective action comprises entering a debug mode or resetting a processor.

9. The method of claim 1, wherein a violation of a physical infrastructure constraint is detected before the violation causes any physical component to fail.

10. The method of claim 1, wherein the microcontroller and the machine check architecture are provided on a system on a chip.

11. A machine inspection architecture, the machine inspection architecture comprising: a receiver that receives reports of violations of physical infrastructure constraints from the microcontroller; a trigger that triggers a machine check exception in response to receiving the report such that the violation of the physical infrastructure constraint is recorded; and An initiator that initiates a corrective action based on the recording of the violation of the physical infrastructure constraint.

12. The machine inspection architecture of claim 11, wherein the microcontroller comprises a system management unit.

13. The machine check architecture of claim 12, wherein the violation of the physical infrastructure constraint is defined in terms of at least one of voltage, power, continuous current, peak current, or temperature.

14. The machine check architecture of claim 11, wherein the violation of the physical infrastructure constraint comprises a violation of an electrical design current constraint.

15. The machine inspection architecture of claim 11, wherein the corrective action comprises at least one of logging, reporting, or profiling a telemetry event.

16. The machine check architecture of claim 11, wherein the machine check architecture is configured to initiate the corrective action as part of a prediction and prevention engine.

17. The machine check architecture of claim 11, wherein the machine check architecture is configured to initiate the corrective action as part of a debug diagnostic.

18. The machine check architecture of claim 11, wherein the corrective action comprises entering a debug mode or resetting a processor.

19. The machine check architecture of claim 11, wherein the microcontroller is configured to detect the violation of the physical infrastructure constraint before the violation causes any physical component to fail.

20. A system, comprising: a microcontroller that issues a report of a violation of a physical infrastructure constraint; A machine check architecture, the machine check architecture: triggering a machine check exception in response to receiving the report such that the violation of the physical infrastructure constraint is logged; and Corrective action is initiated based on the recording of the violation of the physical infrastructure constraint.