Hardware control update of physical operating parameters for field failure detection

By dynamically adjusting operating parameters after hardware component deployment for testing, the problems of low hardware component fault detection efficiency and frequent system restarts in existing technologies are solved, achieving fast and accurate fault detection and extended lifespan.

CN114902190BActive Publication Date: 2025-12-23NVIDIA CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202180007988.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-03-20
Filing Date
2021-03-19
Publication Date
2025-12-23
Estimated Expiration
2041-03-19

AI Technical Summary

Technical Problem

Existing technologies struggle to efficiently detect potential faults after hardware components are deployed, and frequent system restarts reduce reliability, making it impossible to effectively identify the degradation rate and fault boundaries of hardware components.

Method used

By adjusting physical operating parameters, such as power supply voltage, clock speed, and noise level, without restarting the computing platform, the IST controller dynamically updates these parameters to test hardware components, reducing testing time and identifying fault boundaries.

Benefits of technology

It enables rapid detection of potential hardware component failures without affecting system reliability, reduces testing time and system restarts, and improves the lifespan of hardware components and the accuracy of fault detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114902190B_ABST
    Figure CN114902190B_ABST
Patent Text Reader

Abstract

When values of physical operating parameters can be changed without rebooting the computing platform, latency of in-system test (IST) execution of hardware components of a field (deployed) computing platform can be reduced. Tests (e.g., patterns or vectors) are executed for changed values of physical operating parameters (e.g., supply voltage, clock speed, temperature, noise magnitude / duration, operating current, etc.) to provide the ability to detect faults in hardware components.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to in-system test (IST) execution of field hardware components, and more specifically to hardware-controlled updates of physical operating parameters for field failure detection. BACKGROUND

[0002] Computing devices, such as chips and other circuitry, are typically tested by manufacturers prior to deployment in the field in order to verify that they are working properly and to detect manufacturing defects. For example, a computing device can be tested prior to deployment using automated test equipment (ATE). However, due to various factors (e.g., environmental hazards, aging, etc.), some devices develop defects after deployment, and in many applications it is important to have field failure detection capabilities. For example, autonomous functional safety requirements specify that a component will adhere to a fault tolerance time interval (FTTI) of 100 milliseconds, which represents the allowed time between the occurrence of a permanent fault and the execution of a remedial action.

[0003] Conventionally, in-system test (IST) can be used to detect the occurrence of a permanent fault when it occurs in order to adhere to the FTTI. However, while a hardware component, such as an integrated circuit (IC), can be tested by IST, there can be potential defects in the component or connections to the component that develop over time. In order to better detect when a hardware component is approaching failure or to estimate a rate of degradation, IST is performed for different supply voltage values to identify a minimum operating supply voltage (Vmin) for the hardware component. Each time the value is changed for IST, a time-consuming system restart is required. Additionally, restarting the system generally decreases the reliability of the entire system due to power cycling. There is a need to address these issues and / or other issues associated with the prior art. SUMMARY

[0004] When the value of a physical operating parameter can be changed without restarting a computing platform, the latency of in-system test (IST) execution of a hardware component of a field (deployed) computing platform can be reduced. A test (e.g., a pattern or vector) is performed for a changed value of a physical operating parameter (e.g., supply voltage, clock speed, temperature, noise size / duration, operating current, etc.), thereby providing the ability to detect a fault in the hardware component.

[0005] A method, computer-readable medium, and system for hardware control updates of physical operating parameters for in-field fault detection are disclosed. In one embodiment, the method includes executing at least a portion of a test on a hardware component of a field computing platform to produce a first test result, wherein a first value is used for a physical operating parameter applied to the hardware component during execution of the test, storing, by an IST controller within the field computing platform, the first test result in a memory, updating the physical operating parameter to use a second value based on the first test result in response to a command generated by the IST controller, and resuming execution of the test on the hardware component using the second value for the physical operating parameter to produce a second test result.

[0006] In one embodiment, the physical operating parameter is at least one of: a supply voltage, a supply current, a clock speed, a noise magnitude, a noise duration, or a temperature.

[0007] In one embodiment, the second value is determined by the IST controller. In another embodiment, the second value is determined by a central processing unit (CPU) coupled to the hardware component.

[0008] In one embodiment, the first test result indicates a pass, and the method further includes determining that the second value is a cutoff value for the physical operating parameter when the second test result indicates a fail.

[0009] In one embodiment, the test includes at least one of a permanent fault test or a functional test. In one embodiment, the test includes a permanent fault test represented as one or more structural vectors. In another embodiment, the test includes a functional test represented as one or more functional vectors.

[0010] In one embodiment, the method includes waiting for a predetermined duration after resuming execution of the test on the hardware component before checking the second test result, and rebooting the field computing platform when the predetermined duration expires and execution of the test is not complete.

[0011] In one embodiment, the method includes updating the physical operating parameter to use the second value includes transmitting the command from the IST controller to an external component that provides the physical operating parameter to the hardware component.

[0012] In one embodiment, a system includes an in-system test (IST) controller within a field computing platform and coupled to a memory, the IST controller configured to: execute at least a portion of a test on a hardware component of the field computing platform to produce a first test result, wherein a first value is used for a physical operating parameter of the hardware component applied during execution of the test, the IST controller storing the first test result in the memory, updating the physical operating parameter to use a second value based on the first test result in response to a command generated by the IST controller, and resuming execution of the test on the hardware component using the second value of the physical operating parameter to produce a second test result.

[0013] In one embodiment, the system includes at least one of an autonomous or semi-autonomous vehicle, an autonomous or semi-autonomous machine, an autonomous or semi-autonomous robot, an autonomous or semi-autonomous industrial robot, a manned or unmanned aircraft, or a manned or unmanned watercraft.

[0014] In one embodiment, the system includes at least one of a computing server system, a data center, a system on a chip (SoC), or an embedded system.

[0015] In one embodiment, a non-transitory computer-readable medium stores computer instructions that, when executed by one or more processors, cause the one or more processors to perform the steps of: executing at least a portion of a test on a hardware component of a field computing platform to produce a first test result, wherein a first value is used for a physical operating parameter of the hardware component applied during execution of the test, storing the first test result in a memory by an in-system test (IST) controller within the field computing platform, updating the physical operating parameter to use a second value based on the first test result in response to a command generated by the IST controller, and resuming execution of the test on the hardware component using the second value of the physical operating parameter to produce a second test result. BRIEF DESCRIPTION OF DRAWINGS

[0016] Figure 1A A block diagram of a system is shown in accordance with one embodiment.

[0017] Figure 1B A flowchart of a method for performing in-system testing is shown in accordance with one embodiment.

[0018] Figure 2A A flowchart of a method for testing a hardware component using different operating parameter values is shown in accordance with one embodiment.

[0019] Figure 2B A block diagram of a system is shown in accordance with one embodiment.Figure 2A a flowchart of steps of the method shown in FIG. 4.

[0020] Figure 2C is a graph showing the relationship between the operating speed of a hardware component and the cutoff value of a physical operating parameter according to one embodiment.

[0021] Figure 3 shows a parallel processing unit according to one embodiment.

[0022] Figure 4A shows a general processing cluster within a parallel processing unit according to one embodiment Figure 3

[0023] Figure 4B shows a memory partition unit of a parallel processing unit according to one embodiment Figure 3

[0024] Figure 5A shows a streaming multiprocessor according to one embodiment Figure 4A

[0025] Figure 5B is a conceptual diagram of a processing system implemented using a parallel processing unit (PPU) according to one embodiment Figure 3

[0026] Figure 5C shows an exemplary system that can implement various architectures and / or functionality of various previous embodiments.

[0027] Figure 6 is a block diagram of an example system architecture for an example autonomous vehicle according to one embodiment. DETAILED DESCRIPTION

[0028] Conventionally, hardware components are characterized at manufacturing time to determine physical operating parameter values for correct operation (i.e., all tests pass) over different operating conditions (e.g., temperature, operating time, etc.) and manufacturing process variations. For example, a functional test mode can be applied to a batch of hardware components over varying time periods and field operating conditions are simulated, including simulating aging by heating (e.g., baking) to produce degradation data for each physical operating parameter. The goal of the characterization is to determine baseline operating parameter values to ensure proper functioning of the hardware components as they are deployed in the field over time. Typically, the worst degradation measurements obtained for a subset of the batch are used to establish a degradation margin included in the baseline operating parameter values.

[0029] ​​​​IST can be used to determine operating parameter values for the hardware components as they are deployed in the field, effectively characterizing each hardware component individually for the actual operating conditions that the hardware component is subjected to. IST can also be used to test functional safety - e.g., that the components are operating and functioning properly or within expected tolerances. A benefit of characterizing each hardware component in the field is that the operating parameters can be set to values specific to the particular hardware component, rather than baseline operating parameter values determined at the time of manufacture. Using operating parameter values specific to the particular hardware component can improve performance while reducing power consumption compared to using baseline values. Testing the deployed components also has the benefit of detecting when the components can have a fault and / or need repair, corrective action, or otherwise be used unsafely.

[0030] IST is used to identify a boundary between pass and fail values for each operating parameter, where the boundary and corresponding “cutoff” value is specific to the hardware component and the field operating conditions. For example, Vmin is the lowest value of the supply voltage for which the test passes, as determined during IST of the hardware component. When Vmin is lower than the baseline value of the supply voltage, power consumption is reduced. When Vmin is higher than the baseline value of the supply voltage, the hardware component can continue to operate without failure and need replacement. In other words, the field service life of the hardware component can be increased compared to simply using a baseline supply voltage.

[0031] As previously explained, in the IST process, the conventional system needs to be rebooted in order to change one or more operating parameter values. For example, a reboot is needed to change from a functional mode to a test mode after the operating parameter values are changed by a central processing unit (CPU) external to the hardware component. As further described herein, the reboot can be avoided during IST by enabling the hardware component to directly adjust one or more operating parameter values rather than having the operating parameter values adjusted by the CPU.

[0032] Figure 1A A block diagram of a field computing platform 100 is shown in accordance with one embodiment. The computing platform includes a CPU 110, a device 140, a voltage regulator 130, and a memory 135. In one embodiment, the computing platform 100 includes a system on a chip (SoC), a multi-chip module (MCM), a printed circuit board (PCB), or any other feasible implementation. The computing platform 100 can be implemented as at least a portion of an advanced driver assistance system (ADAS), an autonomous driving system, and / or any form of vehicle that can or can not include autonomous and / or semi-autonomous functionality. The computing platform 100 can be implemented as at least a portion of an autonomous or semi-autonomous vehicle, an autonomous or semi-autonomous machine, an autonomous or semi-autonomous robot, a manned or unmanned aircraft, or a manned or unmanned watercraft. The computing platform 100 can be implemented as at least a portion of a server cluster. Alternatively, the computing platform 100 can be implemented within an embedded system.

[0033] In the context of the following description, device 140 is a hardware component that can include an integrated circuit, a SoC, a memory, an interface, a logic circuit, a semiconductor chip, and / or other physical components or combinations thereof. In one embodiment, memory 135 includes a dynamic random access memory (DRAM) that can be separate from or integrated into device 140. In one embodiment, voltage regulator 130 is configured by commands, instructions, and / or signals to provide a supply voltage to device 140. In one embodiment, voltage regulator 130 is included within a power management device. In one embodiment, voltage regulator 130 includes multiple regulators for generating separate supply voltages for different rails within device 140, and each regulator can be individually configured by commands.

[0034] Although computing platform 100 is described in the context of processing units, IST controller 120 within device 140 can be implemented as a program, custom circuit, or by a combination of custom circuit and program. In one embodiment, device 140 is a parallel processing unit 300 as shown. Further, one of ordinary skill in the art will understand that any system that performs the operations of the framework is within the scope and spirit of embodiments of the present disclosure. Figure 3 The parallel processing unit 300 shown. Further, one of ordinary skill in the art will understand that any system that performs the operations of the framework is within the scope and spirit of embodiments of the present disclosure.

[0035] IST controller 120 performs in-situ testing of the hardware components (e.g., circuits, logic, interfaces, memories, etc.) of field computing platform 100. In one embodiment, IST controller 120 performs fault and / or functional testing on at least a portion of computing platform 100. Test patterns 122 (e.g., vectors) stored in memory 135 are applied to the portion of computing platform 100 being tested. In one embodiment, a vector is a binary representation of both the code and data required to execute the vector. The location of one or more of test patterns 122 and test results 124 stored in memory 135 can be predetermined. In one embodiment, test results 124 are stored within IST controller 120.

[0036] The test patterns 122 can be loaded into the memory 135 from another storage resource, such as system memory or flash storage. In another embodiment, at least a portion of the tests applied by the IST controller 120 (e.g., one or more test patterns) are dynamically generated according to an algorithm, rather than read from the memory 135. In one embodiment, the test interface 125 receives the test patterns 122 from the IST controller 120 and sends vectors to the test circuitry (e.g., scan flip-flops) for testing at least a portion of the device 140. The IST controller 120 communicates through the test interface 125 to control the execution of scan testing, memory built-in self-test (MBIST), and other signals and circuitry (e.g., clocks, resets, and pads). In one embodiment, the test interface 125 is implemented according to the Institute of Electrical and Electronics Engineers (IEEE) 1500 interface architecture standard. According to aspects of the present disclosure, scan pattern testing for logic gates, such as fast scan pattern (e.g., FTM2CLK) testing, can be used to detect subtle gate defects, even if they are for individual gates. Subtle gate defects can degrade over time to cause permanent faults, which can be detected during IST.

[0037] The IST controller 120 controls the execution and sequencing of these test patterns 122, including control of the test interface 125 and the voltage regulator 130. The interface 115 provides a communication path between the IST controller 120 (through the test interface 125) and the voltage regulator 130. In one embodiment, the interface 115 is an Inter-Integrated Circuit (I2C) interface. In other embodiments, the interface 115 provides a communication path for the IST controller 120 to control one or more operating parameters in addition to, or instead of, the supply voltage. Examples of other operating parameters include values for supply current, operating clock speed, input resistance, input impedance, noise magnitude / duration, and / or temperature.

[0038] In one embodiment, the interface 115 is coupled to clock generation circuitry, which can be configured to provide a clock for the device 140 according to a run-time clock speed specified by a command received from the IST controller 120. In one embodiment, the interface 115 can be coupled to a current regulator, which can be configured to regulate the current for the device 140 according to a command received from the IST controller 120. In one embodiment, the interface 115 is coupled to circuitry for regulating the input resistance, input impedance, noise, and / or temperature of the device 140, each of which can be configured according to a command received from the IST controller 120.

[0039] During in-field failure and / or functional testing, the IST controller 120 adjusts at least one physical operating parameter, such as the supply voltage. The IST controller 120 can send commands to the voltage regulator 130 through the test interface 125 and the interface 115 to adjust the supply voltage by controlling the voltage regulator 130 to change the value of the supply voltage provided to the device 140. In contrast, in conventional computing platforms, the CPU 110 updates the voltage regulator 130 via the dashed arrow during functional test mode, and typically requires a reboot of the computing platform 100 to reload software for failure testing before entering the test mode to perform IST. Because the IST controller 120 can send commands to the voltage regulator 130 to update the voltage level and avoid a full system reboot, test latency is reduced. Reducing the amount of time required to complete in-field IST is beneficial because the computing platform 100 is unavailable for other processing during IST. Reducing the number of system reboots is beneficial because the occurrence of failures due to manufacturing defects (e.g., bump contact reliability) typically increases with the number of system reboots.

[0040] When execution of the tests (e.g., one or more of the test modes 122) is completed during IST, the test results 124 are stored in the memory 135. The test results 124 indicate whether the tests passed or failed. The test results 124 can be read by the IST controller 120 and / or the CPU 110 and used to determine the next value of the operating parameter and / or the sequence of one or more of the test modes 122 to apply. In one embodiment, one or more operating conditions (e.g., noise, temperature, load, voltage level, etc.) are measured and used to determine the next value of the operating parameter and / or the sequence of test modes. In one embodiment, during testing, one or more operating conditions are measured and an algorithm is used to compensate for any deviations from an expected range of operating conditions. In one embodiment, at the start of testing, the IST controller 120 initializes the timer 105 and waits to read the test results 124 until after the timer 105 expires. If the test results 124 are not ready after the timer expires, the IST controller 120 can indicate that a reboot is required before IST can continue. In one embodiment, code to start the CPU 110 and / or the computing platform 100 is stored in the memory 135.

[0041] In accordance with the desires of the user, more illustrative information regarding various optional architectures and features will now be set forth in connection with the foregoing framework, which can implement these optional architectures and features. It should be strongly noted that the following information is set forth in connection with the foregoing framework in order to illustrate and provide a more detailed description of the various features and structures. The following features can optionally be combined with or without the other features described.

[0042] The IST controller 120 or the CPU 110 can be configured to update values of the operating parameters based on in-field IST to improve performance of the computing platform 100. In an example, when the computing platform 100 is deployed in a data center, a particular workload is run during the day and has a lighter workload during the night. Thus, the workload profile is not evenly distributed over each 24-hour period as can have been assumed during the baseline characterization. In one embodiment, the IST controller 120 performs tests during the day to capture day workload operating conditions in the test results 124 and determines values of the operating parameters to use during the day. Similarly, the IST controller 120 performs tests during the night to capture night workload operating conditions in the test results 124 and determines values of the operating parameters to use during the night. Thus, the operating parameter values of the device 140 can be adjusted for particular operating conditions. In contrast, simply using the baseline operating parameter values can result in increased device failures during the day and power consumption and / or performance inefficiencies during the night.

[0043] Figure 1B A flowchart of a method 150 for performing in-system testing is shown, according to one embodiment. While the method 150 is described in the context of a controller (e.g., processing unit), the method 150 can also be performed by a program, custom circuitry, or by a combination of custom circuitry and program. For example, the method 150 can be performed by a GPU (graphics processing unit), CPU (central processing unit), or any processor capable of implementing IST. In one embodiment, the device 140 is the parallel processing unit 300 shown. Further, one of ordinary skill in the art will appreciate that any system performing the method 150 is within the scope and spirit of embodiments of the present disclosure. Figure 3 The parallel processing unit 300 is shown. Further, one of ordinary skill in the art will appreciate that any system performing the method 150 is within the scope and spirit of embodiments of the present disclosure.

[0044] At step 155, at least a portion of the test is performed on a hardware component of the in-field computing platform 100 (e.g., the device 140) to produce first test results, where the first values are for application to a physical operating parameter of the in-field hardware component. In one embodiment, the full failure or functional test can be subdivided into a plurality of portions, each including one or more test patterns from the test patterns 122. Instead of applying the full test during IST and producing a pass / fail test result when execution of the full test is complete, a test result can be produced for each portion. In one embodiment, execution of the test is stopped when a failure is detected. Thus, test time can be reduced when a failure is detected for one of the portions and execution of the entire test is avoided. In one embodiment, the full test is a functional test and the test is performed by the CPU 110 to test the device 140. In another embodiment, the test portions are applied by the IST controller 120.

[0045] At step 160, the first test result is stored in memory 135 within computing platform 100. In one embodiment, the first test result is stored in memory 135 by IST controller 120. The first test result can indicate that the portion of the test passed or failed when executed by device 140.

[0046] At step 165, the physical operating parameter is updated to a second value based at least on the first test result in response to a command generated by IST controller 120. The second value can be the same or different compared to the first value. In one embodiment, when the first test result indicates a failure, IST controller 120 can increase or decrease the value of the physical operating parameter, thereby updating the physical operating parameter to a value that the same test or a different test has previously passed during IST.

[0047] For example, to determine Vmin, IST controller 120 uses a first value for the supply voltage that is expected to pass the entire test based on the baseline operating parameter value and previous test results. When the entire test (all test patterns 122) passes using the first value for the supply voltage, IST controller 120 decreases the supply voltage to a second value and begins executing at least a portion of the entire test. IST controller 120 can be configured to execute the entire test to produce a test result, or IST controller 120 can be configured to execute one or more portions (e.g., slices) of the entire test to produce a test result for each portion.

[0048] The duration of IST can be reduced by implementing a search algorithm in the microcode to be executed by IST controller 120 to dynamically apply the test patterns 122 in order based on test results 124 for different values of at least one of the physical operating parameters. In contrast, in a conventional IST process, the entire test is executed for each of the different values of the operating parameter before the test results are examined.

[0049] To reduce the duration of IST, once a failed test result occurs, IST controller 120 can identify the pass / fail boundary corresponding to Vmin by testing using a supply voltage higher than the supply voltage corresponding to the failed test result. Additional test patterns can also be applied based on the results of the initial testing to more fully test around the pass / fail boundary.

[0050] When the lowest supply voltage corresponding to the test results is identified, the entire test or a portion of the entire test can be repeated one or more times to confirm that the supply voltage is at Vmin (the lowest supply voltage at which all test patterns 122 are successfully passed). In another embodiment, the IST controller 120 can start the test using a low supply voltage and gradually increase the supply voltage, rather than searching for Vmin by gradually decreasing the supply voltage to find the first failure. However, starting at a higher supply voltage increases the likelihood that the execution of the test will complete and not get stuck or hung. Moreover, starting at a higher supply voltage ensures that data within the device 140 is more reliably cleared and avoids the risk of exposing any secure data due to a failure during a secure boot at a speculative (e.g., too low) supply voltage.

[0051] At step 170, execution of the test of the field hardware component is resumed using the second value of the physical operating parameter to produce a second test result. After completion of the execution of the entire test (or another portion of the test), the IST controller 120 can return to step 160 to store the second test result in the test results 124.

[0052] In addition to implementing a search algorithm, the IST controller 120 can be configured to determine whether the test has completed execution and got stuck. Specifically, when the supply voltage is too low, it is possible that the execution of the test will not end. Figure 2A A flowchart of a method 200 for testing a hardware component using different operating parameter values is shown in accordance with one embodiment. While the method 200 is described in the context of a controller (e.g., processing unit), the method 150 can also be performed by a program, custom circuit, or by a combination of custom circuit and program. For example, the method 200 can be performed by a GPU, CPU, or any processor capable of implementing IST. Moreover, one of ordinary skill in the art will appreciate that any system performing the method 200 is within the scope and spirit of embodiments of the present disclosure.

[0053] As previously combined with the description of the method 150, the IST controller 120 can be configured to perform the method 200. For example, the IST controller 120 can be configured to perform the method 200 using the test patterns 122, the test results 124, and the supply voltage 126. Figure 1BThe described, execution steps 155, 160, and 165 are performed. At step 170, execution of the test on the field hardware component is resumed using the second value of the physical operating parameter to produce a second test result. When the test is restarted, the timer 105 is initialized. In one embodiment, the timer 105 is initialized to a predetermined value that is specific to the test and is an appropriate duration for generating the second test result. When the timer 105 expires during step 170, the IST controller 120 determines whether the entire test or a portion of the entire test is complete execution at step 260. In one embodiment, each test result in the test results 124 includes a test completion status indicator and a test pass status indicator, where the test completion status indicates whether the test is complete (1) and the test pass status indicates whether the test passed (1) or failed (0). In one embodiment, the test pass status is only valid if the test completion status is set to 1.

[0054] If at step 260, the IST controller 120 determines that the test is not complete, then at step 275, the IST controller 120 instructs that the computing platform should be rebooted. In another embodiment, the CPU 110 performs steps 260 and 275. If at step 260, the IST controller 120 determines that the test has completed, then at step 265, the IST controller 120 updates the physical operating parameter or the test based on the completed test result. In conjunction with Figure 2B Step 265 is described in more detail. Updating the test includes identifying a portion of the entire test to be performed. For example, the IST controller 120 or the CPU 110 can be configured to determine a sequence of the test patterns 122. In one embodiment, the IST controller 120 or the CPU 110 can perform the same portion of the test or the entire test with or without updating the physical operating parameter. At step 270, execution of the test on the hardware component is resumed. Steps 260, 275, 265, and 270 can be repeated until the IST of the hardware component is complete. A portion of the test results 124, the operating parameter values, and / or operating conditions can be stored for use by the IST controller 120 and / or field characterization of the hardware component.

[0055] Figure 2B A flowchart of step 265 of the method 200 shown in FIG. 2 is shown in accordance with one embodiment. Figure 2A A flowchart of step 265 of the method 200 shown in FIG. 2 is shown in accordance with one embodiment.

[0056] If the test fails, the IST controller 120 updates the operating parameter by reverting to the previous value before proceeding to step 270. Otherwise, if the test passes, then at step 268 the IST controller 120 determines whether the test should be performed for another iteration. If the IST controller 120 determines that the test should be performed for another iteration, then the execution of the test is resumed at step 270. Otherwise, at step 269 the IST controller 120 updates the operating parameter to a next value different from the value used to produce the test result before proceeding to step 270.

[0057] The IST controller 120 provides various techniques for reducing the duration of the IST for in-field characterization of hardware components. In particular, the IST controller 120 can reduce the duration of the IST during which the control operating parameter values are controlled by removing the latency introduced for rebooting the computing platform 100 for each update of the operating parameter. Further, the IST controller 120 can also reduce the duration of the IST by implementing a search algorithm for pass / fail boundaries without having to perform the entire test at each value of a particular operating parameter. Further, the IST controller 120 can perform a signature comparison for each test result to determine whether each test passes or fails.

[0058] For speculative supply voltage values, when data can be more reliably cleared by increasing the supply voltage level and performing a flush, the risk of exposing on-die data is reduced. Conversely, when a reboot occurs at a speculative supply voltage value, data can not be reliably cleared and is at risk of being exposed.

[0059] Reducing the IST duration provides the ability to determine the cutoff value of an operating parameter in-field under actual usage conditions. The cutoff value can be used to estimate the degradation specific to a particular hardware component. The measured cutoff value for a particular hardware component can be less conservative than a baseline value for the operating parameter. As a result, performance can be improved and / or the lifetime of the hardware component can be extended. Further, the determination of the cutoff value enables the measurement of degradation over time and the prediction of functional failures.

[0060] U.S. Patent Application Serial No. 16 / 601,900, filed October 15, 2019, entitled “Enhanced In-System Test Coverage Based on Detecting Component Degradation,” describes a technique for determining a rate of degradation of a physical operating parameter of a hardware component from results of tests performed for different values of the physical operating parameter (e.g., supply voltage). The IST techniques described herein can be used to generate degradation data for hardware components in the field. The IST techniques described herein can be used to perform in-field testing for applications for which early detection of degradation faults is necessary for improved safety and / or operational availability.

[0061] Figure 2C FIG. 205 is a graph 205 showing a relationship between operating speed of a hardware component and a cutoff value of a physical operating parameter, according to one embodiment. Each point in graph 205, such as points 202 and 204, can represent a cutoff value of a respective hardware component that is tested prior to deployment on a computing platform (such as at a factory) or early on after deployment on a computing platform (such as computing platform 100). In at least some embodiments, the cutoff value of each hardware component can be determined using the methods described herein. As shown, point 202 can correspond to a hardware component with an operating speed of 1800 and a cutoff value of 0.74V. Point 204 can correspond to a hardware component with an operating speed of 1650 and a cutoff value of 0.79V. Threshold line 206 is fitted to the points of graph 205, showing that as operating speed increases, the cutoff value tends to decrease.

[0062] IST controller 120 can utilize the relationship of graph 205 in determining one or more values of one or more physical operating parameters of a test, allowing the cutoff value to be identified in fewer iterations of the test and taken into account in determining the rate of degradation. IST controller 120 can use operating conditions, such as temperature, resulting in values of the physical operating parameter that are different for different test runs. For example, a temperature of 45 degrees Celsius results in a value of 460mV for a first test, and a temperature of 75 degrees Celsius results in a value of 440mV for a second test. As described herein, temperature, operating speed, and / or other characteristics can be captured in tables used by IST controller 120 to look up values of physical operating parameters for a test and / or to calculate these values.

[0063] Returning to Figure 2CFigure 205 illustrates threshold lines 206, 208, 210, and 212, which may correspond to applied values ​​of physical operating parameters for different operating conditions and / or reference degradation rate values. Threshold lines 208, 210, and 212 may correspond to thresholds on cutoff values ​​of hardware components, which correspond to points in Figure 205 at different operating durations after deployment, and may be offset from threshold line 206 to capture the degradation of hardware components over time. In at least one embodiment, if the performance characteristics (in this case, the cutoff value) are too high for the aging and / or usage of the hardware component to trigger one or more remedial actions via the remedial action manager, the degradation rate analyzer determines that the hardware component includes a potential defect.

[0064] like Figure 2C As shown, during initial deployment, if the cutoff value exceeds the corresponding value on threshold line 208 at or before deployment, the degradation rate analyzer can determine that the hardware component contains a potential defect. If the cutoff value exceeds the corresponding value on threshold line 210 from operational deployment to 5000 hours of operation, the degradation rate analyzer can determine that the hardware component contains a potential defect. If the cutoff value exceeds the corresponding value on threshold line 212 from operational deployment to 5000 hours of operation, the degradation rate analyzer can determine that the hardware component contains a potential defect. Although in Figure 2C The figure shows the number of hours the device has been running, but another form of measurement, such as the usage or aging of the hardware components, can be used.

[0065] In at least one embodiment, the degradation rate determiner provides the degradation rate analyzer with measurements of the use or aging of hardware components, and values ​​of performance characteristics derived from test runs. The degradation rate analyzer uses the measurements to calculate and / or look up a reference degradation rate value and determine whether the value of the performance characteristic exceeds the reference degradation rate value. If the degradation rate analyzer determines that the reference degradation rate value is exceeded, an indication can be included in the analysis results provided to the remediation action manager to initiate one or more actions. Examples of remediation actions include disabling one or more functions of the computing platform 100, such as functions implemented using hardware components. In the computing platform 100, which is... Figure 6 In embodiments of the autonomous driving system of vehicle 900, one or more remedial actions may include disabling autonomous driving of vehicle 900 and / or one or more ADAS features. Further examples of remedial measures include causing an indicator to display a degradation rate exceeding a reference degradation rate.

[0066] Parallel processing architecture

[0067] Figure 3A parallel processing unit (PPU) 300 according to one embodiment is illustrated. In one embodiment, the PPU 300 is a multi-threaded processor implemented on one or more integrated circuit devices. The PPU 300 is a latency-hiding architecture designed for parallel processing of many threads. A thread (i.e., an execution thread) is an instance of a set of instructions configured to be executed by the PPU 300. In one embodiment, the PPU 300 is a graphics processing unit (GPU) configured to implement a graphics rendering pipeline for processing three-dimensional (3D) graphics data to generate two-dimensional (2D) image data for display on a display device, such as a liquid crystal display (LCD) device. In other embodiments, the PPU 300 may be used to perform general-purpose computing. Although an exemplary parallel processor is provided herein for illustrative purposes, it should be specifically noted that this processor is illustrated for illustrative purposes only, and any processor may be used to supplement and / or replace this processor.

[0068] One or more PPU 300s can be configured to accelerate thousands of high-performance computing (HPC), data center, and machine learning applications. PPU 300s can be configured to accelerate numerous deep learning systems and applications, including autonomous vehicle platforms, deep learning, high-precision speech, image, and text recognition systems, intelligent video analytics, molecular simulations, drug discovery, disease diagnosis, weather forecasting, big data analytics, astronomy, molecular dynamics simulations, financial modeling, robotics, factory automation, real-time language translation, online search optimization, and personalized user recommendations, among others.

[0069] like Figure 3 As shown, PPU 300 includes an input / output (I / O) unit 305, a front-end unit 315, a scheduler unit 320, a job allocation unit 325, a hub 330, a crossbar (Xbar) 370, one or more general purpose processing clusters (GPCs) 350, and one or more memory partitioning units 380. PPU 300 can be connected to a host processor or other PPU 300 via one or more high-speed NVLink 310 interconnects. PPU 300 can be connected to a host processor or other peripheral devices via interconnect 302. PPU 300 can also be connected to a local memory 304 comprising multiple memory devices. In one embodiment, local memory 304 may include multiple DRAM devices. The DRAM devices may be configured as a high-bandwidth memory (HBM) subsystem, wherein multiple DRAM dies are stacked within each device.

[0070] The NVLink 310 interconnect enables system scaling and includes one or more PPUs 300 in conjunction with one or more CPUs, supports cache coherency between PPUs 300 and CPUs, and CPU mastering. Data and / or commands can be sent by the NVLink 310 through the hub 330 to other units of the PPU 300, or from them, such as one or more copy engines, video encoders, video decoders, power management units, etc. (not explicitly shown). In conjunction with Figure 5B The NVLink 310 is described in more detail.

[0071] The I / O unit 305 is configured to send and receive communications (e.g., commands, data, etc.) from a host processor (not shown) over the interconnect 302. The I / O unit 305 can communicate directly with the host processor via the interconnect 302 or can do so by way of one or more intermediate devices such as a memory hub. In one embodiment, the I / O unit 305 can communicate with one or more other processors (e.g., one or more PPUs 300) via the interconnect 302. In one embodiment, the I / O unit 305 implements a Peripheral Component Interconnect Express (PCIe) interface for communications over a PCIe bus, and the interconnect 302 is a PCIe bus. In alternative embodiments, the I / O unit 305 can implement another known interface for communicating with external devices.

[0072] The I / O unit 305 decodes data packets received via the interconnect 302. In one embodiment, the data packets represent commands configured to cause the PPU 300 to perform various operations. The I / O unit 305 sends the decoded commands, as specified by the commands, to various other units of the PPU 300. For example, some commands can be sent to the front-end unit 315. Other commands can be sent to the hub 330 or other units of the PPU 300, such as one or more copy engines, video encoders, video decoders, power management units, etc. (not explicitly shown). In other words, the I / O unit 305 is configured to route communications between and among various logical units of the PPU 300.

[0073] In one embodiment, a program executed by the host processor encodes a command stream in a buffer that provides a workload to the PPU 300 for processing. The workload can include a number of instructions and data to be processed by those instructions. The buffer is a region of memory that is accessible (e.g., read / write) by both the host processor and the PPU 300. For example, the I / O unit 305 can be configured to access the buffer in system memory connected to the interconnect 302 via memory requests transmitted over the interconnect 302. In one embodiment, the host processor writes the command stream to the buffer and then sends a pointer to the start of the command stream to the PPU 300. The front-end unit 315 receives the pointer to the command stream(s). The front-end unit 315 manages the command stream(s), reading commands from the stream(s) and forwarding the commands to various units of the PPU 300.

[0074] The front-end unit 315 is coupled to a scheduler unit 320 that configures various GPCs 350 to process tasks defined by one or more streams. The scheduler unit 320 is configured to track state information related to various tasks managed by the scheduler unit 320. The state can indicate which GPC 350 a task is assigned to, whether the task is active or inactive, a priority associated with the task, etc. The scheduler unit 320 manages execution of a plurality of tasks on the one or more GPCs 350.

[0075] The scheduler unit 320 is coupled to a work distribution unit 325 that is configured to dispatch tasks for execution on the GPCs 350. The work distribution unit 325 can track a number of scheduled tasks received from the scheduler unit 320. In one embodiment, the work distribution unit 325 manages a pending task pool and an active task pool for each GPC 350. The pending task pool can include a number of slots (e.g., 32 slots) that hold tasks assigned to be processed by a particular GPC 350. The active task pool can include a number of slots (e.g., 4 slots) for tasks that are actively being processed by a GPC 350. When a GPC 350 completes processing of a task, the task is evicted from the active task pool for the GPC 350, and one of the other tasks from the pending task pool is selected and scheduled for execution on the GPC 350. If the active task on a GPC 350 has idled, such as while waiting for a data dependency to be resolved, then the active task can be evicted from the GPC 350 and returned to the pending task pool while another task is selected and scheduled for execution on the GPC 350 from the pending task pool.

[0076] The work distribution unit 325 communicates with one or more GPCs 350 via an XBar (crossbar switch) 370. The XBar 370 is an interconnect network that couples many units of the PPU 300 to other units of the PPU 300. For example, the XBar 370 can be configured to couple the work distribution unit 325 to a specific GPC 350. Although not explicitly shown, one or more other units of the PPU 300 can also be connected to the XBar 370 via a hub 330.

[0077] Tasks are managed by scheduler unit 320 and dispatched to GPC 350 by work allocation unit 325. GPC 350 is configured to process tasks and generate results. Results may be consumed by other tasks within GPC 350, routed to different GPCs 350 via XBar 370, or stored in memory 304. Results may be written to memory 304 via memory partitioning unit 380, which implements a memory interface for reading data from and writing data to memory 304. Results may be sent to another PPU 300 or CPU via NVLink 310. In one embodiment, PPU 300 includes U memory partitioning units 380, which is equal to the number of independent and different memory devices 304 coupled to PPU 300. The following will be combined with... Figure 4B The memory partition unit 380 is described in more detail.

[0078] In one embodiment, the host processor executes a driver kernel that implements an application programming interface (API), enabling the execution of one or more applications on the host processor to schedule operations for execution on the PPU 300. In one embodiment, multiple computing applications are executed concurrently by the PPU 300, and the PPU 300 provides isolation, Quality of Service (QoS), and independent address spaces for the multiple computing applications. Applications can generate instructions (e.g., API calls) that cause the driver kernel to generate one or more tasks for execution by the PPU 300. The driver kernel outputs the tasks to one or more streams being processed by the PPU 300. Each task may include one or more associated thread groups, referred to herein as a warp. In one embodiment, a warp includes 32 associated threads that can execute in parallel. Cooperative threads can refer to multiple threads that include instructions for executing tasks and can exchange data via shared memory. Figure 5A A more detailed description of threads and cooperative threads.

[0079] Figure 4A An embodiment is shown. Figure 3 The PPU 300 and GPC 350. For example... Figure 4AAs shown, each GPC 350 includes a number of hardware units for processing tasks. In one embodiment, each GPC 350 includes a pipeline manager 410, a pre-raster operations unit (PROP) 415, a raster engine 425, a work distribution crossbar (WDX) 480, a memory management unit (MMU) 490, and one or more data processing clusters (DPCs) 420. It should be understood that Figure 4A The GPC 350 can include other hardware units or other variations of the units shown in FIG. 4B in place of or in addition to those shown. Figure 4A Figure 4A

[0080] In one embodiment, the operations of the GPC 350 are controlled by the pipeline manager 410. The pipeline manager 410 manages configuration of one or more DPCs 420 for processing tasks assigned to the GPC 350. In one embodiment, the pipeline manager 410 can configure at least one of the one or more DPCs 420 to implement at least a portion of a graphics rendering pipeline. For example, a DPC 420 can be configured to execute a vertex shading program on the programmable streaming multi-processor (SM) 440. The pipeline manager 410 can also be configured to route data from a work distribution unit 325 to appropriate logical units in the GPC 350. For example, some data packets can be routed to the fixed function hardware units in the PROP 415 and / or the raster engine 425 while other data packets can be routed to a DPC 420 for processing by the primitive engine 435 or the SM 440. In one embodiment, the pipeline manager 410 can configure at least one of the one or more DPCs 420 to implement a neural network model and / or a compute pipeline.

[0081] The PROP unit 415 is configured to route data generated by the raster engine 425 and the DPCs 420 to a render operation (ROP) unit in conjunction with Figure 4B as described in more detail below. The PROP unit 415 can also be configured to perform optimizations for color blending, organize pixel data, perform address translations, and / or the like.

[0082] ​​The raster engine 425 includes several fixed function hardware units configured to perform various raster operations. In one embodiment, the raster engine 425 includes a setup engine, a coarse raster engine, a cull engine, a clip engine, a fine raster engine, and a tile aggregate engine. The setup engine receives transformed vertices and generates a plane equation associated with a geometric primitive defined by the vertices. The plane equation is sent to the coarse raster engine to generate coverage information (e.g., an x, y coverage mask of a tile) for the primitive. The output of the coarse raster engine is sent to the cull engine, where fragments associated with primitives that fail a z-test are culled, and to the clip engine, where fragments that are outside a view frustum are clipped. Those fragments that survive clipping and culling can be passed to the fine raster engine to generate attributes for the pixel fragments based on the plane equation generated by the setup engine. The output of the raster engine 425 includes, for example, fragments to be processed by a fragment shader implemented within the DPCs 420.

[0083] Each DPC 420 included in the GPC 350 includes a M Pipe Controller (MPC) 430, a primitive engine 435, and one or more SMs 440. The MPC 430 controls the operation of the DPC 420, routing received commands to the appropriate unit of the DPC 420. For example, commands associated with a vertex will be routed to the primitive engine 435, which is configured to fetch vertex attributes associated with the vertex from the memory 304. In contrast, commands associated with a shader program can be sent to the SM 440.

[0084] The SM 440 includes a programmable streaming processor configured to process tasks represented by multiple threads. Each SM440 is multithreaded and configured to execute multiple threads (e.g., 32 threads) from a specific thread group concurrently. In one embodiment, the SM 440 implements a SIMD (Single Instruction, Multiple Data) architecture, where each thread in a thread group (e.g., a thread bundle) is configured to process a different dataset based on the same instruction set. All threads in the thread group execute the same instructions. In another embodiment, the SM 440 implements a SIMT (Single Instruction, Multiple Threads) architecture, where each thread in a thread group is configured to process a different dataset based on the same instruction set, but where individual threads in the thread group are allowed to diverge during execution. In one embodiment, a program counter, call stack, and execution state are maintained for each thread bundle, enabling concurrency between thread bundles and serial execution within a thread bundle when threads within a thread bundle diverge. In another embodiment, a program counter, call stack, and execution state are maintained for each individual thread, thereby achieving equal concurrency among all threads within and between thread bundles. When maintaining the execution state for each individual thread, threads executing the same instructions can converge and execute in parallel to achieve maximum efficiency. The following section combines... Figure 5A A more detailed description of the SM 440.

[0085] MMU 490 provides an interface between GPC 350 and memory partitioning unit 380. MMU 490 can provide virtual address to physical address translation, memory protection, and arbitration of memory requests. In one embodiment, MMU 490 provides one or more translation back buffers (TLBs) for performing translations from virtual addresses to physical addresses in memory 304.

[0086] Figure 4B An embodiment is shown. Figure 3 The PPU 300's memory partition unit 380. For example... Figure 4B As shown, the memory partition unit 380 includes a raster operation (ROP) unit 450, a secondary (L2) cache 460, and a memory interface 470. The memory interface 470 is coupled to the memory 304. The memory interface 470 can implement 32, 64, 128, or 1024-bit data buses for high-speed data transfer. In one embodiment, the PPU 300 incorporates U memory interfaces 470, one memory interface 470 for each pair of memory partition units 380, wherein each pair of memory partition units 380 is connected to a memory device of the corresponding memory 304. For example, the PPU 300 can connect to up to Y memory devices 304, such as high-bandwidth memory stacks or synchronous dynamic random access memory of Graphics Dual Data Rate Version 5, or other types of persistent memory.

[0087] In one embodiment, memory interface 470 implements an HBM2 memory interface and Y is equal to half of U. In one embodiment, HBM2 memory stacks are located on the same physical package as PPU 300, providing significant power and area savings compared to a conventional GDDR5 SDRAM system. In one embodiment, each HBM2 stack includes four memory dies and Y is equal to 4, where the HBM2 stack includes two 128-bit channels per die for a total of 8 channels and a data bus width of 1024 bits.

[0088] In one embodiment, memory 304 supports single error correction double error detection (SECDED) error correcting code (ECC) to protect data. For computing applications sensitive to data corruption, ECC provides higher reliability. In large cluster computing environments, PPU 300 processes very large data sets and / or long-running applications, where reliability is especially important.

[0089] In one embodiment, PPU 300 implements a multi-level memory hierarchy. In one embodiment, memory partition unit 380 supports a unified memory to provide a single unified virtual address space for a CPU and PPU 300 memory, enabling data sharing between virtual memory systems. In one embodiment, the frequency of access to memory locations by PPU 300 that are located on other processors is tracked such that memory pages are moved to the physical memory of the PPU 300 that accesses the pages most frequently. In one embodiment, NVLink 310 supports address translation services, which allow PPU 300 to access page tables stored in host memory, and provide full access to a CPU's memory by a PPU 300.

[0090] In one embodiment, a copy engine transfers data between multiple PPUs 300 or between a PPU 300 and a CPU. The copy engine can generate a page fault for an address that is not mapped into a page table. Memory partition unit 380 can then service the page fault, map an address into a page table, after which the copy engine can perform the transfer. In a conventional system, memory is pinned (e.g., unpageable) for multiple copy engine operations between multiple processors, which significantly reduces available memory. Due to hardware paging, an address can be passed to the copy engine without worrying about whether a memory page is resident, and whether the copy process is transparent.

[0091] Data from memory 304 or other system memory can be retrieved by memory partitioning unit 380 and stored in L2 cache 460, which is located on-chip and shared among the various GPCs 350. As shown, each memory partitioning unit 380 includes a portion of the L2 cache 460 associated with the corresponding memory 304. Lower-level caches can then be implemented in multiple cells within the GPC 350. For example, each SM 440 can implement a Level 1 (L1) cache. The L1 cache is a dedicated memory for a specific SM 440. Data from L2 cache 460 can be fetched and stored in each L1 cache for processing within the functional units of the SM 440. L2 cache 460 is coupled to memory interface 470 and XBar 370.

[0092] ROP unit 450 performs graphic raster operations related to pixel color, such as color compression and pixel blending. ROP unit 450 also performs depth testing in conjunction with raster engine 425, receiving the depth of sample locations associated with pixel fragments from the culling engine of raster engine 425. The depth of the sample location associated with the fragment is tested relative to the corresponding depth in the depth buffer. If the fragment passes the depth test for the sample location, ROP unit 450 updates the depth buffer and sends the result of the depth test to raster engine 425. It will be understood that the number of memory partition units 380 may differ from the number of GPCs 350, and therefore each ROP unit 450 may be coupled to each GPC 350. ROP unit 450 tracks data packets received from different GPCs 350 and determines which GPC 350 the result generated by ROP unit 450 is routed to via Xbar 370. Although in Figure 4B ROP unit 450 is included within memory partition unit 380, but in other embodiments, ROP unit 450 may be located outside memory partition unit 380. For example, ROP unit 450 may reside in GPC 350 or another unit.

[0093] Figure 5A An embodiment is shown. Figure 4A The streaming multiprocessor 440. For example... Figure 5A As shown, the SM 440 includes an instruction cache 505, one or more scheduler units 510, a register file 520, one or more processing cores 550, one or more special function units (SFUs) 552, one or more load / store units (LSUs) 554, an interconnect network 580, and a shared memory / L1 cache 570.

[0094] As described above, the work distribution unit 325 dispatches tasks for execution on the GPCs 350 of the PPU 300. A task is assigned to a specific DPC 420 within a GPC 350 and, if the task is associated with a shader program, to an SM 440. The scheduler unit 510 receives tasks from the work distribution unit 325 and manages instructions for one or more thread blocks assigned to the SM 440. The scheduler unit 510 dispatches thread blocks to be executed by thread warps of parallel threads, where each thread block is assigned at least one thread warp. In one embodiment, each thread warp executes 32 threads. The scheduler unit 510 can manage a plurality of different thread blocks, allocating thread warps to different thread blocks and then dispatching instructions from the plurality of different cooperative groups to various functional units (i.e., the cores 550, the SFU 552, and the LSU 554) during each clock cycle.

[0095] Cooperative groups are a programming model for organizing groups of communicating threads that allow developers to express the granularity at which threads are communicating, enabling richer, more efficient parallel decomposition. The cooperative launch API supports synchronicity between thread blocks to execute parallel algorithms. Conventional programming models provide a single, simple construct for synchronizing cooperating threads: a barrier across all threads of a thread block (e.g., the syncthreads() function). However, programmers often want to define groups of threads at a granularity smaller than a thread block and synchronize within the defined groups to enable higher performance, design flexibility, and software reuse in the form of collective group-wide function interfaces.

[0096] Cooperative groups enable programmers to explicitly define groups of threads at sub-block (e.g., as small as a single thread) and multi-block granularities and perform collective operations, such as synchronicity across threads in a cooperative group. The programming model supports clean composition across software boundaries so that libraries and utility functions can safely synchronize in their local environment without making assumptions about convergence. Cooperative group primitives enable new patterns of cooperative parallelism, including producer-consumer parallelism, opportunistic parallelism, and global synchronization across a grid of thread blocks.

[0097] The dispatch unit 515 is configured to transmit instructions to one or more functional units. In this embodiment, the scheduler unit 510 includes two dispatch units 515, which enable two different instructions from the same thread warp to be dispatched during each clock cycle. In alternative embodiments, each scheduler unit 510 can include a single dispatch unit 515 or additional dispatch units 515.

[0098] Each SM 440 includes a register file 520 that provides a set of registers for the functional units of the SM 440. In one embodiment, the register file 520 is divided into registers for each of the functional units such that each functional unit is allocated a dedicated portion of the register file 520. In another embodiment, the register file 520 is partitioned among different thread blocks executed by the SM 440. The register file 520 provides temporary storage for operands of the data

[0099] Each SM 440 includes L processing cores 550. In one embodiment, the SM 440 includes a large number (e.g., 128) of different processing cores 550. Each core 550 can include a fully pipelined, single-precision, double-precision, and / or mixed precision processing unit including floating point and integer arithmetic logic units. In one embodiment, the floating point arithmetic logic units implement the IEEE 754-2008 standard for floating point arithmetic. In one embodiment, the cores 550 include 64 single-precision (32-bit) floating point cores, 64 integer cores, 32 double-precision (64-bit) floating point cores, and 8 tensor cores.

[0100] The tensor cores are configured to perform matrix operations, and in one embodiment, one or more tensor cores are included in the cores 550. Specifically, the tensor cores are configured to perform deep learning matrix operations, such as convolution operations for neural network training and inference. In one embodiment, each tensor core operates on 4x4 matrices and performs matrix multiply and accumulate operations D = A x B + C, where A, B, C, and D are 4x4 matrices.

[0101] In one embodiment, the matrix multiply inputs A and B are 16-bit floating point matrices, while the accumulate matrices C and D can be 16-bit floating point or 32-bit floating point matrices. The tensor cores operate on 16-bit floating point input data as well as 32-bit floating point accumulation. The 16-bit floating point multiplication requires 64 operations to produce a full precision product, which is then accumulated using 32-bit floating point addition with other intermediate products of the 4x4x4 matrix multiplication. In practice, the tensor cores are used to perform larger two-dimensional or higher dimensional matrix operations built up from these smaller elements. APIs, such as the CUDA 9 C++ API, expose specialized matrix load, matrix multiply and accumulate, and matrix store operations to efficiently use the tensor cores from a CUDA-C++ program. At the CUDA level, the thread block level interface assumes 16x16 size matrices across all 32 threads of a thread block.

[0102] Each SM 440 also includes M SFUs 552 that perform special functions, e.g., certain mathematical functions, bit and integer logic functions, etc. In one embodiment, SFUs 552 include tree traversal units configured to traverse a hierarchical tree data structure. In one embodiment, SFUs 552 include a texture unit configured to perform texture map filtering operations. In one embodiment, the texture unit is configured to receive a texture map (e.g., a 2D array of texture pixels) from memory 304 and perform sampling operations to produce sampled texture values for use in a shader program executed by SM 440. In one embodiment, the texture map is stored in shared memory / L1 cache 570. The texture unit implements texture operations such as filtering operations using mipmaps (i.e., texture maps of varying levels of detail) in one embodiment. In one embodiment, each SM 440 includes two texture units.

[0103] Each SM 440 also includes N LSUs 554 that implement load and store operations between shared memory / L1 cache 570 and register file 520. Each SM 440 includes an interconnect network 580 that connects each of the functional units to register file 520 and LSUs 554 to register file 520, shared memory / L1 cache 570. In one embodiment, interconnect network 580 is a cross-bar switch that allows any of the functional units to connect to any of the register files, as well as the LSUs 554 to connect to either or both of the register files and the shared memory / L1 cache 570.

[0104] Shared memory / L1 cache 570 is an on-chip memory unit that allows data storage for instructions and / or data, as understood by one skilled in the art. In one embodiment, shared memory / L1 cache 570 includes 128 KB of storage space and is on the path from SM 440 to memory partition unit 380. Shared memory / L1 cache 570 can be used to cache reads and writes. One or more of shared memory / L1 cache 570, L2 cache 460, and memory 304 are backed up by off-chip memory (not shown).

[0105] Combining data cache and shared memory functionality into a single memory block provides the best overall performance for both types of memory accesses. The capacity can be used by a program as a cache that does not use shared memory. For example, if the shared memory is configured to use half the capacity, then texture and load / store operations can use the remaining capacity. The integration within shared memory / L1 cache 570 causes shared memory / L1 cache 570 to function as a high-throughput pipeline for streaming data and, at the same time, provide high-bandwidth and low-latency access to frequently-reused data.

[0106] When configured for general-purpose parallel computing, a simpler configuration can be used compared to graphics processing. Specifically, Figure 3 The illustrated fixed-function graphics processing units are bypassed, creating a simpler programming model. In a general-purpose parallel computing configuration, work distribution unit 325 assigns and dispatches thread blocks directly to DPCs 420. The threads in a block execute the same program, use the unique thread ID in the computation to ensure each thread generates a unique result, use SM 440 to execute the program and perform the computation, use shared memory / L1 cache 570 to communicate between threads, and use LSU 554 to read and write global memory through shared memory / L1 cache 570 and memory partition unit 380. When configured for general-purpose parallel computing, SM 440 can also write commands that scheduler unit 320 can use to launch new work on DPCs 420.

[0107] PPU 300 can be included in a desktop computer, a laptop computer, a tablet computer, a server, a supercomputer, a smart phone (e.g., a wireless, hand-held device), a personal digital assistant (PDA), a digital camera, a vehicle, a head-mounted display, a hand-held electronic device, etc. In one embodiment, PPU 300 is contained on a single semiconductor die. In another embodiment, PPU 300 is included on a system-on-a-chip (SoC) along with one or more other devices, such as an additional PPU 300, a memory 304, a reduced instruction set computer (RISC) CPU, a memory management unit (MMU), a digital-to-analog converter (DAC), etc.

[0108] In one embodiment, PPU 300 can be included on a graphics card that includes one or more memories. The graphics card can be configured to interface with a PCIe slot on a motherboard of a desktop computer. In yet another embodiment, PPU 300 can be an integrated graphics processing unit (iGPU) or parallel processor contained in a chipset of a motherboard.

[0109] Exemplary Computing System

[0110] Systems with multiple GPUs and CPUs are being used across various industries as developers expose to and leverage greater parallelism in applications such as artificial intelligence computing. High-performance GPU-accelerated systems with tens to thousands of compute nodes are being deployed in data centers, research institutions, and supercomputers to tackle larger problems. As the number of processing devices within high-performance systems increases, communication and data transmission mechanisms need to be scaled to support this increased bandwidth.

[0111] Figure 5B This is based on the use of one embodiment. Figure 3 A conceptual diagram of a processing system 500 implemented by a PPU 300. An exemplary system 565 can be configured to implement the method 150 shown in Figure 1 and / or Figure 2A The method 200 shown is illustrated. The processing system 500 includes a CPU 530, a switch 510, and multiple PPUs 300 and corresponding memory 304. NVLink 310 provides a high-speed communication link between each PPU 300. Although... Figure 5B A specific number of NVLink 310 and interconnect 302 connections are shown, but the number of connections to each PPU 300 and CPU 530 can vary. Switch 510 interfaces between interconnect 302 and CPU 530. PPU 300, memory 304, and NVLink 310 can reside on a single semiconductor platform to form a parallel processing module 525. In one embodiment, switch 510 supports two or more protocols that interface between various different connections and / or links.

[0112] In another embodiment (not shown), NVLinks 310 provide one or more high-speed communication links between each PPU 300 and CPU 530, and switch 510 interfaces between interconnect 302 and each PPU 300. PPU 300, memory 304, and interconnect 302 can be located on a single semiconductor platform to form a parallel processing module 525. In yet another embodiment (not shown), interconnect 302 provides one or more communication links between each PPU 300 and CPU 530, and switch 510 interfaces between each PPU 300 using NVLinks 310 to provide one or more high-speed communication links between the PPUs 300. In another embodiment (not shown), NVLinks 310 provide one or more high-speed communication links between PPU 300 and CPU 530 through switch 510. In yet another embodiment (not shown), interconnect 302 provides one or more communication links directly between each PPU 300. One or more NVLink 310 high-speed communication links can be implemented as a physical NVLink interconnect or an on-chip or on-die interconnect using the same protocol as NVLink 310.

[0113] In the context of this specification, a single semiconductor platform can refer to a sole unitary semiconductor-based integrated circuit that is fabricated in a single fabrication operation. It should be noted that the term single semiconductor platform can also refer to multi-chip modules with increased connectivity which simulate on-chip operation, and make substantial improvements over utilizing a conventional bus implementation. Of course, the various circuits or devices can alternatively be placed individually or in various combinations of two or more of the circuits or devices.

[0114] In one embodiment, the signaling rate of each NVLink 310 is 20 to 25 gigabits / second, and each PPU 300 includes six NVLink 310 interfaces (as shown in FIG. 4). Each NVLink 310 provides a data transfer rate of 25 gigabits / second in each direction, with six links providing 300 gigabits / second. When CPU 530 also includes one or more NVLink 310 interfaces, NVLinks 310 can be dedicated to PPU-to-PPU communication, as shown in FIG. 4, or some combination of PPU-to-PPU and PPU-to-CPU. Figure 5B In one embodiment, the signaling rate of each NVLink 310 is 20 to 25 gigabits / second, and each PPU 300 includes six NVLink 310 interfaces (as shown in FIG. 4). Each NVLink 310 provides a data transfer rate of 25 gigabits / second in each direction, with six links providing 300 gigabits / second. When CPU 530 also includes one or more NVLink 310 interfaces, NVLinks 310 can be dedicated to PPU-to-PPU communication, as shown in FIG. 4, or some combination of PPU-to-PPU and PPU-to-CPU. Figure 5B In one embodiment, the signaling rate of each NVLink 310 is 20 to 25 gigabits / second, and each PPU 300 includes six NVLink 310 interfaces (as shown in FIG. 4). Each NVLink 310 provides a data transfer rate of 25 gigabits / second in each direction, with six links providing 300 gigabits / second. When CPU 530 also includes one or more NVLink 310 interfaces, NVLinks 310 can be dedicated to PPU-to-PPU communication, as shown in FIG. 4, or some combination of PPU-to-PPU and PPU-to-CPU.

[0115] In one embodiment, the NVLink 310 allows direct load / store / atomic access from the CPU 530 to the memory 304 of each PPU 300. In one embodiment, the NVLink 310 supports coherency operations allowing data read from the memory 304 to be stored in the cache hierarchy of the CPU 530, reducing cache access latency for the CPU 530. In one embodiment, the NVLink 310 includes support for address translation services (ATS) allowing the PPU 300 to directly access page tables within the CPU 530. One or more NVLinks 310 can also be configured to operate in a low power mode.

[0116] Figure 5C An exemplary system 565 is shown in which various previously described embodiments of various architectures and / or functionality can be implemented. The exemplary system 565 can be configured to implement the method 150 shown in FIG. 1 and / or the method 200 shown in FIG. 2. Figure 2A

[0117] As shown, a system 565 is provided that includes at least one central processing unit 530 coupled to a communication bus 575. The communication bus 575 can be implemented using any suitable protocol, such as PCI (Peripheral Component Interconnect), PCI-Express, AGP (Accelerated Graphics Port), HyperTransport, or any other bus or point-to-point communication protocol(s). The system 565 also includes a main memory 540. Control logic (software) and data are stored in the main memory 540, which can take the form of random access memory (RAM).

[0118] The system 565 also includes an input device 560, a parallel processing system 525, and a display device 545, such as a conventional CRT (cathode ray tube), LCD (liquid crystal display), LED (light-emitting diode), plasma display, or the like. User input can be received from the input device 560, e.g., keyboard, mouse, touchpad, microphone, etc. Each of the aforementioned modules and / or devices can even be located on a single semiconductor platform, e.g., a system on a chip. Alternatively, various modules can be located on different semiconductor platforms, which can be configured, e.g., to form a

[0119] Furthermore, the system 565 can be coupled to a network (e.g., a telecommunications network, a local area network (LAN), a wireless network, a wide area network (WAN) such as the Internet, a peer-to-peer network, cable network, etc.) for communication purposes through a network interface 535.

[0120] ​The system 565 can also include secondary storage (not shown). Secondary storage includes, for example, a hard disk drive and / or a removable storage drive, representing a floppy disk drive, a magnetic tape drive, a compact disk drive, digital versatile disk (DVD) drive, recording device, universal serial bus (USB) flash drive. The removable storage drive reads from and / or writes to a removable storage unit in a well-known manner.

[0121] Computer programs, or computer control logic algorithms, can be stored in the main memory 540 and / or the secondary memory. These computer programs, when executed, enable the system 565 to perform various functions. The memory 540, the storage, and / or any other storage is a possible example of computer-readable media.

[0122] The architectures and / or functionalities of the various preceding figures can be implemented in the context of a general computer system, a circuit board system, a game console system dedicated for entertainment purposes, a special purpose system, and / or any other desired system. For example, the system 565 can take the form of a desktop computer, laptop computer, tablet computer, server computer, supercomputer, smart telephone (e.g., wireless, hand-held device), personal digital assistant (PDA), digital camera, vehicle, head mounted display, hand-held electronic device, mobile telephone device, television, workstation, game console, embedded system, and / or any other type of logic.

[0123] While various embodiments have been described above, it should be understood that they have been presented by way of example only, and not limitation. Thus, the breadth and scope of the preferred embodiments should not be limited by any of the above described exemplary embodiments, but should instead be defined in accordance with the following claims and their equivalents.

[0124] Graphics Processing Pipeline

[0125] In one embodiment, the PPU 300 includes a graphics processing unit (GPU). The PPU 300 is configured to receive commands that specify processing of graphics data by a shader program. The graphics data can be defined as a set of primitives, such as points, lines, triangles, quads, triangle strips, etc. Typically, a primitive includes data that specifies a plurality of vertices (e.g., in a model space coordinate system) and attributes associated with each vertex of the primitive. The PPU 300 can be configured to process the primitives to generate a frame buffer (e.g., pixel data for each of the pixels of a display).

[0126] An application writes model data (e.g., a collection of vertices and attributes) for a scene into memory, such as system memory or memory 304. The model data defines each of the objects that can be visible on a display. The application then makes an API call to the driver kernel, which requests the model data to be rendered and displayed. The driver kernel reads the model data and writes commands to one or more streams to perform operations to process the model data. These commands can reference different shading programs to be implemented on the SMs 440 of the PPU 300, including one or more of vertex shading, hull shading, domain shading, geometry shading, and pixel shading. For example, one or more of the SMs 440 can be configured to execute a vertex shading program that processes a number of vertices defined by the model data. In one embodiment, different ones of the SMs 440 can be configured to execute different shading programs at the same time. For example, a first subset of the SMs 440 can be configured to execute a vertex shading program while a second subset of the SMs 440 can be configured to execute a pixel shading program. The first subset of the SMs 440 processes the vertex data to produce processed vertex data, which is written to the L2 cache 460 and / or memory 304. After the processed vertex data is rasterized (e.g., converted from three-dimensional data to two-dimensional data in screen space) to produce fragment data, the second subset of the SMs 440 executes the pixel shading to produce processed fragment data, which is then blended with other processed fragment data and written to a frame buffer in memory 304. The vertex shading program and the pixel shading program can execute at the same time, processing different data from the same scene in a pipelined fashion until all of the model data for the scene has been rendered to the frame buffer. The contents of the frame buffer are then transmitted to a display controller to be displayed on a display device.

[0127] Machine Learning

[0128] Deep neural networks (DNNs) developed on processors such as the PPU 300 have been used for a variety of use cases: from self-driving cars to faster drug development, from automatic image captioning in online image databases to intelligent real-time language translation in video chat applications. Deep learning is a technology that models the neural learning process of the human brain, learns continuously, gets smarter over time, and delivers more accurate results faster over time. A child is initially taught by an adult to correctly identify and classify various shapes, and eventually is able to identify shapes without any coaching. Similarly, a deep learning or neural learning system needs to be trained in object identification and classification in order to become more intelligent and efficient in identifying basic objects, occluded objects, and the like, while also assigning context to objects.

[0129] At the simplest level, neurons in the human brain look at the various inputs received, assign a level of importance to each of these inputs, and pass the output to other neurons for processing. An artificial neuron or perceptron is the most basic model of a neural network. In one example, a perceptron can receive one or more inputs that represent various features of an object that the perceptron is being trained to recognize and classify, and each of these features is given a certain weight based on the importance of that feature in defining the shape of the object.

[0130] Deep neural network (DNN) models include multiple layers of connected nodes (e.g., perceptrons, Boltzmann machines, radial basis functions, convolutional layers, etc.) that can be trained with large amounts of input data to solve complex problems quickly and with high accuracy. In one example, the first layer of a DNN model breaks down an input image of a car into individual parts and looks for basic patterns such as lines and corners. The second layer assembles the lines to look for higher-level patterns such as wheels, windshields, and mirrors. The next layer identifies the type of vehicle, and the last few layers generate a label for the input image, identifying a specific make and model of car.

[0131] Once a DNN is trained, it can be deployed and used to recognize and classify objects or patterns in a process known as inference. Examples of inference (the process by which a DNN extracts useful information from a given input) include recognizing handwritten numbers on checks deposited into an ATM machine, recognizing images of friends in a photograph, providing movie recommendations to over 50 million users, recognizing and classifying different types of cars, pedestrians, and road hazards in a self-driving car, or translating human speech in real-time.

[0132] During training, data flows through the DNN in a forward propagation phase until a prediction is made that indicates a label corresponding to the input. If the neural network does not correctly label the input, the error between the correct label and the predicted label is analyzed and the weights are adjusted for each feature during a backward propagation phase until the DNN correctly labels that input and others in the training dataset. Training complex neural networks requires a large amount of parallel computing performance, including floating point multiplication and addition supported by PPU 300. Inference is less computationally intensive than training and is a latency-sensitive process in which a trained neural network is applied to new inputs that it has not seen before to classify images, translate speech, and generally infer new information.

[0133] Neural networks rely heavily on matrix math operations, and complex multi-layer networks require large amounts of floating point performance and bandwidth to improve efficiency and speed. With thousands of processing cores, optimized for matrix math operations, and delivering tens to hundreds of TFLOPS of performance, the PPU 300 is a computing platform capable of delivering the performance required for deep neural network-based artificial intelligence and machine learning applications.

[0134] Figure 6 is a block diagram of an example system architecture of an example autonomous vehicle 900 according to one embodiment of the present disclosure. It should be understood that this and other arrangements described herein are set forth only as examples. Other arrangements and elements (e.g., machines, interfaces, functions, orders, groupings of functions, etc.) can be used in addition to or instead of those shown, and some elements can be omitted altogether. Further, many of the elements described herein are functional entities that can be implemented as discrete or distributed components or in conjunction with other components, and in any suitable combination and location of hardware, firmware, and / or software. Various functions described herein as being performed by entities can be carried out by hardware, firmware, and / or software. For instance, various functions can be implemented by a processor executing instructions stored in memory.

[0135] Figure 6 Each of the components, features, and systems of the vehicle 900 in FIG. 1 is shown as being connected via a bus 902. The bus 902 can include a controller area network (CAN) data interface (alternatively referred to herein as a “CAN bus”). The CAN can be a network within the vehicle 900 that is used to help control various features and functions of the vehicle 900, such as actuation of brakes, acceleration, braking, steering, windshield wipers, etc. The CAN bus can be configured to have dozens or even hundreds of nodes, each with its own unique identifier (e.g., CAN ID). The CAN bus can be read to find steering wheel angle, ground speed, engine revolutions per minute (RPM), button positions, and / or other vehicle status indicators. The CAN bus can be ASIL B compliant.

[0136] Although bus 902 is described herein as a CAN bus, this is not intended to be limiting. For example, FlexRay and / or Ethernet can be used in addition to or in place of a CAN bus. Additionally, although a single line is used to represent bus 902, this is not intended to be limiting. For example, any number of buses 902 can be present, which can include one or more CAN buses, one or more FlexRay buses, one or more Ethernet buses, and / or one or more other types of buses using different protocols. In some examples, two or more buses 902 can be used to perform different functions, and / or can be used for redundancy. For example, a first bus 902 can be used for anti-collision functions, and a second bus 902 can be used for actuation control. In any example, each bus 902 can be in communication with any component of vehicle 900, and two or more buses 902 can be in communication with the same component. In some examples, each SoC 904, each controller 936, and / or each computer within a vehicle can have access to the same input data (e.g., input from sensors of vehicle 900), and can be connected to a common bus, such as a CAN bus.

[0137] Vehicle 900 can include one or more controllers 936, such as those described herein with respect to Figure 6 controllers. Controllers 936 can be used for various functions. Controllers 936 can be coupled to any of the various other components and systems of vehicle 900, and can be used to control vehicle 900, artificial intelligence of vehicle 900, infotainment of vehicle 900, etc.

[0138] Vehicle 900 can include a system on a chip (SoC) 904. SoC 904 can include one or more CPUs 906, one or more GPUs 908, one or more processors 910, one or more caches 912, one or more accelerators 914, one or more data stores 916, and / or other components and features not shown. SoC 904 can be used to control vehicle 900 in various platforms and systems. For example, one or more SoCs 904 can be combined in a system (e.g., a system of vehicle 900) with an HD map 922 that can obtain map refreshes and / or updates from one or more servers via a network interface 924.

[0139] In at least one embodiment, one or more of SoC(s) 904 can include one or more hardware components (e.g., CPU(s) 906, GPU(s) 906, processor(s) 910, cache(s) 912, accelerator(s) 914, and / or data store(s) 916) and / or one or more degradation detection systems.

[0140] CPU(s) 906 can include a CPU cluster or CPU complex (alternatively referred to herein as a “CCPLEX”). CPU(s) 906 can include multiple cores and / or L2 cache. For example, in some embodiments, CPU(s) 906 can include eight cores in a coherent multi-processor configuration. In some embodiments, CPU(s) 906 can include four dual-core clusters with each cluster having a dedicated L2 cache (e.g., 2 MB L2 cache). CPU(s) 906 (e.g., CCPLEX) can be configured to support simultaneous cluster operation such that any combination of clusters of CPU(s) 906 are active at any given time.

[0141] CPU(s) 906 can implement power management capabilities including one or more of the following features: individual hardware blocks can be automatically clock-gated when idle to save dynamic power; each core clock can be gated when the core is not actively executing instructions due to execution of WFI / WFE instructions; each core can be independently power-gated; each core cluster can be independently clock-gated when all cores are clock-gated or power-gated; and / or each core cluster can be independently power-gated when all cores are power-gated. CPU(s) 906 can also implement enhanced algorithms for managing power states with allowed power states and expected wake-up times specified, and hardware / microcode determines the best power state for the core, cluster, and CCPLEX to enter. The processing core can support a simplified power state entry sequence in software that offloads work to microcode.

[0142] GPU 908 can include an integrated GPU (alternatively referred to herein as an “iGPU”). GPU 908 can be programmable and efficient for parallel workloads. In some examples, one or more GPU 908 can use an enhanced tensor instruction set. GPU 908 can include one or more streaming microprocessors, where each streaming microprocessor can include an LI cache (e.g., an LI cache having at least 96 KB of storage capacity), and two or more of the streaming microprocessors can share an L2 cache (e.g., an L2 cache having 512 KB of storage capacity). In some embodiments, one or more GPU 908 can include at least eight streaming microprocessors. GPU 908 can use a compute application programming interface (API). Further, one or more GPU 908 can use one or more parallel computing platforms and / or programming models (e.g., NVIDIA’s CUDA).

[0143] GPU 908 can be power-optimized for best performance in automotive and embedded use cases. For example, GPU 908 can be fabricated on a fin field-effect transistor (FinFET). However, this is not intended to be limiting, and GPU 908 can be fabricated using other semiconductor manufacturing processes. Each streaming microprocessor can incorporate a number of mixed-precision processing cores that are partitioned into a number of blocks. For example, but not by way of limitation, 64 PF32 cores and 32 PF64 cores can be partitioned into four processing blocks. In this example, each processing block can be allocated 16 FP32 cores, 8 FP64 cores, 16 INT32 cores, two mixed-precision NVIDIA Tensor Cores for deep learning matrix operations, an L0 instruction cache, a warp scheduler, a dispatch unit, and / or a 64 KB register file. Further, the streaming microprocessor can include independent parallel integer and floating point data paths to provide efficient execution of workloads through a mix of computing and addressing computations. The streaming microprocessor can include independent thread scheduling capabilities to enable finer-grain synchronization and cooperation between parallel threads. The streaming microprocessor can include a combined LI data cache and shared memory unit to improve performance while simplifying programming.

[0144] GPU 908 can include a high-bandwidth memory (HBM) and / or a 16 GB HBM2 memory subsystem to provide, in some examples, about 900 GB / second peak memory bandwidth. In some instances, in addition to or instead of HBM memory, a synchronous graphics random access memory (SGRAM) can be used, such as a graphics double data rate type five synchronous random access memory (GDDR5).

[0145] GPU 908 can include a unified memory technology that includes access counters to allow memory pages to be migrated more accurately to the processors that access them most frequently, improving efficiency of memory ranges shared between processors. In some examples, address translation services (ATS) support can be used to allow GPU 908 to directly access CPU 906 page tables. In these instances, when a GPU 908 memory management unit (MMU) experiences a miss, an address translation request can be transmitted to CPU 906. In response, CPU 906 can look up a virtual to physical mapping for the address in its page tables and transmit the translation back to GPU 908. As such, the unified memory technology can allow a single unified virtual address space for memory of both CPU 906 and GPU 908, thereby simplifying programming of GPU 808 and porting of applications to GPU 908.

[0146] Further, GPU 908 can include access counters that can track how frequently GPU 908 accesses memory of other processors. The access counters can help ensure that memory pages are moved to the physical memory of the processor that accesses the page most frequently.

[0147] SoC 904 can include any number of caches 912, including those described herein. For example, one or more caches 912 can include an L3 cache that is available to both one or more CPUs 906 and one or more GPUs 908 (e.g., connected to both one or more CPUs 906 and one or more GPUs 908). One or more caches 912 can include a write-back cache that can track the state of a line, for example, by using a cache coherency protocol (e.g., MESI, MSI, etc.). Depending on the embodiment, the L3 cache can include 4MB or more, although smaller cache sizes can be used.

[0148] The SoC 904 can include one or more accelerators 914 (e.g., hardware accelerators, software accelerators, or a combination thereof). For example, the one or more SoCs 904 can include a hardware acceleration cluster that can include optimized hardware accelerators and / or large on-chip memory. The large on-chip memory (e.g., 4 MB of SRAM) can enable the hardware acceleration cluster to accelerate neural networks and other computations. The hardware acceleration cluster can be used to supplement the GPU 908 and offload some of the tasks of the GPU 908 (e.g., freeing up more cycles of the GPU 908 for performing other tasks). As an example, the accelerators 914 can be used for target workloads (e.g., perception, convolutional neural networks (CNNs), etc.) that are stable enough to be amenable to acceleration. The term “CNN” as used herein can include all types of CNNs, including region-based or regional convolutional neural networks (RCNNs) and fast RCNNs (e.g., as used for object detection).

[0149] The accelerators 914 (e.g., hardware acceleration cluster) can include a deep learning accelerator (DLA). The DLA can include one or more tensor processing units (TPUs) that can be configured to provide additional tera operations per second for deep learning applications and inferencing. The TPUs can be accelerators that are configured and optimized for performing image processing functions (e.g., for CNNs, RCNNs, etc.). The DLA can be further optimized for neural network types and specific sets of floating point operations, as well as inferencing. The design of the DLA can provide more performance per mm than general purpose GPUs and greatly outperform CPUs. The TPUs can perform several functions, including single instance convolution functions, support INT8, INT16, and FP16 data types for both features and weights, for example, and post-processor functions.

[0150] The DLA can perform neural networks (especially CNNs) on processed or unprocessed data quickly and efficiently for any of a variety of functions, including, for example, but not limited to: CNNs for object recognition and detection using data from a camera sensor; CNNs for distance estimation using data from a camera sensor; CNNs for emergency vehicle detection and identification and detection using data from a microphone; CNNs for face recognition and vehicle owner identification using data from a camera sensor; and / or CNNs for safety and / or safety related events.

[0151] The DLA can perform any of the functions of the one or more GPU(s) 908, and by using an inference accelerator, for example, a designer can target the one or more DLAs or the one or more GPU(s) 908 for any function. For example, a designer can concentrate the processing of CNNs and floating point operations on the DLA, and leave other functions to the GPU(s) 908 and / or other accelerator(s) 914.

[0152] The accelerator(s) 914 (e.g., hardware acceleration cluster) can include a programmable vision accelerator (PVA), which can be alternatively referred to herein as a computer vision accelerator. The PVA can be designed and configured to accelerate computer vision algorithms for advanced driver assistance systems (ADAS), autonomous driving, and / or augmented reality (AR) and / or virtual reality (VR) applications. The PVA can provide a balance between performance and flexibility. For example, each PVA can include, for example and without limitation, any number of reduced instruction set computer (RISC) cores, direct memory access (DMA), and / or any number of vector processors.

[0153] The RISC cores can interact with image sensors (e.g., image sensors of any of the cameras described herein), image signal processors, and / or the like. Each of the RISC cores can include any amount of memory. Depending on the embodiment, the RISC cores can use any of a variety of protocols. In some examples, the RISC cores can execute a real-time operating system (RTOS). The RISC cores can be implemented using one or more integrated circuit devices, application specific integrated circuits (ASICs), and / or memory devices. For example, the RISC cores can include an instruction cache and / or a tightly coupled RAM.

[0154] The DMA can enable components of the PVA to access system memory independently of the CPU(s) 906. The DMA can support any number of features for providing optimizations to the PVA, including but not limited to supporting multi-dimensional addressing and / or circular addressing. In some examples, the DMA can support up to six or more dimensions of addressing, which can include block width, block height, block depth, horizontal block stride, vertical block stride, and / or depth stride.

[0155] The vector processors can be programmable processors that can be designed to efficiently and flexibly execute programming for computer vision algorithms and provide signal processing capabilities. In some examples, a PVA can include a PVA core and two vector processing subsystem partitions. The PVA core can include a processor subsystem, one or more DMA engines (e.g., two DMA engines), and / or other peripherals. The vector processing subsystems can operate as the main processing engines of the PVA and can include vector processing units (VPUs), instruction caches, and / or vector memories (e.g., VMEMs). The VPU cores can contain digital signal processors, such as single instruction multiple data (SIMD), very long instruction word (VLIW) digital signal processors. The combination of SIMD and VLIW can enhance throughput and speed.

[0156] Each of the vector processors in the vector processor can include an instruction cache and can be coupled to a dedicated memory. Thus, in some examples, each of the vector processors can be configured to execute independently of the other vector processors. In other examples, the vector processors included in a particular PVA can be configured to employ data parallelism. For example, in some embodiments, multiple vector processors included in a single PVA can execute the same computer vision algorithm, but on different regions of an image. In other examples, the vector processors included in a particular PVA can execute different computer vision algorithms on the same image simultaneously, or even different algorithms on sequential images or portions of an image. Among other things, any number of PVAs can be included in the hardware acceleration cluster, and any number of vector processors can be included in each PVA. Furthermore, the PVAs can include additional error-correcting code (ECC) memory to enhance overall system security.

[0157] The accelerator 914 (e.g., hardware acceleration cluster) can include a computer vision network-on-chip and SRAM for providing high bandwidth, low latency SRAM for the accelerator 914. In some examples, the on-chip memory can include at least 4 MB of SRAM, composed of, for example, but not limited to, eight field configurable memory blocks, which can be accessed by the PVAs and the DLAs. Each pair of memory blocks can include an advanced peripheral bus (APB) interface, configuration circuitry, a controller, and a multiplexer. Any type of memory can be used. The PVAs and the DLAs can access the memory via a backbone that provides the PVAs and the DLAs with high speed access to the memory. The backbone can include a computer vision network-on-chip that interconnects (e.g., using APB) the PVAs and the DLAs with the memory.

[0158] The computer vision processor on a network can include an interface that determines both PVA and DLA provide ready and valid signals before transmitting any control signals / addresses / data. Such an interface can provide separate phases and separate channels for transmitting control signals / addresses / data, as well as burst-type communication for continuous data transmission. This type of interface can comply with ISO 26262 or IEC 61508 standards, although other standards and protocols can be used.

[0159] In some examples, one or more SoCs 904 can include a real-time ray tracing hardware accelerator, as described in U.S. Patent Application No. 16 / 101,232, filed August 10, 2018. The real-time ray tracing hardware accelerator can be used to quickly and efficiently determine locations and extents of objects (e.g., within a world model) for purposes of generating real-time visualizations simulations, for RADAR signal interpretation, for sound propagation synthesis and / or analysis, for simulation of SONAR systems, for general wave propagation simulation, for comparison with LIDAR data, for localization and / or other functions, and / or for other uses.

[0160] The accelerator 914 (e.g., hardware accelerator cluster) has wide use for autonomous driving. The PVA can be a programmable vision accelerator that can be used for key processing stages in ADAS and autonomous vehicles. The capabilities of the PVA are a good match for algorithmic domains that require predictable processing at low power and low latency. In other words, the PVA performs well for semi-dense or dense regular computations, even for small data sets, which require predictable runtimes with low latency and low power. Thus, in the context of a platform for autonomous vehicles, the PVA is designed to run classical computer vision algorithms because they are efficient in terms of object detection and in terms of integer math operations.

[0161] For example, according to one embodiment of the present technology, the PVA is used to perform computer stereo vision. In some instances, a semi-global matching based algorithm can be used, although this is not intended to be limiting. Many applications for Level 3-5 autonomous driving require immediate motion estimation / stereo matching (e.g., structure from motion, pedestrian recognition, lane detection, etc.). The PVA can perform computer stereo vision functions on inputs from two monocular cameras.

[0162] In some examples, the PVA can be used to perform dense optical flow. According to processing raw RADAR data (e.g., using a 4D fast Fourier transform) to provide processed RADAR. In other examples, the PVA is used for time-of-flight depth processing, such as by processing raw time-of-flight data to provide processed time-of-flight data.

[0163] The DLA can be used to run any type of network to enhance control and driving safety, including, for example, a neural network that outputs a confidence measure for each object detection. This confidence value can be interpreted as a probability, or as providing a relative “weight” for each detection compared to other detections. The confidence value enables the system to make further decisions about which detections should be considered as true positive detections and not false positive detections. For example, the system can set a threshold for confidence and only consider detections that exceed the threshold as true positive detections. In an automatic emergency braking (AEB) system, a false positive detection would cause the vehicle to automatically perform an emergency brake, which is obviously undesirable. Thus, only the most confident detections should be considered as triggers for AEB. The DLA can run a neural network for regression of a confidence value. The neural network can take as its input at least some subset of parameters such as bounding box size, ground plane estimate (e.g., obtained from another subsystem), vehicle 900 orientation, distance, 3D position estimate of the object obtained from the neural network and / or other sensors (e.g., one or more LIDAR sensors 964 or one or more RADAR sensors 960), inertial measurement unit (IMU) sensor 966 output related to the 3D position estimate of the object obtained from the neural network and / or other sensors, etc.

[0164] The one or more SoCs 904 can include one or more data stores 916 (e.g., memory). The data stores 916 can be on-chip memory of the one or more SoCs 904, which can store neural networks to be executed on the GPU and / or DLA. In some examples, the one or more data stores 916 can be large enough in capacity to store multiple instances of a neural network for redundancy and safety. The data stores 912 can include L2 or L3 caches 912. References to the data stores 916 can include references to memory associated with the PVA, DLA, and / or other accelerators 914, as described herein.

[0165] SoC 904 can include one or more processors 910 (e.g., embedded processors). The processors 910 can include a boot and power management processor, which can be a special purpose processor and subsystem to handle boot power and management functions and related security enforcement. The boot and power management processor can be part of the SoC 904 boot sequence and can provide runtime power management services. The boot power and management processor can provide clock and voltage programming, assistance in system low power state transitions, management of one or more SoC 904 thermal and temperature sensors, and / or management of one or more SoC 904 power states. Each temperature sensor can be implemented as a ring oscillator whose output frequency is proportional to temperature, and the SoC 904 can use the ring oscillator to detect the temperature of the CPU 906, GPU 908, and / or accelerators 914. If it is determined that the temperature exceeds a threshold, the boot and power management processor can enter a temperature fault routine and place the SoC 904 in a lower power state and / or place the vehicle 900 in a driver-to-safe-stop mode (e.g., bring the vehicle 900 to a safe stop).

[0166] The processors 910 can also include a set of embedded processors that can act as an audio processing engine. The audio processing engine can be an audio subsystem that enables full hardware support for multi-channel audio on multiple interfaces, as well as a broad and flexible range of audio I / O interfaces. In some examples, the audio processing engine is a dedicated processor core with a digital signal processor (with dedicated RAM).

[0167] The processors 910 can also include an always-on processor engine that can provide the necessary hardware features to support low-power sensor management and wake-up use cases. The always-on processor engine can include a processor core, tightly coupled RAM, supporting peripherals (e.g., timers and interrupt controllers), different I / O controller peripherals, and routing logic.

[0168] The processors 910 can also include a security cluster engine that includes a dedicated processor subsystem for handling security management for automotive applications. The security cluster engine can include two or more processor cores, tightly coupled RAM, supporting peripherals (e.g., timers, interrupt controllers, etc.), and / or routing logic. In a secure mode, the two or more cores can operate in a lockstep mode and act as a single core with comparison logic to detect any differences between their operations.

[0169] The processor 910 can include a video image compositor, which can be a processing block (e.g., implemented on a microprocessor) that implements video post-processing functions required by a video playback application to produce the final image for the player window. The video image compositor can perform lens distortion correction on the wide field of view camera 970, surround camera 974, and / or cabin interior monitoring camera sensors. The cabin interior monitoring camera sensors are preferably monitored by a neural network running on another instance of the advanced SoC that is configured to identify events in the cabin and respond accordingly.

[0170] The video image compositor can also be configured to perform stereoscopic correction on input stereoscopic lens frames. The video image compositor can also be used for user interface composition when the operating system desktop is in use, and does not require the GPU 908 to continuously render new surfaces. The video image compositor can be used to offload the GPU 908 to improve performance and responsiveness even when the GPU 908 is powered on and actively doing 3D rendering.

[0171] The SoC 904 can also include a Mobile Industry Processor Interface (MIPI) camera serial interface for receiving video and input from cameras, a high-speed interface, and / or a video input block that can be used for camera and related pixel input functions. The SoC 904 can also include an input / output controller that can be controlled by software and can be used to receive I / O signals that are not committed to a specific role.

[0172] The SoC 904 can also include a wide range of peripheral interfaces to enable communication with peripherals, audio codecs, power management, and / or other devices. The SoC 904 can be used to process data from cameras (e.g., over a gigabit multimedia serial link and an Ethernet connection), sensors (e.g., LIDAR sensor 964, RADAR sensor 960, etc.), data from the bus 902 (e.g., speed of the vehicle 900, steering wheel position, etc.), data from the GNSS sensor 958 (e.g., over an Ethernet or CAN bus connection). The SoC 904 can also include a dedicated high-performance mass storage controller that can include its own DMA engine and can be used to free the CPU 906 from routine data management tasks.

[0173] SoC 904 can be an end-to-end platform with a flexible architecture spanning automation levels 3-5, thereby providing a comprehensive functional safety architecture that leverages computer vision and ADAS technology for a flexible, reliable driving software stack with deep learning tools. SoC 904 can be faster, more reliable, and even more energy and space efficient than conventional systems. For example, one or more accelerators 914, when combined with one or more CPUs 906, one or more GPUs 908, and one or more data stores 916, can provide a fast, efficient platform for level 3-5 autonomous vehicles.

[0174] Vehicle 900 can also include a network interface 924, which can include one or more wireless antennas 926 (e.g., one or more wireless antennas for different communication protocols, such as cellular antennas, Bluetooth antennas, etc.). Network interface 924 can be used to enable wireless connections over the Internet with a cloud (e.g., with server 978 and / or other network devices), with other vehicles, and / or with computing devices (e.g., a client device of a passenger). For communication with other vehicles, a direct link can be established between two vehicles and / or an indirect link can be established (e.g., across a network and over the Internet). A vehicle-to-vehicle communication link can be used to provide the direct link. The vehicle-to-vehicle communication link can provide vehicle 900 information about vehicles in the vicinity of vehicle 900 (e.g., vehicles in front of, to the side of, and / or behind vehicle 900). This functionality can be part of a cooperative adaptive cruise control functionality of vehicle 900.

[0175] Network interface 924 can include a SoC that provides modulation and demodulation functionality and enables one or more controllers 936 to communicate over a wireless network. Network interface 924 can include a radio frequency front end for upconversion from baseband to radio frequency and downconversion from radio frequency to baseband. Frequency conversion can be performed through well-known processes and / or can be performed using a superheterodyne process. In some examples, radio frequency front end functionality can be provided by a separate chip. Network interface can include wireless functionality for communicating over LTE, WCDMA, UMTS, GSM, CDMA2000, Bluetooth, Bluetooth LE, Wi-Fi, Z-Wave, ZigBee, LoRaWAN, and / or other wireless protocols.

[0176] Vehicle 900 can also include one or more data stores 928, which can include off-chip (e.g., off-SoC 904) storage. Data stores 928 can include one or more storage elements, including RAM, SRAM, DRAM, VRAM, flash memory, hard disks, and / or other components and / or devices that can store at least one bit of data.

[0177] The vehicle 900 can also include one or more GNSS sensors 958. The GNSS sensors 958 (e.g., GPS and / or assisted GPS sensors) are used to assist mapping, perception, occupancy grid generation, and / or path planning functions. Any number of GNSS sensors 958 can be used, including, for example and without limitation, a GPS using a USB connector with an Ethernet-to-serial (RS-232) bridge.

[0178] The vehicle 900 can also include one or more RADAR sensors 960. The RADAR sensors 960 can be used by the vehicle 900 for long-range vehicle detection, even in darkness and / or adverse weather conditions. The RADAR functional safety level can be ASIL B. In some examples, the one or more RADAR sensors 960 can use the CAN and / or the bus 902 (e.g., for communicating data generated by the one or more RADAR sensors 960) to control and access object tracking data, with Ethernet access to raw data. A wide variety of RADAR sensor types can be used. For example and without limitation, the RADAR sensors 960 can be suitable for front, rear, and side use of RADAR. In some examples, one or more pulsed Doppler RADAR sensors are used.

[0179] The RADAR sensors 960 can include different configurations, such as long-range with narrow field of view, short-range with wide field of view, short-range side coverage, etc. In some examples, long-range RADAR can be used for adaptive cruise control functionality. Long-range RADAR systems can provide a wide field of view implemented by two or more independent scans, such as over a 250 m range. The RADAR sensors 960 can help distinguish between static and moving objects, and can be used by the ADAS system for emergency brake assist and forward collision warning. Long-range RADAR sensors can include a single-station multi-mode RADAR with multiple (e.g., six or more) fixed RADAR antennas and high-speed CAN and FlexRay interfaces. In examples with six antennas, the central four antennas can create a focused beam pattern designed to record the surroundings of the vehicle 900 at higher speeds with minimal interference from traffic in adjacent lanes. The other two antennas can expand the field of view so that vehicles entering or leaving the lane of the vehicle 900 can be detected quickly.

[0180] As an example, a mid-range RADAR system can include a range of up to 860 m (front) or 80 m (rear), and a field of view of up to 42 degrees (front) or 850 degrees (rear). A short-range RADAR system can include, but is not limited to, a RADAR sensor designed to be mounted at both ends of the rear bumper. When mounted at both ends of the rear bumper, such a RADAR sensor system can produce two beams that continuously monitor the blind spots behind and alongside the vehicle.

[0181] A short-range RADAR system can be used in an ADAS system for blind spot detection and / or lane change assist.

[0182] The vehicle 900 can also include one or more ultrasonic sensors 962. The ultrasonic sensors 962, which can be positioned in the front, rear, and / or sides of the vehicle 900, can be used for parking assist and / or to create and update an occupancy grid. A variety of ultrasonic sensors 962 can be used, and different ultrasonic sensors 962 can be used for different detection ranges (e.g., 2.5 m, 4 m). The ultrasonic sensors 962 can operate at a functional safety level of ASIL B.

[0183] The vehicle 900 can include a LIDAR sensor 964. The LIDAR sensor 964 can be used for object and pedestrian detection, emergency braking, collision avoidance, and / or other functions. The LIDAR sensor 964 can be at a functional safety level of ASIL B. In some examples, the vehicle 900 can include multiple LIDAR sensors 964 (e.g., two, four, six, etc.) that can use Ethernet (e.g., to provide data to a Gigabit Ethernet switch).

[0184] In some examples, one or more LIDAR sensors 964 can be capable of providing a 360-degree field of view of object lists and their distances. For example, a commercially available LIDAR sensor 964 can have an advertised range of approximately 800 m, with an accuracy of 2 cm - 3 cm, and support an 800 Mbps Ethernet connection. In some examples, one or more flush-mounted LIDAR sensors 964 can be used. In such examples, the one or more LIDAR sensors 964 can be implemented as small devices that can be embedded into the front, rear, sides, and / or corners of the vehicle 900. In such examples, the one or more LIDAR sensors 964 can provide a horizontal field of view of up to 820 degrees and a vertical field of view of 35 degrees, with a range of 200 m even for low reflectivity objects. A front-mounted LIDAR sensor 964 can be configured for a horizontal field of view between 45 degrees and 135 degrees.

[0185] In some examples, LIDAR technology can also be used, such as 3D Flash LIDAR. 3D Flash LIDAR uses a flash of laser light as a source of emission, illuminating up to approximately 200 m around the vehicle. The Flash LIDAR unit includes a receiver that records the laser pulse transit time and reflected light on each pixel, which in turn corresponds to the distance from the vehicle to the object.

[0186] The vehicle can also include one or more IMU sensors 966. In some examples, the IMU sensors 966 can be located at the center of the rear axle of the vehicle 900. The IMU sensors 966 can include, for example and without limitation, accelerometers, magnetometers, gyroscopes, magnetic compasses, and / or other sensor types. In some examples, such as in six-axis applications, the one or more IMU sensors 966 can include accelerometers and gyroscopes, while in nine-axis applications, the one or more IMU sensors 966 can include accelerometers, gyroscopes, and magnetometers.

[0187] In some embodiments, the IMU sensors 966 can be implemented as a miniature high-performance GPS-aided inertial navigation system (GPS / INS) that combines micro-electromechanical systems (MEMS) inertial sensors, high-sensitivity GPS receivers, and advanced Kalman filtering algorithms to provide estimates of position, velocity, and attitude. As such, in some examples, the one or more IMU sensors 966 can enable the vehicle 900 to estimate heading by directly observing and correlating changes in velocity from GPS to the one or more IMU sensors 966 without input from magnetic sensors. In some examples, the IMU sensors 966 and the GNSS sensors 958 can be combined in a single integrated unit.

[0188] The vehicle can include microphones 996 placed in and / or around the vehicle 900. The microphones 996 can be used, among other things, for emergency vehicle detection and identification.

[0189] The vehicle can also include any number of camera types, including one or more stereo cameras 968, one or more wide-view cameras 970, one or more infrared cameras 972, one or more surround-view cameras 974, one or more long-range and / or mid-range cameras 998, and / or other camera types. The cameras can be used to capture image data around the entire perimeter of the vehicle 900. The types of cameras used depend on the embodiment and requirements of the vehicle 900, and any combination of camera types can be used to provide the necessary coverage around the vehicle 900. Further, the number of cameras can vary depending on the implementation. For example, the vehicle can include six cameras, seven cameras, ten cameras, twelve cameras, and / or another number of cameras. The cameras can support, for example and without limitation, Gigabit Multimedia Serial Link (GMSL) and / or Gigabit Ethernet.

[0190] The vehicle 900 can also include a vibration sensor 942. The vibration sensor 942 can measure vibrations of components of the vehicle, such as axles. For example, changes in vibration can indicate changes in the road. In another example, when two or more vibration sensors 942 are used, differences between the vibrations can be used to determine road friction or slippage (e.g., when the vibration difference is between a power driven axle and a free spinning axle).

[0191] The vehicle 900 can include an ADAS system 938. In some examples, the ADAS system 938 can include a SoC. The ADAS system 938 can include adaptive / automatic / autonomous cruise control (ACC), cooperative adaptive cruise control (CACC), forward collision warning (FCW), automatic emergency braking (AEB), lane departure warning (LDW), lane keep assist (LKA), blind spot warning (BSW), rear cross-traffic warning (RCTW), collision warning system (CWS), lane centering (LC), and / or other features and functionality.

[0192] In at least one embodiment, any of the different one or more actions from the remedial action manager can be with respect to one or more of the ADAS system 938 and / or the above-mentioned functionality (e.g., indicators, disablement, logging, etc.).

[0193] The ACC system can use one or more RADAR sensors 960, one or more LIDAR sensors 964, and / or one or more cameras. The ACC system can include longitudinal ACC and / or lateral ACC. Longitudinal ACC monitors and controls the distance to the vehicle directly in front of the vehicle 900 and automatically adjusts the vehicle speed to maintain a safe distance from the vehicle in front. Lateral ACC makes distance keeping and informs the vehicle 900 to make a lane change as needed. Side ACC is related to other ADAS applications such as LCA and CWS.

[0194] CACC uses information from other vehicles that can be received from other vehicles indirectly via a network connection (e.g., over the Internet) or via a wireless link via the network interface 924 and / or the wireless antenna 926. The direct link can be provided by a vehicle-to-vehicle (V2V) communication link, while the indirect link can be an infrastructure-to-vehicle (12V) communication link. Generally, the V2V communication concept provides information about the vehicle directly in front (e.g., the vehicle directly in front and in the same lane as the vehicle 900), while the 12V communication concept provides information about more forward traffic. The CACC system can include one or both of the 12V and V2V information sources. Given the information about vehicles in front of the vehicle 900, CACC can be more reliable and potentially improve traffic flow and reduce congestion on the road.

[0195] FCW systems are designed to warn the driver of a hazard so that the driver can take corrective action. FCW systems use a front-facing camera and / or RADAR sensor 960 coupled to a dedicated processor, DSP, FPGA, and / or ASIC that is electrically coupled to driver feedback such as a display, speaker, and / or vibration assembly. FCW systems can provide warnings such as in the form of a sound, visual warning, vibration, and / or rapid brake pulse.

[0196] AEB systems detect an impending forward collision with another vehicle or other object and can automatically apply brakes if the driver does not take corrective action within specified time or distance parameters. AEB systems can use one or more front-facing cameras and / or one or more RADAR sensors 960 coupled to a dedicated processor, DSP, FPGA, and / or ASIC. When an AEB system detects a hazard, it typically first warns the driver to take corrective action to avoid a collision, and if the driver does not take corrective action, the AEB system can automatically apply the brakes in an effort to prevent or at least mitigate the impact of a predicted collision. AEB systems can include technologies such as dynamic brake support and / or imminent collision braking.

[0197] The vehicle 900 can also include an infotainment SoC 930 (e.g., an in-vehicle infotainment system (IVI)). Although illustrated and described as a SoC, the infotainment system can not be a SoC and can include two or more discrete components. The infotainment SoC 930 can include a combination of hardware and software that can be used to provide audio (e.g., music, personal digital assistant, navigation instructions, news, radio, etc.), video (e.g., TV, movies, streaming, etc.), telephony (e.g., hands-free calling), network connectivity (e.g., LTE, WiFi, etc.), and / or information services (e.g., navigation system, rear park assist, radio data system, vehicle related information such as fuel level, total distance traveled, brake fuel level, oil level, doors open / close, air filter information, etc.) to the vehicle 900. For example, the infotainment SoC 930 can include a radio, disc player, navigation system, video player, USB and Bluetooth connectivity, garage, in-car entertainment, WiFi, steering wheel audio controls, hands-free voice controls, heads-up display (HUD), HMI display 934, telematics equipment, control panel (e.g., for controlling and / or interacting with different components, features, and / or systems), and / or other components.

[0198] The infotainment SoC 930 can include GPU functionality. The infotainment SoC 930 can communicate with other devices, systems, and / or components of the vehicle 900 over the bus 902 (e.g., a CAN bus, Ethernet, etc.). In some instances, the infotainment SoC 930 can be coupled to a supervisory MCU such that the GPU of the infotainment system can perform some self-driving functions in the event of a failure of the primary controller(s) 936 (e.g., a primary computer and / or a backup computer of the vehicle 900). In such examples, the infotainment SoC 930 can place the vehicle 900 into a driver-to-safe-stop mode as described herein.

[0199] The vehicle 900 can also include an instrument cluster 932 (e.g., a digital instrument cluster, an electronic instrument cluster, a digital instrument cluster, etc.). The instrument cluster 932 can include a controller and / or supercomputer (e.g., a discrete controller or supercomputer). The instrument cluster 932 can include a set of gauges such as a speedometer, a fuel level, an oil pressure, a tachometer, an odometer, a turn indicator, a gear indicator, one or more seatbelt warning lights, one or more parking brake warning lights, one or more engine malfunction lights, a supplemental restraint system (SRS) system information, lighting controls, safety system controls, navigation information, etc. In some examples, information can be displayed and / or shared between the infotainment SoC 930 and the instrument cluster 932. In other words, the instrument cluster 932 can be included as part of the infotainment SoC 930 or vice versa.

[0200] In at least one embodiment, one or more indicators provided by the remedial action manager can be presented and / or displayed using one or more of the infotainment SoC 930, the instrument cluster 932, or the HMI display 934.

[0201] It should be noted that the techniques described herein can be embodied in executable instructions stored in a computer-readable medium, which are used or in conjunction with a processor-based instruction execution machine, system, device, or apparatus. Those skilled in the art will understand that, for some embodiments, different types of computer-readable media may be included for storing data. As used herein, “computer-readable medium” includes one or more suitable media for storing executable instructions of a computer program such that an instruction execution machine, system, device, or apparatus can read (or retrieve) the instructions from the computer-readable medium and execute the instructions for performing the described embodiments. Suitable storage formats include one or more of electronic, magnetic, optical, and electromagnetic formats. A non-exhaustive list of conventional exemplary computer-readable media includes: portable computer disks; random access memory (RAM); read-only memory (ROM); erasable programmable read-only memory (EPROM); flash memory devices; and optical storage devices, including portable compact discs (CDs), portable digital video discs (DVDs), etc.

[0202] It should be understood that the arrangement of components shown in the accompanying drawings is for illustrative purposes and other arrangements are possible. For example, one or more of the elements described herein may be implemented wholly or partially as electronic hardware components. Other elements may be implemented in software, hardware, or a combination of software and hardware. Furthermore, some or all of these other elements may be combined, some may be omitted entirely, and additional components may be added while still achieving the functionality described herein. Thus, the subject matter described herein can be embodied in many different variations, and all such variations are considered to be within the scope of the claims.

[0203] To facilitate understanding of the subject matter described herein, numerous aspects are described in terms of sequences of actions. Those skilled in the art will recognize that different actions can be performed by dedicated circuitry or circuits, by program instructions executed by one or more processors, or by a combination of both. The description of any sequence of actions herein is not intended to imply that a specific order for performing that sequence must be followed. Unless otherwise indicated herein or where the context clearly conflicts, all methods described herein can be performed in any suitable order.

[0204] The use of the terms “a” and “an” and “the” and similar referents in the context of describing the subject matter (especially in the context of the following claims) are to be construed to cover both the singular and the plural, unless otherwise indicated herein or clearly contradicted by context. The use of the term “at least one” followed by a list of one or more items (for example, “at least one of A and B”) is to be construed to mean one or more items from the listed items (A or B or both A and B) individually or any combination thereof. Furthermore, to the extent that the term “includes” is used in either the detailed description or the claims, such term is intended to be inclusive in a manner similar to the term “comprising” as “comprising” is interpreted when employed as a transitional word in the introductory clauses of the claims. Additionally, the indefinite articles “a” and “an” are used in the detailed description and in the claims solely to limit the context of the items being described to the items specifically recited in the immediately preceding sentence. The indefinite articles “a” and “an” are not intended to be construed as limiting the subject matter to the items specifically identified in the immediately preceding sentence. Furthermore, the use of the term “based on” and other like phrases (e.g., “based on the [item or term]”) is not meant to be construed as limiting the subject matter to only the item or term immediately following the phrase. The language of the specification is intended to be interpreted broadly in accordance with the principles of the subject matter.

Claims

1. A computer-implemented method comprising: executing at least a portion of a test on a hardware component of an in-field computing platform to produce a first test result, wherein the test comprises a first sequence of test patterns, a first value being used for a physical operating parameter of the hardware component applied during execution of the test; storing, by an in-system test (IST) controller within the in-field computing platform, the first test result in a memory; updating, in response to a command generated by the IST controller, the physical operating parameter to use a second value based on the first test result; dynamically determining a second sequence of the test patterns to apply to the hardware component, wherein the second sequence comprises at least one test pattern from the first sequence of the test patterns; and resuming execution of a second portion of the test comprising the second sequence of the test patterns on the hardware component using the second value of the physical operating parameter to produce a second test result.

2. The computer-implemented method of claim 1, wherein the physical operating parameter is at least one of: a supply voltage; a supply current; a clock speed; a noise magnitude; a noise duration; or a temperature.

3. The computer-implemented method of claim 1, wherein the second value is determined by the IST controller.

4. The computer-implemented method of claim 1, wherein the second value is determined by a central processing unit (CPU) coupled to the hardware component.

5. The computer-implemented method of claim 1, wherein the first test result indicates a pass, and further comprising, when the second test result indicates a fail, determining that the second value is a cutoff value for the physical operating parameter.

6. The computer-implemented method of claim 1, wherein the test comprises at least one of: a permanent fault test; or a functional test.

7. The computer-implemented method of claim 6, wherein the test comprises a permanent fault test represented as one or more structural vectors.

8. The computer-implemented method of claim 6, wherein the test comprises a functional test represented as one or more functional vectors.

9. The computer-implemented method of claim 1, further comprising: waiting for a predetermined duration of time after resuming execution of the test on the hardware component before checking the second test result; and rebooting the in-field computing platform when the predetermined duration of time expires and execution of the test is not complete. transmitting the command from the IST controller to an external component that provides the physical operating parameter to the hardware component.

10. The computer-implemented method of claim 1, wherein updating the physical operating parameter to use the second value comprises:

11. A system comprising: an in-system test (IST) controller within an in-field computing platform and coupled to a memory, the IST controller configured to: ​ performing at least a portion of a test on a hardware component of the in-field computing platform to produce a first test result, wherein the test includes a first sequence of test patterns, a first value is used for a physical operating parameter of the hardware component applied during performance of the test; storing, by the IST controller, the first test result in the memory; updating, in response to a command generated by the IST controller, the physical operating parameter to use a second value based on the first test result; determining dynamically a second sequence of the test patterns to apply to the hardware component, wherein the second sequence includes at least one test pattern from the first sequence of the test patterns; and resuming performance of a second portion of the test including the second sequence of the test patterns on the hardware component using the second value of the physical operating parameter to produce a second test result.

12. The system of claim 11, wherein the physical operating parameter is at least one of: a supply voltage; a supply current; a clock speed; a noise magnitude; a noise duration; or a temperature.

13. The system of claim 11, wherein the second value is determined by the IST controller.

14. The system of claim 11, wherein the first test result indicates a pass, and the IST controller is further configured to determine that the second value is a cutoff value for the physical operating parameter when the second test result indicates a fail.

15. The system of claim 11, wherein the IST controller is further configured to update the physical operating parameter to use the second value by transmitting the command to an external component that provides the physical operating parameter to the hardware component.

16. The system of claim 11, wherein the system includes at least one of: an autonomous or semi-autonomous vehicle; an autonomous or semi-autonomous machine; an autonomous or semi-autonomous industrial robot; an autonomous or semi-autonomous robot; a manned or unmanned aerial vehicle; or a manned or unmanned watercraft.

17. The system of claim 11, wherein the system includes at least one of: a computing server system; a data center; a system on a chip (SoC); or an embedded system.

18. A non-transitory computer-readable medium storing computer instructions that, when executed by one or more processors, cause the one or more processors to perform the steps of: performing at least a portion of the test on a hardware component of an on-site computing platform to produce a first test result, wherein, the test includes a first sequence of test patterns, a first value is used for a physical operating parameter of the hardware component applied during performance of the test; storing, by an in-system test (IST) controller within the in-field computing platform, the first test result in a memory; updating, in response to a command generated by the IST controller, the physical operating parameter to use a second value based on the first test result; determining dynamically a second sequence of the test patterns to apply to the hardware component, wherein the second sequence includes at least one test pattern from the first sequence of the test patterns; and resuming performance of a second portion of the test including the second sequence of the test patterns on the hardware component using the second value of the physical operating parameter to produce a second test result. resuming execution of a second portion of the test comprising the second sequence of test patterns on the hardware component using the second value of the physical operating parameter to produce a second test result.

19. The non-transitory computer readable medium of claim 18, wherein the physical operating parameter is at least one of: a supply voltage; a supply current; a clock speed; a noise magnitude; a noise duration; or a temperature.

20. The non-transitory computer-readable medium of claim 18, wherein updating the physical operating parameter to use the second value comprises: transmitting the command from the IST controller to an external component that provides the physical operating parameter to the hardware component.

Citation Information

Patent Citations

  • Method for programmable timeouts of tree traversal mechanisms in hardware

    US10885698B2

  • Enhanced in-system test coverage based on detecting component degradation

    US11693753B2

  • Input voltage reduction for processing devices

    US20170357310A1