System, apparatus and method to debug accelerator hardware

Through example debugging circuits and applications, efficient debugging of hardware accelerators is achieved, which solves the time-consuming and complex problems in existing technologies and improves the debugging efficiency and accuracy of AI/ML models.

CN120653535APending Publication Date: 2025-09-16INTEL CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510975793.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2021-09-23
Filing Date
2022-08-23
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Debugging hardware accelerators is a time-consuming and complex process, especially in AI/ML models, where parallel core execution and a lack of dedicated hardware support make it difficult to effectively locate errors and isolate performance bottlenecks.

Method used

Using sample debug circuits and debug applications, it is possible to stop hardware accelerator output at a specified breakpoint, step through subsequent output transactions, stop AI/ML model execution in response to specific data generation, and output transactions to identify data at a specified point in time, reducing the difficulty of error location and performance optimization.

Benefits of technology

It improves the efficiency and accuracy of hardware accelerator debugging, reduces the generation of erroneous output, and simplifies the problem location and correction process in a parallel core environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120653535A_ABST
    Figure CN120653535A_ABST
Patent Text Reader

Abstract

Methods, apparatus, systems, and articles of manufacture to debug hardware accelerators, such as neural network accelerators, for performing artificial intelligence computing workloads are disclosed. An example apparatus includes a core having a core input and a core output to execute executable code based on a machine learning model to generate a data output based on a data input, and a debug circuit coupled to the core. The debug circuit is configured to detect a breakpoint associated with a machine learning model, and compile executable code based on at least one of the machine learning model or the breakpoint. In response to triggering of the breakpoint, the debug circuitry is to stop execution of the executable code and output data such as data inputs, data outputs, and breakpoints for debugging the hardware accelerator.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates generally to hardware accelerators, and more particularly to systems, apparatuses, and methods for debugging hardware accelerators. Background Art

[0002] In recent years, the demand for computationally intensive processing capabilities, such as artificial intelligence / machine learning and image processing, has expanded beyond high-powered dedicated desktop hardware and has become a necessity for personal and / or other mobile devices. Hardware accelerators can be included in such devices to enable these capabilities. Debugging such hardware accelerators is a time-consuming and complex task. BRIEF DESCRIPTION OF THE DRAWINGS

[0003] Figure 1 is an illustration of an example computing system including an example accelerator circuit with example debug circuitry to enable improved debugging of the accelerator circuit.

[0004] Figure 2 yes Figure 1 A block diagram of an example implementation of an accelerator circuit and a debug circuit.

[0005] Figure 3 yes Figure 1 A block diagram of another example implementation of an accelerator circuit and a debug circuit is shown.

[0006] Figure 4 yes Figure 1 A block diagram of yet another example implementation of an accelerator circuit and a debug circuit.

[0007] Figure 5 yes Figure 1 A block diagram of another example implementation of an accelerator circuit and a debug circuit is shown.

[0008] Figure 6 yes Figure 1 A block diagram of yet another example implementation of an accelerator circuit and a debug circuit.

[0009] Figure 7 yes Figure 1 A block diagram of another example implementation of an accelerator circuit and a debug circuit is shown.

[0010] Figure 8A yes Figure 1 A block diagram of an example implementation of a debug circuit for debugging Figure 1 An example accelerator circuit for a read operation.

[0011] Figure 8B yes Figure 1 A block diagram of an example implementation of a debug circuit for debugging Figure 1 The write operation of the accelerator circuit.

[0012] Figure 8C yes Figure 1 A block diagram of another example implementation of a debug circuit for debugging Figure 1 The accelerator circuit read operation.

[0013] Figure 8D yes Figure 1 A block diagram of another example implementation of a debug circuit for debugging Figure 1 The write operation of the accelerator circuit.

[0014] Figure 9 corresponds to Figure 1 、 2 , 3, 4, 5, 6, 7, 8A, 8B, 8C and / or 8D.

[0015] Figure 10 corresponds to Figure 1 A second example workflow of example operations for another example implementation of an accelerator circuit.

[0016] Figure 11 is representative of the execution by an example processor circuit to implement Figure 1 Flowcharts of example machine-readable instructions and / or example operations for an example computing system.

[0017] Figure 12 is representative of an example processor circuit executed to implement Figure 1 Another flow chart of example machine-readable instructions and / or example operations for an example computing system.

[0018] Figure 13 is representative of an example processor circuit executed to implement Figure 1 、 2 , 3, 4, 5, 6, 7, 8A, 8B, 8C, 8D and / or 10 are flowcharts of example machine-readable instructions and / or example operations of example debug circuits and / or example accelerator circuits.

[0019] Figure 14 is a block diagram of an example processing platform including a processor circuit configured to perform Figure 11-13 Example machine readable instructions and / or example operations to implement Figure 1 、 2 , 3, 4, 5, 6, 7, 8A, 8B, 8C, 8D and / or 10 example debug circuits and / or example accelerator circuits.

[0020] Figure 15 yes Figure 14 A block diagram of an example implementation of a processor circuit.

[0021] Figure 16 yes Figure 14 A block diagram of another example implementation of a processor circuit.

[0022] Figure 17 is a block diagram of an example software distribution platform that distributes software to client devices associated with end users and / or customers, retailers, and / or original equipment manufacturers (OEMs). DETAILED DESCRIPTION

[0023] The figures are not to scale. In general, throughout the accompanying drawings and the accompanying written description, the same reference numerals will be used to refer to the same or similar parts. As used herein, unless otherwise indicated, a connection reference (e.g., attached, coupled, connected and engaged) may include intermediate members between the elements referenced by the connection reference and / or relative movement between those elements. As such, a connection reference does not necessarily imply that two elements are directly connected and / or are in a fixed relationship to each other. As used herein, stating that any component is "in contact" with another component is defined as meaning that there is no intermediate component between the two components.

[0024] Unless otherwise specifically stated, descriptors such as "first," "second," "third," etc., as used herein, are not intended to imply or otherwise indicate any meaning of priority, physical order, arrangement in a list, and / or ordering in any manner, but are merely used as labels and / or arbitrary names to distinguish elements to facilitate understanding of the disclosed examples. In some examples, the descriptor "first" may be used to refer to an element in the detailed description, while the same element may be referred to using different descriptors, such as "second" or "third," in the claims. In such cases, it should be understood that such descriptors are only used to clearly identify, for example, those elements that might otherwise share the same name.

[0025] As used herein, the term "communication" (including variations thereof) encompasses direct communication and / or indirect communication through one or more intermediate components, and does not require direct physical (e.g., wired) communication and / or continuous communication, but additionally includes selective communication at periodic intervals, scheduled intervals, non-periodic intervals and / or one-time events.

[0026] As used herein, a "processor circuit" is defined as comprising (i) one or more specialized circuits configured to perform specific operations and including one or more semiconductor-based logic devices (e.g., electrical hardware implemented by one or more transistors), and / or (ii) one or more general-purpose semiconductor-based circuits programmed with instructions to perform specific operations and including one or more semiconductor-based logic devices (e.g., electrical hardware implemented by one or more transistors). Examples of processor circuits include programmed microprocessors, field programmable gate arrays (FPGAs) that can instantiate instructions, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), XPUs, or microcontrollers and integrated circuits such as application-specific integrated circuits (ASICs). For example, an XPU can be implemented by a heterogeneous computing system comprising multiple types of processor circuits (e.g., one or more FPGAs, one or more CPUs, one or more GPUs, one or more DSPs, etc., and / or combinations thereof) and (one or more) application programming interfaces (APIs) that can dispatch (one or more) computing tasks to the processing circuit(s) of the multiple types that are best suited to perform the computing tasks.

[0027] Typical computing systems (including personal computers and / or mobile devices) implement computationally intensive tasks such as advanced image processing or computer vision algorithms to automate tasks that human vision can perform. For example, computer vision tasks may include acquiring, processing, analyzing, and / or understanding digital images. Some such tasks facilitate, in part, extracting dimensional data from digital images to produce numerical and / or symbolic information. Computer vision algorithms can use numerical and / or symbolic information to make decisions and / or otherwise perform operations associated with, among other things, three-dimensional (3-D) pose estimation, event detection, object recognition, video tracking, and the like. In order to support augmented reality (AR), virtual reality (VR), robotics, and / or other applications, it is important to perform such tasks accordingly quickly (e.g., substantially in real time or near real time) and efficiently, where such tasks are performed by the example hardware accelerators disclosed herein.

[0028] Artificial intelligence / machine learning (AI / ML) models, such as neural networks (e.g., convolutional neural networks (CNNs or ConvNets)), can be utilized to implement computationally intensive tasks, such as advanced image processing or computer vision algorithms. Neural networks, such as CNNs, are deep artificial neural networks (ANNs) that are typically used to classify images, cluster images by similarity (e.g., photo search), and / or perform object recognition within images using convolutions. Thus, neural networks can be used to identify faces, individuals, street signs, animals, and the like included in an input image by passing the outputs of one or more filters corresponding to image features on the input image (e.g., horizontal lines, two-dimensional (2-D) shapes, etc.) to identify matches of the image features within the input image. Example hardware accelerators, as disclosed herein, can achieve such identification by processing a large number of inputs (e.g., AI / ML inputs) to generate outputs (e.g., AI / ML outputs), which can be used to achieve the identification.

[0029] Hardware accelerators that are customized, tailored, and / or otherwise optimized for implementing neural networks are referred to as neural network accelerators. Other types of AI / ML accelerators have the potential to improve the performance of specific types of AI / ML models. Such neural network accelerators, and / or hardware accelerators more generally, are becoming increasingly complex and difficult to debug in an attempt to improve and / or otherwise optimize the efficiency and performance that AI / ML models can achieve. As AI / ML datasets increase in size, debugging hardware accelerators is an increasingly time-consuming and complex task. Debugging is utilized in examples where the output of a hardware accelerator does not meet expectations, or where a particular configuration of the hardware accelerator and / or input (e.g., a configuration image) may cause a system hang or pipeline stall of the hardware accelerator.

[0030] Debugging can also be used to improve the performance of hardware accelerators. For example, improving the frames per second executed by a neural network accelerator may require a large number of compiler adjustments and modifications to identify pipeline or processing bottlenecks. The examples disclosed herein change the typical hardware debugging paradigm. For example, debugging hardware is typically designed for conventional microprocessor architectures that execute relatively long programs, where each debug instruction only works on a few small operands. However, with the advent of hardware accelerators such as graphics processor units (GPUs) and neural network accelerators, the ratio between debug instructions and operands has inverted. For example, hardware accelerators do not have dedicated hardware support for debugging purposes. In some such examples, a software application (e.g., a software debugger) used to debug a hardware accelerator may be designed to execute a relatively small program, but the operands (e.g., tensors in the CNN example) that each debug instruction operates on are quite large in number.

[0031] Without dedicated hardware debugging capabilities, the time required to debug a hardware accelerator can increase exponentially. For example, a single pass through a ResNet-50 neural network with an input image of size 224×224×3 (e.g., 150,000 inputs) generates over 10,500,000 activations, which traverse the network's 50 layers to produce a single output. However, newer neural network architectures can have even greater complexity and, thus, generate over 10,500,000 activations across more than 50 layers of the architecture. In some such examples, trying to find a bug in a set of 10,500,000 numbers spanning 50 layers is an increasingly difficult and time-consuming endeavor, especially if the network execution is decomposed into multiple smaller workloads. Further debugging difficulties arise in examples where a workload (e.g., a hardware accelerator workload, an AI / ML workload, etc.) is scheduled for execution by multiple cores (e.g., hardware accelerator cores) and thus runs or executes in parallel. In some of these examples, when multiple cores work in parallel, the potential for errors due to core interactions and workload synchronization is quite high.

[0032] As a result, identifying vulnerabilities, errors, etc. associated with the execution of AI / ML models may require a human to tediously infer the configuration of the hardware accelerator or other issues by inspecting the generated output. Advantageously, the examples disclosed herein include systems, apparatuses, methods, and articles of manufacture for debugging hardware accelerators by leveraging improved data-centric controllability through hardware accelerator operation to locate vulnerabilities and / or isolate performance bottlenecks.

[0033] Examples disclosed herein include systems, apparatus, methods, and articles of manufacture for debugging a hardware accelerator for improved performance and reduced generation of erroneous output. In some disclosed examples, the hardware accelerator includes example debug circuitry (or debugger circuitry) that can be instantiated to stop the output of the hardware accelerator at a specified breakpoint and step through one or more subsequent output transactions. In some disclosed examples, an example debug application (or debugger application) can program and / or instantiate the debug circuitry, and / or more generally, program and / or instantiate the hardware accelerator to stop execution of an AI / ML model on a per-workload, per-core basis in response to detecting specific generated data and / or in response to determining that an output transaction is associated with a certain address and / or address range. In some disclosed examples, the debug circuitry can output a readout of the output transaction (e.g., each output transaction, if instantiated as such) to identify data generated at a specified point in time during execution of the AI / ML model.

[0034] In some examples, if the address spaces in the hardware accelerator workload configurations overlap incorrectly, output data may be overwritten. In situations where many different output streams from a single workload and different workloads from different cores are running in parallel, the likelihood of accidental overwrites increases. In some such examples, a software debugger can be used to analyze the generated output and root cause issues, but such an effort is difficult and consumes a lot of time. Advantageously, the example debugging circuitry disclosed herein reduces the difficulty and time consumption of such an effort.

[0035] In some examples, due to incorrect configuration, accelerator outputs may be sent to a completely different address space outside the actual provisioned accelerator memory. If the address to which the output is sent is unknown, the software debugger may be deficient in locating the output. For example, a software debugger can analyze memory contents, but if the memory contents do not match expectations or have not yet been written, the software debugger may not be able to determine whether a memory transaction was issued or whether a memory transaction was issued to an incorrect address outside the observable address space. Advantageously, the example debug circuits disclosed herein overcome such deficiencies.

[0036] In some examples, machine learning models are modified through changes in compiler software to better understand and pinpoint the root cause of a problem. However, implementing custom modifications in software for debugging purposes is extremely time-consuming, especially if the problem only arises due to parallel core execution. Advantageously, the example debugging circuitry disclosed herein overcomes this drawback.

[0037] In some examples, isolating specific erroneous data generated during operation on a network with millions of output points for analysis can be a tedious task if there is no hardware support that can automatically detect specific data segments, halt execution, and signal the user to obtain further instructions. In some such examples, the hardware accelerator may not have the ability to detect writes to certain addresses or address ranges, resulting in deficiencies in isolating unintended writes. Advantageously, the example debug circuits disclosed herein overcome such deficiencies.

[0038] Figure 11 is a diagram of an example computing environment 100 that includes an example computing system 102, including an example central processing unit (CPU) 104, an example field programmable gate array (FPGA) 106, a first example accelerator circuit 108 (identified by accelerator circuit A), and a second example accelerator circuit 110 (identified by accelerator circuit B). In the illustrated example, the first accelerator circuit 108 and the second accelerator circuit 110 include an example debug circuit 112. In the illustrated example, the CPU 104 and the FPGA 106 include and / or otherwise instantiate an example debug application 114 (identified by debug application). In this example, the computing system 102 includes an example interface circuit 116, an example memory 118, an example power supply 120, and an example data storage 122.

[0039] In the illustrated example, the data store 122 includes example machine learning (ML) model(s) 124 and example breakpoint(s) 126. For example, the ML model(s) 124 may include one or more ML models, and one or more of the ML models may be of different types from one another. The breakpoint(s) 126 may include one or more breakpoints that, when triggered, activated, and / or otherwise invoked by the debug circuitry 112 and / or more generally by the first accelerator circuitry 108 and / or the second accelerator circuitry 110, may halt execution of an executable object, which may be implemented by one of the following: an executable binary object, executable code (e.g., executable machine-readable code), an executable file (e.g., an executable binary file), an executable program, executable instructions (e.g., executable machine-readable instructions), etc., corresponding to one of the ML model(s) 124. In some examples, the breakpoint(s) 126 may include a breakpoint at the start of a workload, a breakpoint on a specific data item being written or in the process of being written, a breakpoint on a specific address or address range to which it is written, a breakpoint on a specific data item being read from memory 118 into the accelerator circuit 108, 110, a breakpoint on a specific address or address range being read from memory 118, a breakpoint on the generation of a specific internal data item to the accelerator circuit 108, 110, and the like.

[0040] exist Figure 1In the illustrated example of FIG100, the CPU 104, the FPGA 106, the first accelerator circuit 108, the second accelerator circuit 110, the debug circuit 112, the debug application 114, the interface circuit 116, the memory 118, the power supply 120, and the data storage 122 communicate with one or more of each other via an example bus 128. For example, the bus 128 can be implemented using at least one of an Inter-Integrated Circuit (I2C) bus, a Serial Peripheral Interface (SPI) bus, a Peripheral Component Interconnect (PCI) bus, or a Peripheral Component Interconnect Express (PCIe) bus. Additionally or alternatively, the bus 128 can be implemented using any other type of computing or electrical bus. Further depicted in the computing environment 100 are an example user interface 130, an example network 132, and an example external computing system 134.

[0041] In some examples, the computing system 102 is a system on a chip (SoC), which represents one or more integrated circuits (ICs) (e.g., compact ICs) that combine components of a computer or other electronic system in a compact format. For example, the computing system 102 can be implemented using a combination of one or more types of processor circuits, hardware logic, and / or hardware peripherals and / or interfaces. Additionally or alternatively, the computing system 102 can include (one or more) input / output (I / O) ports and / or auxiliary storage devices. For example, the computing system 102 can include a CPU 104, an FPGA 106, a first accelerator circuit 108, a second accelerator circuit 110, a debug circuit 112, an interface circuit 116, a memory 118, a power supply 120, a data storage 122, a bus 128, (one or more) I / O ports, and / or auxiliary storage devices, all on the same substrate (e.g., a silicon substrate, a semiconductor-based substrate, etc.). In some examples, the computing system 102 includes digital, analog, mixed-signal, radio frequency (RF), or other signal processing functionality.

[0042] Figure 1 The example FPGA 106 is a field programmable logic device (FPLD). For example, once configured, the FPGA 106 can instantiate the debug application 114. Alternatively, one or more of the FPGA 106, the first accelerator circuit 108, and / or the second accelerator circuit 110 can be different types of hardware, such as a digital signal processor (DSP), an application-specific integrated circuit (ASIC), and / or a programmable logic device (PLD).

[0043] exist Figure 1In the illustrated example, the first accelerator circuit 108 is an artificial intelligence (AI) accelerator. For example, the first accelerator circuit 108 can implement a hardware accelerator configured to accelerate AI tasks or workloads, such as neural networks (e.g., convolutional neural networks (CNNs), deep neural networks (DNNs), artificial neural networks (ANNs), etc.), machine vision, machine learning, etc. In some examples, the first accelerator circuit 108 can implement a sparse accelerator (e.g., a sparse hardware accelerator). In some examples, the first accelerator circuit 108 can implement a vision processing unit (VPU) to perform machine or computer vision computing tasks, and / or train and / or execute neural networks. In some examples, the first accelerator circuit 108 can train and / or execute CNNs, DNNs, ANNs, recurrent neural networks (RNNs), etc., and / or combinations thereof.

[0044] exist Figure 1 In the illustrated example, the second accelerator circuit 110 is a graphics processor unit (GPU). For example, the second accelerator circuit 110 can be a GPU that generates computer graphics, performs general-purpose computations, executes vector workloads, and the like. In some examples, the second accelerator circuit 110 is another instance of the first accelerator circuit 108. For example, the second accelerator circuit 110 can be an AI accelerator. In some such examples, the computing system 102 (or portion(s) thereof, such as the CPU 104) can provide portion(s) of the AI / ML workload to be executed in parallel by the first accelerator circuit 108 and the second accelerator circuit 110.

[0045] exist Figure 1 In the illustrated example of FIG1 , interface circuitry 116 is hardware that can implement one or more interfaces (e.g., computing interfaces, network interfaces, etc.). For example, interface circuitry 116 can be hardware, software, and / or firmware that implements a communication device (e.g., a network interface card (NIC), a smart NIC, a gateway, a switch, etc.), such as a transmitter, a receiver, a transceiver, a modem, a residential gateway, a wireless access point, and / or a network interface to facilitate exchanging data with an external machine (e.g., any kind of computing device) via network 132. In some examples, interface circuitry 116 communicates via Bluetooth. The communication can be performed by a wireless connection, an Ethernet connection, a digital subscriber line (DSL) connection, a wireless fidelity (Wi-Fi) connection, a telephone line connection, a coaxial cable system, a satellite system, a field line wireless system, a cellular telephone system, an optical connection (e.g., a fiber optic connection), etc. For example, the interface circuit 116 can be implemented by any type of interface standard, such as Bluetooth. interface, Ethernet interface, Wi-Fi interface, universal serial bus (USB), near field communication (NFC) interface, PCI interface and / or PCIe interface.

[0046] The memory 118 of the illustrated example can be implemented by at least one volatile memory (e.g., synchronous dynamic random access memory (SDRAM), dynamic random access memory (DRAM), RAMBUS dynamic random access memory (RDRAM), etc.) and / or at least one non-volatile memory (e.g., flash memory).

[0047] Computing system 102 includes a power supply 120 for delivering power to the hardware of computing system 102. In some examples, power supply 120 may implement a power delivery network. For example, power supply 120 may implement an alternating current to direct current (AC / DC) power supply, a direct current to direct current (DC / DC) power supply, or the like. In some examples, power supply 120 may be coupled to an electrical grid infrastructure, such as an AC mains (e.g., a 110 volt (V) AC mains, a 220V AC mains, or the like). Additionally or alternatively, power supply 120 may be implemented by one or more batteries. For example, power supply 120 may be a limited energy device, such as a lithium-ion battery or any other rechargeable battery or power source. In some such examples, power supply 120 may be rechargeable using a power adapter or converter (e.g., an AC / DC power converter), a wall outlet (e.g., a 110 VAC wall outlet, a 220V AC wall outlet, or the like), a portable energy storage device (e.g., a portable power bank, a portable power cell, or the like), or the like.

[0048] Figure 1The computing system 102 of the illustrated example includes a data store 122 for recording data (e.g., ML model(s) 124, breakpoint(s) 126, etc.). The data store 122 of this example can be implemented by volatile memory and / or non-volatile memory (e.g., flash memory). The data store 122 can additionally or alternatively be implemented by one or more double data rate (DDR) memories, such as DDR, DDR2, DDR3, DDR4, mobile DDR (mDDR), etc. The data store 122 can additionally or alternatively be implemented by one or more mass storage devices, such as hard disk drive(s) (HDD), compact disk (CD) drive(s), digital versatile disk (DVD) drive(s), solid state disk (SSD) drive(s), etc. Although the data store 122 is illustrated as a single data store in the illustrated example, the data store 122 can be implemented by any number and / or type(s) of data stores. Furthermore, the data stored in the data store 122 may be in any data format, such as, for example, binary data, comma-delimited data, tab-delimited data, Structured Query Language (SQL) structures, executable objects (eg, executable binary objects, configuration images, etc.), and the like.

[0049] exist Figure 1 In the illustrated example of FIG, computing system 102 communicates with user interface 130. For example, user interface 130 may be implemented by a graphical user interface (GUI), an application user interface, etc., which may be presented to a user in circuitry with computing system 102 and / or on a display device otherwise in communication with computing system 102. In this example, user interface 130 may implement debugging application 114. For example, a user (e.g., a developer, an IT administrator, a customer, etc.) may utilize debugging application 114 to control computing system 102, configure, train, execute, and / or debug ML model(s) 124, generate and / or modify breakpoint(s) 126, etc. by interacting with user interface 130. Alternatively, computing system 102 may include and / or otherwise implement user interface 130.

[0050] exist Figure 1In the illustrated example of , network 132 is the Internet. However, network 132 of this example can be implemented using any suitable wired and / or wireless network(s), including, for example, one or more data buses, one or more local area networks (LANs), one or more wireless LANs, one or more cellular networks, one or more private networks, one or more public networks, one or more edge networks, etc. In some examples, network 132 enables computing system 102 to communicate with external computing system(s) in external computing system 134.

[0051] exist Figure 1 In the illustrated example of , the external computing system 134 includes and / or otherwise implements one or more computing devices on which the ML model(s) 124 are to be executed. In this example, the external computing system 134 includes an example desktop computer 136, an example mobile device (e.g., a smartphone, an internet-enabled smartphone, etc.) 138, an example laptop computer 140, an example tablet computer (e.g., a tablet computer, an internet-enabled tablet computer, etc.) 142, and an example server (e.g., an edge server, a rack-mounted server, a virtualized server, etc.) 144. In some examples, the example computing system 134 may be implemented using a server-side architecture. Figure 1 134. Additionally or alternatively, external computing systems 134 may include, correspond to, and / or otherwise represent any other type and / or number of computing devices. For example, one or more of external computing systems 134 may be virtualized computing systems.

[0052] In some examples, one or more of external computing systems 134 executes ML model(s) in ML model(s) 124 to process a computational workload (e.g., an AI / ML workload). For example, mobile device 138 may be implemented as a cell phone or mobile phone having processor circuitry (e.g., a CPU, GPU, VPU, AI or neural network specific processor, etc.) on a single SoC to process AI / ML workloads using ML model(s) in ML model(s) 124. In some examples, desktop computer 136, mobile device 138, laptop computer 140, tablet computer 142, and / or server 144 may be implemented as computing device(s) having processor circuitry (e.g., a CPU, GPU, VPU, AI or neural network specific processor, etc.) on one or more SoCs to process AI / ML workloads using ML model(s) in ML model(s) 124. In some examples, server 144 may implement one or more servers (e.g., physical servers, virtualized servers, etc., and / or combinations thereof) that may implement data facilities, cloud services (e.g., public or private cloud providers, cloud-based repositories, etc.), etc. to process (one or more) AI / ML workloads using (one or more) ML models in (one or more) ML models 124.

[0053] exist Figure 1 In the illustrated example of , the debug application 114 obtains (one or more) ML models 124 and compiles and / or otherwise generates output, such as an executable binary object, that can be executed on the first accelerator circuit 108 and / or the second accelerator circuit 110 to perform accelerator operations, such as AI / ML workloads. For example, the debug application 114 can implement a compiler (e.g., an accelerator compiler, an AI / ML compiler, a neural network compiler, etc.). In some such examples, the debug application 114 can compile a configuration image based on the (one or more) ML models 124 and / or the (one or more) breakpoints 126 for implementation on the (one or more) accelerator circuits 108, 110. For example, the configuration image can be implemented by an executable binary object that includes AI / ML configuration data (e.g., register configuration, activation data, activation sparse data, weight data, weight sparse data, hyperparameters, etc.), the AI / ML operations to be executed (e.g., convolutions, neural network layers, etc.).

[0054] exist Figure 1In the illustrated example of , the debug application 114 can instruct, direct, and / or otherwise invoke one or more of the accelerator circuits 108, 110 to execute one or more of the ML model(s) 124, and the debug application 114 can configure the debug circuit 112 to debug the execution(s) of the ML model(s) 124. AI, including machine learning (ML), deep learning (DL), and / or other artificial machine-driven logic, enables a machine (e.g., a computer, logic circuit, etc.) to process input data using a model to generate output based on patterns and / or associations previously learned by the model via a training process. For example, the machine learning model(s) 124 can be trained using data to recognize patterns and / or associations and follow such patterns and / or associations when processing the input data so that other input(s) result in output(s) consistent with the recognized patterns and / or associations.

[0055] There are many different types of machine learning models and / or machine learning architectures. In some examples, the debugging application 114 generates (one or more) machine learning models 124 as (one or more) neural network models. The debugging application 114 can instruct the interface circuit 116 to transmit (one or more) machine learning models 124 to (one or more) external computing systems in the external computing system 134. Using the neural network model enables the accelerator circuits 108, 110 to execute AI / ML workloads. In general, machine learning models / architectures suitable for use in the example methods disclosed herein include recurrent neural networks. However, other types of machine learning models, such as supervised learning ANN models, clustering models, classification models, etc., and / or combinations thereof, may be used in addition or alternatively. Example supervised learning ANN models may include two-layer (2-layer) radial basis neural networks (RBNs), learning vector quantization (LVQ) classification neural networks, etc. Example clustering models may include k-means clustering, hierarchical clustering, mean-shift clustering, density-based clustering, etc. Example classification models may include logistic regression, support vector machines or networks, naive Bayes, etc. In some examples, debugging application 114 may compile and / or otherwise generate the machine learning model(s) in machine learning model(s) 124 as lightweight machine learning models.

[0056] In general, implementing an ML / AI system involves two phases—a learning / training phase and an inference phase. In the learning / training phase, a training algorithm is used to train (one or more) machine learning models 124 to operate according to patterns and / or associations based on, for example, training data. In general, (one or more) machine learning models 124 include internal parameters (e.g., configuration data) that guide how input data is transformed into output data, such as through a series of nodes and connections within (one or more) machine learning models 124. Additionally, hyperparameters are used as part of the training process to control how learning is performed (e.g., learning rate, number of layers to be used in the machine learning model, etc.). Hyperparameters are defined as training parameters that are determined before the training process is started.

[0057] Different types of training can be performed based on the type and / or expected output of the ML / AI model. For example, the debugging application 114 can invoke supervised training to select parameters of (one or more) machine learning models 124 using inputs and corresponding expected (e.g., labeled) outputs (e.g., by iterating on combinations of selected parameters) to reduce model error. As used herein, "label" refers to the expected output of the machine learning model (e.g., a classification, an expected output value, etc.). Alternatively, the debugging application 114 can invoke unsupervised training (e.g., used in deep learning, a subset of machine learning, etc.), which involves inferring patterns from inputs to select parameters of (one or more) machine learning models 124 (e.g., without the benefit of expected (e.g., labeled) outputs).

[0058] In some examples, the debugging application 114 uses unsupervised clustering of operational observables to train the machine learning model(s) 124. However, the debugging application 114 can additionally or alternatively use any other training algorithm, such as stochastic gradient descent, simulated annealing, particle swarm optimization, evolutionary algorithms, genetic algorithms, nonlinear conjugate gradient, etc.

[0059] In some examples, debugging application 114 may train machine learning model(s) 124 until the error level no longer decreases. In some examples, debugging application 114 may train machine learning model(s) 124 locally on computing system 102 and / or remotely at an external computing system (e.g., one or more of external computing system(s) 134) communicatively coupled to computing system 102. In some examples, debugging application 114 trains machine learning model(s) 124 using hyperparameters that control how learning is performed (e.g., learning rate, number of layers to use in the machine learning model, etc.). In some examples, debugging application 114 may use hyperparameters that control model performance and training speed, such as learning rate and regularization parameter(s). Debugging application 114 may select such hyperparameters, for example, through trial and error, to achieve optimal model performance. In some examples, debugging application 114 utilizes Bayesian hyperparameter optimization to determine an optimal and / or otherwise improved or more efficient network architecture to avoid model overfitting and improve the overall applicability of machine learning model(s) 124. Alternatively, the debugging application 114 may use any other type of optimization. In some examples, the debugging application 114 may perform retraining. The debugging application 114 may perform such retraining in response to override(s) by a user of the computing system 102, receipt of new training data, in response to debugging of the accelerator circuits 108, 110, etc.

[0060] In some examples, the debugging application 114 facilitates the training of the machine learning model(s) 124 using training data. In some examples, the debugging application 114 utilizes training data derived from locally generated data. In some examples, the debugging application 114 utilizes training data derived from externally generated data. In some examples where supervised training is used, the debugging application 114 can label the training data. Labels are applied to the training data manually by a user or by an automated data pre-processing system. In some examples, the debugging application 114 can pre-process the training data using, for example, an interface (e.g., interface circuitry 116). In some examples, the debugging application 114 subdivides the training data into a first portion of data for training the machine learning model(s) 124 and a second portion of data for validating the machine learning model(s) 124.

[0061] Once training is complete, the debugging application 114 can deploy the machine learning model(s) 124 to serve as an executable construct that processes inputs and provides outputs based on the network of nodes and connections defined in the machine learning model(s) 124. The debugging application 114 can store the machine learning model(s) 124 in the data store 122. In some examples, the debugging application 114 can invoke the interface circuitry 116 to transfer the machine learning model(s) 124 to the external computing system(s) in the external computing system 134. In some such examples, in response to transferring the machine learning model(s) 124 to the external computing system(s) in the external computing system 134, the external computing system(s) in the external computing system 134 can execute the machine learning model(s) 124, thereby executing the AI / ML workload with at least one of improved efficiency or performance. Advantageously, in response to debugging the ML model(s) 124, the debugging application 114 can publish and / or otherwise promote the ML model(s) 124 that are more accurate than the previous implementation.

[0062] Once trained, the deployed machine learning model(s) in the machine learning model(s) 124 can operate in the inference phase to process data. In the inference phase, data to be analyzed (e.g., live data) is input to the machine learning model(s) 124, and the machine learning model(s) 124 execute to create output. This inference phase can be thought of as the AI ​​"thinking," which generates output based on what it has learned from training (e.g., applying the learned patterns and / or associations to live data by executing the machine learning model(s) 124). In some examples, the input data undergoes pre-processing before being used as input to the machine learning model(s) 124. Furthermore, in some examples, the output data can undergo post-processing after it is generated by the machine learning model(s) 124 to transform the output into a useful result (e.g., a display of data, detection and / or identification of an object, instructions to be executed by a machine, etc.).

[0063] In some examples, the output of the deployed machine learning model(s) in machine learning model(s) 124 can be captured and provided as feedback. By analyzing the feedback, the accuracy of the deployed machine learning model(s) in machine learning model(s) 124 can be determined. If the feedback indicates that the accuracy of the deployed model is below a threshold or other criterion, the feedback and updated training dataset, hyperparameters, etc. can be used to trigger training of an updated model to generate an updated, deployed model.

[0064] In some examples, the debug application 114 can configure the debug circuitry 112 to debug and / or troubleshoot unexpected accelerator performance or ML model execution. For example, the debug circuitry 112 can receive input(s) to be processed by the accelerator circuitry 108, 110 (e.g., ML input(s)). In some such examples, in response to the breakpoint(s) 126 not being triggered based on the input(s) (e.g., value(s) of the input(s), address(es) of the input(s), etc.), the accelerator circuitry 108, 110 can pass the input(s) to the core of the accelerator circuitry 108, 110, and the debug circuitry 112 can thereby operate in a bypass operation mode. In some examples, in response to the breakpoint(s) 126 being triggered based on the input(s), the debug circuitry 112 can perform debug operations, which can include reading accelerator transactions, reading triggered breakpoint(s), modifying breakpoint(s), modifying input(s), etc., and / or combinations thereof. Advantageously, the debug circuitry 112 may reduce debug time associated with the accelerator circuits 108 , 110 and / or the ML model(s) 124 by halting execution of the accelerator pipeline in response to a breakpoint being triggered based on the input(s) to the ML model(s) 124 .

[0065] In some examples, debug circuitry 112 may receive output(s) (e.g., ML output(s)) generated by accelerator circuitry 108, 110 in response to execution of ML model(s) 124. In some such examples, in response to breakpoint(s) 126 not being triggered based on the output(s) (e.g., value(s) of the output(s), address(es) of the output(s), etc.), accelerator circuitry 108, 110 may pass the output(s) to memory 118 and thereby may operate in a bypass mode of operation. In some examples, in response to breakpoint(s) 126 being triggered based on the output(s), debug circuitry 112 may perform debug operations, which may include reading accelerator transactions, reading triggered breakpoint(s), modifying breakpoint(s), modifying input(s), etc., and / or combinations thereof. Advantageously, the debug circuitry 112 may reduce debug time associated with the accelerator circuits 108 , 110 and / or the ML model(s) 124 by halting execution of the accelerator pipeline in response to a breakpoint being triggered based on the output(s) to the ML model(s) 124 .

[0066] Figure 2 is a block diagram of a first example accelerator circuit debugging system 200, the first example accelerator circuit debugging system 200 including Figure 1 Debugging application 114, Figure 1 Memory 118 and third example accelerator circuit 202. In some examples, Figure 2 The third accelerator circuit 202 may be Figure 1 Example implementations of the first accelerator circuit 108 and / or the second accelerator circuit 110 .

[0067] exist Figure 2 In the illustrated example of , the memory 118 includes (one or more) example machine learning inputs 204 and (one or more) example machine learning outputs 206. For example, the (one or more) machine learning inputs 204 may be Figure 1 the data processed by the ML model(s) 124, Figure 1 The ML model(s) 124 can be instantiated by the third accelerator circuitry 202 to generate machine learning output(s) 206. In some such examples, the machine learning input(s) 204 can be numerical data, categorical data, time series data, textual data, one or more portions of digital images and / or videos, sensor data, and / or any other type of data that can be processed and / or analyzed by a machine learning model (e.g., data associated with autonomous motion, robotic control, Internet of Things (IoT) data, and / or the like). In some examples, the machine learning output(s) 206 can be numerical data, categorical data, time series data, textual data, and / or any combination thereof. For example, the third accelerator circuitry 202 can output numerical data from a multiply-accumulate (MAC) circuit of the third accelerator circuitry 202.

[0068] The third accelerator circuit 202 includes example debug circuits 208, 210 and example cores (e.g., core circuits) 212, 214. For example, the third accelerator circuit 202 includes two or more instances of the debug circuits 208, 210 and two or more instances of the cores 212, 214. Alternatively, the third accelerator circuit 202 may include fewer instances of the debug circuits 208, 210 and / or the cores 212, 214. In some examples, the debug circuits 208, 210 may be Figure 1 An example implementation of the debug circuit 112 is provided.

[0069] The debug circuitry 208, 210 of the illustrated example includes example debug register(s) 216. In some examples, the debug register(s) 216 may include one or more registers, which may be implemented using vector register(s), single instruction multiple data (SIMD) register(s), general purpose register(s), flag register(s), segment register(s), machine-specific register(s), instruction pointer register(s), control register(s), debug register(s), memory management register(s), machine check register(s), and the like. The debug register(s) 216 may store data values ​​corresponding to configuration parameters, settings, and the like of the debug circuitry 208, 210. For example, the debug register(s) 216 may store value(s) representing a breakpoint to be triggered by the debug circuitry 208, 210 and / or the cores 212, 214. In some examples, debug register(s) 216 may store value(s) corresponding to machine learning input(s) in machine learning input(s) 204, address(es) and / or address range(s) associated with machine learning input(s) in machine learning input(s) 204, machine learning output(s) in machine learning output(s) 206, address(es) and / or address range(s) associated with machine learning output(s) in machine learning output(s) 206, the like, and / or combinations thereof.

[0070] The debug circuitry 208, 210 of the illustrated example includes an example debug interface 218. In some examples, the debug interface 218 can be implemented using an I2C bus, an SPI bus, a PCI bus, a PCIe bus, and / or any other type of electrical, hardware, or computing bus. In some examples, the debug application 114 can transmit data to and / or store or write data in debug register(s) 216 of the debug circuitry 208, 210 and / or more generally to the debug circuitry 208, 210 via the debug interface 218. In some examples, the debug application 114 can receive data from the debug register(s) 216 and / or more generally from the debug circuitry 208, 210 via the debug interface 218.

[0071] The cores 212, 214 of the illustrated example include example execution circuitry 220. In some examples, the execution circuitry 220 can be implemented using circuitry that can generate (one or more) machine learning outputs 206 based on (one or more) machine learning inputs 204. For example, the execution circuitry 220 can implement Figure 1 In some examples, the execution circuitry 220 may be implemented using MAC circuitry, datapath unit (DPU) circuitry, arithmetic logic circuitry (e.g., one or more arithmetic logic units (ALUs)), etc., and / or combinations thereof. For example, the debug application 114 may compile an executable binary object based on the machine learning model(s) 124 and provide the executable binary object to the cores 212, 214 via the core interface 224. In some examples, the execution circuitry 220 may be configured based on the executable binary object to implement the machine learning model(s) 124. In some examples, the debug application 114 may compile the executable binary object by including one or more breakpoints in the executable binary object. In some examples, the cores 212, 214 may store the value(s) of the one or more breakpoints in the configuration register(s) 222 to be called by the execution circuitry 220, and / or, more generally, store the one or more breakpoints.

[0072] The illustrated example cores 212, 214 include example configuration register(s) 222 (identified by configuration register(s). In some examples, the configuration register(s) 222 may include one or more registers that may be implemented using vector register(s), SIMD register(s), general purpose register(s), flag register(s), segment register(s), machine-specific register(s), instruction pointer register(s), control register(s), debug register(s), memory management register(s), machine check register(s), and the like. The configuration register(s) 222 may store data values ​​corresponding to configuration parameters, settings, and the like for the configuration circuitry 220 and / or more generally for the cores 212, 214. For example, the configuration register(s) 222 may store values ​​from an executable binary object to configure the execution circuitry 220 and / or, more generally, configure the cores 212, 214 to implement the machine learning model(s) 124. In some examples, configuration register(s) 222 may store value(s) representing breakpoints to be triggered by debug circuitry 208, 210 and / or cores 212, 214. In some examples, configuration register(s) 222 may store value(s) corresponding to machine learning input(s) of machine learning input(s) 204, address(es) and / or address range(s) associated with machine learning input(s) of machine learning input(s) 204, machine learning output(s) of machine learning output(s) 206, address(es) and / or address range(s) associated with machine learning output(s) of machine learning output(s) 206, the like, and / or combinations thereof.

[0073] The cores 212, 214 of the illustrated example include an example core interface 224. In some examples, the core interface 224 can be implemented using an I2C bus, an SPI bus, a PCI bus, a PCIe bus, and / or any other type of electrical, hardware, or computing bus. In some examples, the debug application 114 can transfer data to and / or store or write data in configuration register(s) 222 of the cores 212, 214 and / or more generally to the cores 212, 214 via the core interface 224. In some examples, the debug application 114 can receive data from the configuration register(s) 222 and / or more generally from the cores 212, 214 via the core interface 224.

[0074] exist Figure 2 In the illustrated example, the third accelerator circuit 202 is implemented with debug circuits 208, 210 that are instantiated to trigger and / or otherwise invoke breakpoint(s) based on the machine learning input(s) 204. In the illustrated example, the input(s) (e.g., input terminal(s), input connection(s), etc.) of the debug circuits 208, 210 are coupled to the output(s) of the memory 118. The output(s) (e.g., output terminal(s), output connection(s), etc.) of the debug circuits 208, 210 are coupled to the input(s) of the core(s) 212, 214. For example, the output(s) of the debug circuits 208, 210 are coupled to the example bus 226, the execution circuitry 220, the configuration register(s) 222, and / or the core interface 224 of the core(s) 212, 214. In some examples, bus 226 can be implemented using an I2C bus, an SPI bus, a PCI bus, a PCIe bus, and / or any other type of electrical, hardware, or computing bus. Output(s) of cores 212, 214 are coupled to input(s) of memory 118. In some examples, debug interface 218 and core interface 224 are instantiated to communicate with debug application 114.

[0075] In an example operation, the execution circuit 220 may execute the executable binary object to implement Figure 1 In response to execution of the executable binary object, the execution circuitry 220 may perform a read operation to obtain one of the machine learning input(s) 204. In an example operation, the execution circuitry 220 may generate a request to read one of the machine learning input(s) 204 at one or more addresses of the memory 118. The debug circuitry 208, 210 may obtain the request and the one or more addresses.

[0076] In some examples, the debug circuitry 208, 210 may provide a request to the memory 118 in response to the breakpoint not being triggered. For example, the debug circuitry 208, 210 may determine that the one or more addresses do not match an address associated with the breakpoint. The debug circuitry 208, 210 may receive the requested machine learning input(s) from the memory 118. In an example operation, the debug circuitry 208, 210 may provide the requested machine learning input(s) from the machine learning input(s) 204 to the cores 212, 214 in response to the breakpoint not being triggered. The execution circuitry 220 may generate the machine learning output(s) 206 based on the machine learning input(s) 204. The execution circuitry 220 may write the machine learning output(s) 206 to the memory 118.

[0077] In some examples, debug circuitry 208, 210 can trigger a breakpoint in response to a determination that: (i) one or more addresses match an address (or address range) associated with the breakpoint and / or (ii) the requested machine learning input(s) in (one or more) machine learning inputs 204 match a value(s) associated with the breakpoint. For example, debug circuitry 208, 210 can trigger a breakpoint in response to a second determination that the address (or address range) at which the value(s) of machine learning input(s) 204 are stored matches a value of the breakpoint. In some examples, debug circuitry 208, 210 can compare a first value(s) of machine learning input(s) 204 with a second value(s) of debug register(s) 216. In some examples, debug circuitry 208, 210 can trigger a breakpoint (e.g., a debug breakpoint) in response to the first value(s) matching the second value(s). For example, debug circuitry 208, 210 can trigger a breakpoint in response to a first determination that the value(s) of machine learning input(s) 204 match a value of the breakpoint. In example operation, the debug circuitry 208 , 210 may halt execution of the executable binary object by the core 212 , 214 in response to a breakpoint being triggered.

[0078] In some examples, the debug application 114 can perform a debug operation and / or cause a debug operation to be performed in response to one or more breakpoints being triggered. For example, the debug application 114 can query at least one of the cores 212, 214 or the debug circuits 208, 210 for the invoked breakpoint(s). In some examples, the debug application 114 can retrieve and / or otherwise access at least one of the following: the machine learning input(s) 204, the machine learning output(s) 206, or the associated memory address(es) of the machine learning input(s) 204 and / or the machine learning output(s) 206 (e.g., an address or address range at which the machine learning input(s) 204 are read from the memory 118, or an address or address range at which the machine learning output(s) 206 are to be written to the memory 118). In some examples, the debug circuitry 208, 210, and / or more generally the third accelerator circuitry 202, can output at least one of the machine learning input(s) 204, the machine learning output(s) 206, or the associated memory address(es) of the machine learning input(s) 204 and / or the machine learning output(s) 206 via the debug interface 218. In some examples, the debug application 114 can determine the completion progress of the workload(s) (e.g., the machine learning workload(s)) executed by the cores 212, 214 by querying the cores 212, 214 to obtain data indicating at which portion of the execution of the executable binary object one or more breakpoints were triggered.

[0079] In some examples, debug application 114 can perform and / or cause the execution of debug operations in response to one or more breakpoints being triggered, and the debug operations can include adjusting and / or modifying data values. For example, debug application 114 can change the value(s) of machine learning input(s) 204 stored in memory 118, debug circuitry 208, 210, and / or cores 212, 214.

[0080] In some examples, the debugging application 114 can perform a debugging operation and / or cause the execution of a debugging operation in response to one or more breakpoints being triggered, and the debugging operation can include an increment operation of the executable binary object. For example, the debugging application 114 can instruct the debugging circuitry 208, 210 and / or more generally the third accelerator circuit 202 to perform an increment operation of the executable object (e.g., an increment accelerator operation, a single-step operation of the accelerator circuitry 108, 110, etc.). In some such examples, the increment operation can include one or more read operations, one or more write operations, and / or one or more compute operations. For example, the debugging application 114 can instruct the debugging circuitry 208, 210 to obtain a first input of (one or more) machine learning inputs 204 and / or read the first input out to the debugging application 114 via the debugging interface 218. In some such examples, the debugging application 114 can instruct the debugging circuitry 208, 210 to determine whether the first input triggers one or more breakpoints. In some such examples, the debug application 114 can instruct the debug circuitry 208, 210 to provide a first input to the execution circuitry 220 of the cores 212, 214 to generate a first output of the machine learning output(s) 206. In some examples, the debug application 114 can instruct the cores 212, 214 to read the first output out to the debug application 114 via the core interface 224. Advantageously, the debug application 114 can debug the third accelerator circuitry 202 in an incremental manner, thereby identifying erroneous hardware accelerator operations with improved accuracy and granularity compared to previous implementations.

[0081] Advantageously, the debug circuitry 208, 210 may be implemented to accelerate software and compiler development for hardware accelerators such as the third accelerator circuitry 202. For example, as the complexity of machine learning models such as neural networks continues to increase, efforts to pinpoint any issues (e.g., bugs, performance bottlenecks, etc.) in the execution of those machine learning models in hardware accelerators are growing. Figure 1 With debug circuitry 112, the task of isolating any problems can be greatly simplified. Advantageously, debug circuitry 208, 210 can obviate the need for dedicated external debug equipment to perform debug operations on the hardware accelerator. For example, a compiler engineer can command full use of debug circuitry 208, 210 by establishing a communication channel (e.g., an SPI communication channel, an I2C communication channel, a communication channel utilizing one or more application programming interfaces (APIs), etc.) into the hardware accelerator.

[0082] In some examples, there is a software model that allows for pre-computation of expected hardware outputs for given inputs to machine learning model(s) 124. In some such examples, for each workload identified by an executable binary object, memory transactions can be obtained from debug circuitry 208, 210 and matched against the expected outputs from the software model. In some such examples, debug application 114 can identify issues based on detected mismatches.

[0083] Figure 3 is a block diagram of a second example accelerator circuit debugging system 300, the second example accelerator circuit debugging system 300 including Figure 1 Debugging application 114, Figure 1 Memory 118 and fourth example accelerator circuit 302. In some examples, Figure 3 The fourth accelerator circuit 302 may be Figure 1 Example implementations of the first accelerator circuit 108 and / or the second accelerator circuit 110 .

[0084] The fourth accelerator circuit 302 includes Figure 2 The debugging circuits 208 and 210 include Figure 2 The fourth accelerator circuit 302 includes the debug register(s) 216 and the debug interface 218. Figure 2 The cores 212 and 214 include Figure 2 Execution circuitry 220 , configuration register(s) 222 , core interface 224 , and bus 226 .

[0085] exist Figure 3In the illustrated example, the fourth accelerator circuit 302 is implemented with debug circuitry 208, 210 that is instantiated to trigger and / or otherwise invoke breakpoint(s) based on output(s) from the cores 212, 214 and / or more generally based on the machine learning output(s) 206. In the illustrated example, input(s) of the cores 212, 214 are coupled to output(s) of the memory 118. For example, input(s) of the execution circuitry 220, configuration register(s) 222, core interface 224, and / or bus 226 may be coupled to output(s) of the memory 118. Output(s) of the cores 212, 214 are coupled to input(s) of the debug circuitry 208, 210. For example, output(s) of the cores 212, 214 are coupled to debug register(s) 216 and / or debug interface 218 of the debug circuitry 208, 210. The output(s) of debug circuitry 208, 210 are coupled to the input(s) of memory 118. In some examples, debug interface 218 and core interface 224 are instantiated to communicate with debug application 114.

[0086] In example operations, bus 226 and / or more generally, cores 212, 214 may obtain machine learning input(s) 204 from memory 118. Execution circuitry 220 may generate machine learning output(s) 206 based on the machine learning input(s) 204. Execution circuitry 220 may provide, convey, and / or otherwise output the machine learning output(s) 206 to debug circuitry 208, 210.

[0087] In example operation, the debug circuitry 208, 210 may output and / or otherwise write the machine learning output(s) 206 to the memory 118 in response to determining that the machine learning output(s) 206 or data associated therewith (e.g., a memory address, a memory address range, etc.) do not trigger a breakpoint. In example operation, the debug circuitry 208, 210 may stop execution of an ongoing workload by the cores 212, 214 by not performing a read operation from the cores 212, 214 in response to determining that one or more breakpoints are triggered based on the machine learning output(s) 206 or data associated therewith. In example operation, the debug application 114 may perform one or more debug operations in response to a determination that one or more breakpoints are triggered.

[0088] Figure 4 is a block diagram of a third example accelerator circuit debugging system 400, the third example accelerator circuit debugging system 400 including Figure 1Debugging application 114, Figure 1 Memory 118 and fifth example accelerator circuit 402. In some examples, Figure 4 The fifth accelerator circuit 402 may be Figure 1 Example implementations of the first accelerator circuit 108 and / or the second accelerator circuit 110 .

[0089] The fifth accelerator circuit 402 includes Figure 2 The debugging circuits 208 and 210 include Figure 2 The fifth accelerator circuit 402 includes the debug register(s) 216 and the debug interface 218. Figure 2 The cores 212 and 214 include Figure 2 Execution circuitry 220 , configuration register(s) 222 , core interface 224 , and bus 226 .

[0090] The fifth accelerator circuit 402 of the illustrated example includes a circuit coupled to Figure 2 214. In some examples, the debug circuits 404, 406 may be Figure 110. The fifth accelerator circuit 402 is implemented with debug circuits 208, 210, 404, 406 that are instantiated to trigger and / or otherwise invoke breakpoint(s) based on at least one of: (i) input(s) to the cores 212, 214, and / or more generally, the machine learning input(s) 204, or (ii) output(s) from the cores 212, 214, and / or more generally, the machine learning output(s) 206. In the illustrated example, the input(s) of the debug circuits 208, 210 are coupled to the output(s) of the memory 118. For example, the debug register(s) 216 and / or the input(s) of the debug interface 218 are coupled to the output(s) of the memory 118. The output(s) of debug circuitry 208, 210 are coupled to the input(s) of cores 212, 214. For example, the output(s) of debug circuitry 208, 210 are coupled to the input(s) of execution circuitry 220, configuration register(s) 222, core interface 224, and / or bus 226. The output(s) of cores 212, 214 are coupled to the input(s) of debug circuitry 404, 406. For example, the output(s) of cores 212, 214 are coupled to the debug register(s) 216 and / or debug interface 218 of debug circuitry 404, 406. The output(s) of debug circuitry 404, 406 are coupled to the input(s) of memory 118. In some examples, debug interface 218 and core interface 224 are instantiated to communicate with debug application 114.

[0091] In example operation, the cores 212, 214 may execute executable binary objects to implement Figure 1One of the (one or more) ML models 124. In response to the execution of the executable binary object, the core 212, 214 may request (one or more) machine learning inputs from the memory 118. In example operation, the debug circuitry 208, 210 may obtain the request. In some examples, the debug circuitry 208, 210 may trigger a breakpoint in response to a determination that an address (or address range) associated with the read request matches an address (or address range) of a breakpoint. In some such examples, the debug circuitry 208, 210 may stop the execution of the executable binary object by preventing subsequent read and / or write operations from being completed (and thereby creating backpressure in the accelerator pipeline). In response to the triggering of the breakpoint based on the address (or address range), the debug application 114 may perform one or more debug operations, which may include obtaining the read request, triggering the address, address range, etc. of the breakpoint, performing an increment operation, etc., and / or a combination thereof.

[0092] In some examples, the debug circuitry 208, 210 may not trigger a breakpoint in response to a determination that an address (or address range) associated with the read request does not match an address (or address range) of a breakpoint. In response to a determination that a breakpoint is not triggered based on the address (or address range) of the read request, the debug circuitry 208, 210 may obtain (one or more) machine learning inputs 204 from the memory 118. In some examples, the debug circuitry 208, 210 may identify that one or more breakpoints are triggered based on the machine learning inputs 204. In response to the identification that one or more breakpoints in the one or more breakpoints are triggered based on the machine learning inputs 204, the debug application 114 may perform one or more debug operations, which may include obtaining (one or more) machine learning inputs 204 that trigger one or more breakpoints in the one or more breakpoints, performing an increment operation, etc., and / or combinations thereof.

[0093] In an example operation, the debug circuitry 208, 210 may output the machine learning input(s) 204 to the cores 212, 214 in response to a determination that the machine learning input(s) 204 did not trigger a breakpoint. For example, the debug circuitry 208, 210 may output the machine learning input(s) 204 to the cores 212, 214 to implement a breakpoint by generating the machine learning output(s) 206 based on the machine learning input(s) 204. Figure 1 The cores 212 , 214 may output the machine learning output(s) 206 to the debug circuits 404 , 406 .

[0094] In example operation, the debug circuitry 404, 406 may determine that one or more breakpoints are triggered based on the machine learning output(s) 206, which may include the value of the machine learning output(s) 206, the address(es) of the memory 118 to which the values ​​may be written, etc. In response to a determination that one or more of the one or more breakpoints are triggered based on the machine learning output(s) 206, the debug application 114 may perform one or more debug operations, which may include obtaining the machine learning output(s) 206 that triggered the one or more breakpoint(s), performing an increment operation, etc., and / or combinations thereof.

[0095] Figure 5 is a block diagram of a fourth example accelerator circuit debugging system 500, the fourth example accelerator circuit debugging system 500 including Figure 1 Debugging application 114, Figure 1 Memory 118 and sixth example accelerator circuit 502. In some examples, Figure 5 The sixth accelerator circuit 502 may be Figure 1 Example implementations of the first accelerator circuit 108 and / or the second accelerator circuit 110 .

[0096] exist Figure 5 In the illustrated example, the sixth accelerator circuit 502 includes example cores 504, 506, including at least a first example core 504 and a second example core 506. In the illustrated example, the cores 504, 506 include Figure 2 The debugging circuits 208 and 210 include Figure 2 The cores 504, 506 of the illustrated example include the debug register(s) 216 and the debug interface 218. Figure 2 Execution circuitry 220 , configuration register(s) 222 , core interface 224 , and bus 226 .

[0097] The sixth accelerator circuit 502 is implemented with debug circuitry 208, 210 that is instantiated to trigger and / or otherwise invoke breakpoint(s) based on input(s) to the cores 504, 506 and / or, more generally, the machine learning input(s) 204. In the illustrated example, the input(s) of the debug circuitry 208, 210 are coupled to output(s) of the memory 118. For example, the input(s) of the debug register(s) 216 and / or the debug interface 218 are coupled to output(s) of the memory 118. The output(s) of the debug circuitry 208, 210 are coupled to input(s) of the execution circuitry 220, configuration register(s) 222, core interface 224, and / or bus 226. The output(s) of the cores 504, 506 are coupled to input(s) of the memory 118. For example, the output(s) of execution circuitry 220 are coupled to the input(s) of memory 118. In some examples, debug interface 218 and core interface 224 are instantiated to communicate with debug application 114.

[0098] Figure 6 is a block diagram of a fifth example accelerator circuit debugging system 600, the fifth example accelerator circuit debugging system 600 including Figure 1 Debugging application 114, Figure 1 Memory 118 and seventh example accelerator circuit 602. In some examples, Figure 6 The seventh accelerator circuit 602 may be Figure 1 Example implementations of the first accelerator circuit 108 and / or the second accelerator circuit 110 .

[0099] exist Figure 6 In the illustrated example, the seventh accelerator circuit 602 includes example cores 604, 606, including at least a first example core 604 and a second example core 606. In the illustrated example, the cores 604, 606 include Figure 2 The debugging circuits 208 and 210 include Figure 2 The cores 604, 606 of the illustrated example include the debug register(s) 216 and the debug interface 218. Figure 2 Execution circuitry 220 , configuration register(s) 222 , core interface 224 , and bus 226 .

[0100] The seventh accelerator circuit 602 is implemented with debug circuits 208, 210 that are instantiated to trigger and / or otherwise invoke breakpoint(s) based on output(s) from the cores 604, 606 and / or, more generally, the machine learning output(s) 206. In the illustrated example, the input(s) of the cores 604, 606 are coupled to the output(s) of the memory 118. For example, the input(s) of the execution circuit 220 are coupled to the output(s) of the memory 118. The output(s) of the execution circuit 220, the configuration register(s) 222, and / or the core interface 224 are coupled (e.g., via the bus 226) to the input(s) of the debug circuits 208, 210. The output(s) of the debug circuits 208, 210 are coupled to the input(s) of the memory 118. For example, debug register(s) 216, debug interface 218, and / or more generally, output(s) of debug circuitry 208, 210 are coupled to input(s) of memory 118. In some examples, debug interface 218 and core interface 224 are instantiated to communicate with debug application 114.

[0101] Figure 7 is a block diagram of a sixth example accelerator circuit debugging system 700, the sixth example accelerator circuit debugging system 700 including Figure 1 Debugging application 114, Figure 1 Memory 118 and eighth example accelerator circuit 702. In some examples, Figure 7 The eighth accelerator circuit 702 may be Figure 1 Example implementations of the first accelerator circuit 108 and / or the second accelerator circuit 110 .

[0102] The eighth accelerator circuit 702 includes example cores 704, 706. The cores 704, 706 include Figure 2 The debugging circuits 208 and 210 include Figure 2 debug register(s) 216 and debug interface 218, and Figure 4 The debug circuits 404 and 406 of the core 704 and 706 include Figure 2The eighth accelerator circuitry 702 is implemented with debug circuitry 208, 210, 404, 406 that is instantiated to trigger and / or otherwise invoke breakpoint(s) based on at least one of: (i) input(s) to the core(s) 704, 706, and / or more generally, the machine learning input(s) 204, or (ii) output(s) from the core(s) 704, 706, and / or more generally, the machine learning output(s) 206.

[0103] exist Figure 7 In the illustrated example of FIG, input(s) of debug circuitry 208, 210 are coupled to output(s) of memory 118. For example, input(s) of debug register(s) 216 and / or debug interface 218 are coupled to output(s) of memory 118. Output(s) of debug circuitry 208, 210 are coupled (e.g., via bus 226) to input(s) of execution circuitry 220, configuration register(s) 222, and / or core interface 224. Output(s) of execution circuitry 220, configuration register(s) 222, core interface 224, and / or bus 226 are coupled to input(s) of debug circuitry 404, 406. For example, output(s) of execution circuitry 220, configuration register(s) 222, core interface 224, and / or bus 226 are coupled to debug register(s) 216 and / or debug interface 218 of debug circuitry 404, 406. The output(s) of debug circuitry 404, 406 are coupled to the input(s) of memory 118. In some examples, debug interface 218 and core interface 224 are instantiated to communicate with debug application 114.

[0104] In an example operation, the cores 704, 706 may execute executable binary objects to implement Figure 1 In response to execution of the executable binary object, the core 704, 706 may request one of the machine learning input(s) 204 from the memory 118. In example operation, the debug circuitry 208, 210 may obtain the request. In some examples, the debug circuitry 208, 210 may trigger a breakpoint in response to a determination that the request or data associated therewith (e.g., an address or address range at which the requested one of the machine learning input(s) 204 is stored in the memory 118) triggers a breakpoint.

[0105] In some examples, the debug circuitry 208, 210 may obtain the machine learning input(s) 204 from the memory 118 in response to a determination that the request did not trigger a breakpoint. The debug circuitry 208, 210 may identify that one or more breakpoints are triggered based on the machine learning input(s) 204. In response to the identification(s) that the breakpoint(s) of the one or more breakpoints are triggered based on the machine learning input(s) 204, the debug application 114 may perform one or more debug operations, which may include obtaining the machine learning input(s) 204 that triggered the breakpoint(s) of the one or more breakpoints, performing an increment operation, the like, and / or combinations thereof.

[0106] In example operation, the debug circuitry 208, 210 may output the machine learning input(s) 204 to the execution circuitry 220 to implement the process by generating the machine learning output(s) 206 based on the machine learning input(s) 204. Figure 1 The execution circuitry 220 may output the machine learning output(s) 206 to the debug circuitry 404, 406.

[0107] In example operation, the debug circuitry 404, 406 may determine that one or more breakpoints are triggered based on the machine learning output(s) 206. In response to the determination(s) that the breakpoint(s) of the one or more breakpoints are triggered based on the machine learning output(s) 206, the debug application 114 may perform one or more debug operations, which may include obtaining the machine learning output(s) 206 that triggered the breakpoint(s) of the one or more breakpoints, performing an increment operation, the like, and / or combinations thereof.

[0108] Figure 8A is a block diagram of a seventh example accelerator circuit debugging system 800, the seventh example accelerator circuit debugging system 800 including Figure 1 Debugging application 114, Figure 1 The memory 118 and the ninth example accelerator circuit 802. The ninth accelerator circuit 802 includes an example debug circuit 804. In some examples, Figure 8A The ninth accelerator circuit 802 may be Figure 1 In some examples, the debug circuit 804 may be an example implementation of the first accelerator circuit 108 and / or the second accelerator circuit 110. Figure 1The debug application 114, the memory 118 and / or the ninth accelerator circuit 802 may be instantiated by a processor circuit such as a central processing unit that executes instructions. Additionally or alternatively, Figure 8A The debug application 114, memory 118, ninth accelerator circuit 802, and / or debug circuit 804 may be instantiated by an ASIC or FPGA configured to perform operations corresponding to instructions.

[0109] The ninth accelerator circuit 802 of the illustrated example includes a first example execution circuit thread 806 (identified by execution circuit thread 0), a second example execution circuit thread 808 (identified by execution circuit thread N), and (one or more) example configuration registers 810. In some examples, the first execution circuit thread 806 and / or the second execution circuit thread 808 may be Figure 2-7 808. In some examples, the configuration register(s) 810 may be a multi-threaded execution circuit comprising N+1 threads (e.g., threads 0-N), which may include a first execution circuit thread 806 and a second execution circuit thread 808. In the illustrated example, the debug circuit 804 is instantiated at the beginning of the accelerator pipeline by intercepting and / or analyzing input(s) read from the memory 118, the input(s) to be routed to the first execution circuit thread 806 and / or the second execution circuit thread 808 via the debug circuit 804. Figure 2 In some examples, the first execution circuit thread 806, the second execution circuit thread 808, and the configuration register(s) 810 may be Figure 2-4 The first core 212 and / or the second core 214, Figure 5 The first core 504 and / or the second core 506, Figure 6 The first core 604 and / or the second core 606, and / or Figure 7 Example implementations of the first core 704 and / or the second core 706 .

[0110] The debug circuitry 804 of the illustrated example includes a first example interface circuit 812, a first example comparator circuit 814, a first example breakpoint register(s) 816, a second example interface circuit 818, a second example comparator circuit 820, a second example breakpoint register(s) 822, an example control circuit 824, an example multiplexer circuit 826, an example counter circuit 828, and an example shift register 830. In the illustrated example, the communicative coupling(s) between the first execution circuit thread 806 and the first interface circuit 812 may implement a first example communication channel 832. In the illustrated example, the communicative coupling(s) between the first interface circuit 812, the memory 118, and / or the multiplexer circuit 826 may implement a second example communication channel 834. In the illustrated example, the communicative coupling(s) between the second execution circuit thread 808 and the second interface circuit 818 may implement a third example communication channel 836. In the illustrated example, the communicative coupling(s) between the second interface circuit 818 , the memory 118 , and / or the multiplexer circuit 826 may implement a fourth example communication channel 838 .

[0111] In the illustrated example, the input(s) of the first execution circuit thread 806 and the second execution circuit thread 808 are coupled to the output(s) of the configuration register(s) 810. The input(s) and / or output(s) of the first execution circuit thread 806 are coupled to the corresponding output(s) and / or input(s) of the first interface circuit 812. The input(s) and / or output(s) of the first interface circuit 812 are coupled to the corresponding output(s) and / or input(s) of the memory 118, the first comparator circuit 814, the control circuit 824, and / or the multiplexer circuit 826. The input(s) and / or output(s) of the first comparator circuit 814 are coupled to the corresponding output(s) and / or input(s) of the control circuit 824 and / or the first breakpoint register(s) 816. The input(s) and / or output(s) of the first breakpoint register(s) 816 are coupled to the corresponding output(s) and / or input(s) of the control circuit 824. The input(s) and / or output(s) of configuration register(s) 810 are coupled to corresponding output(s) and / or input(s) of control circuitry 824 .

[0112] In the illustrated example, the input(s) and / or output(s) of the second execution circuit thread 808 are coupled to corresponding output(s) and / or input(s) of the second interface circuit 818. The input(s) and / or output(s) of the second interface circuit 818 are coupled to corresponding output(s) and / or input(s) of the memory 118, the second comparator circuit 820, the control circuit 824, and / or the multiplexer circuit 826. The input(s) and / or output(s) of the second comparator circuit 820 are coupled to corresponding output(s) and / or input(s) of the control circuit 824 and / or the second breakpoint register(s) 822. The input(s) and / or output(s) of the second breakpoint register(s) 822 are coupled to corresponding output(s) and / or input(s) of the control circuit 824. The output(s) of the counter circuit 828 are coupled to the input(s) (e.g., select input(s), control input(s), etc.) of the multiplexer circuit 826. The output(s) of the multiplexer circuit 826 are coupled to the input(s) of the shift register 830. The input(s) of the shift register 830 are coupled to the output(s) of the control circuit 824. The output(s) of the shift register 830 are coupled to the input(s) of the configuration register(s) 810.

[0113] In the illustrated example, configuration register(s) 810 are instantiated to communicate with the debug application 114. For example, the debug application 114 can read data from and / or write data to the configuration register(s) 810. In some such examples, the debug application 114 can write an executable binary object to the configuration register(s) 810, which causes the configuration register(s) 810 to configure the first execution circuit thread 806 and / or the second execution circuit thread 808. In some such examples, the executable binary object can include one or more breakpoints, which can be written to the configuration register(s) 810. In some examples, the debug application 114 can write commands, instructions, and the like, such as read instructions, single-step instructions, and resume instructions, to the configuration register(s) 810. In some examples, the debug application 114 can write a breakpoint to the configuration register(s) 810, which can cause the configuration register(s) 810 to provide the breakpoint to the control circuit 824 via the example breakpoint configuration instruction 848 (identified by BPCONFIG).

[0114] In an example operation, the debug application 114 may compile an executable binary object (eg, a configuration image) that, when executed and / or instantiated by the ninth accelerator circuit 802, may implement Figure 1 The debug application 114 may transmit, convey, and / or otherwise provide the executable binary object to the configuration register(s) 810. In response to receiving the executable binary object, the ninth accelerator circuit 802 may load values ​​based on the executable binary object into the configuration register(s) 810. In response to loading the values, the configuration register(s) 810 may configure the corresponding execution circuit thread(s) in the first execution circuit thread 806 and / or the second execution circuit thread 808 to implement the machine learning model(s) in the machine learning model(s) 124.

[0115] In an example operation, the first execution circuit thread 806 and / or the second execution circuit thread 808 may initiate execution of the accelerator workload based on the executable binary object according to the hardware arrangement, configuration, settings, etc. For example, the first execution circuit thread 806 and / or the second execution circuit thread 808 may obtain Figure 2 (one or more) machine learning inputs 204, and generating based on (one or more) machine learning inputs 204 Figure 2 (one or more) machine learning outputs 206. In some such examples, first execution circuit thread 806 can execute the executable binary object or (one or more) portions thereof, and / or second execution circuit thread 808 can execute the executable binary object or (one or more) portions thereof.

[0116] In an example operation, the first execution circuit thread 806 may request data related to executing an executable binary object (e.g., one or more machine learning inputs in the machine learning input(s) 204). The first execution circuit thread 806 may generate an example request signal 840 (identified by REQ / ADR) to the first interface circuit 812, which may include an address to be read from the memory 118. In response to the request signal 840 not triggering a breakpoint, the first interface circuit 812 may provide the request signal 840 to the memory 118 via the second communication channel 834 to facilitate the memory read operation. The memory 118 may generate a first example ready signal 842 (identified by RDY) to indicate to the first execution circuit thread 806 that the memory 118 is ready to provide the requested data. The first execution circuit thread 806 may generate a second example ready signal 844 (identified by RDY) to indicate to the memory 118 that the first execution circuit thread 806 is ready to receive the requested data. The memory 118 may provide the requested data via an example response signal 846 (identified by RSP / DATA).

[0117] In some examples, the debug application 114 can instantiate the debug circuit 804 to trigger a breakpoint (e.g., a breakpoint event, a debug event, etc.) based on the input(s), output(s), or associated memory address(es) of the ninth accelerator circuit 802. In some examples, the debug application 114 can compile the executable binary object to include one or more first breakpoints that, when called, stop execution of the executable binary object or portion(s) thereof. For example, the debug application 114 can compile the executable binary object to trigger the first breakpoint on a per-workload basis, which can be achieved when the first breakpoint corresponds to a specific or targeted workload. In some such examples, the debug circuit 114 can load the first breakpoint into one or more configuration registers in the configuration register(s) 810. In some such examples, the configuration register(s) 810 can provide the first breakpoint to the control circuit 824 via the BP CONFIG 848 instruction. In some such examples, control circuitry 824 may provide the first breakpoint to first breakpoint register(s) 816 and second breakpoint register(s) 822 .

[0118] In some examples, the first comparator circuit 814 can compare incoming data from the first interface circuit 812 and the first communication channel 832 with the first breakpoint from the first breakpoint register(s) 816. In some such examples, the first comparator circuit 814 can indicate to the control circuit 824 that the first breakpoint is triggered based on the comparison (e.g., the incoming data matches the data associated with the first breakpoint). In some such examples, the control circuit 824 can determine that the first execution circuit thread 806 has executed the target workload based on the first breakpoint triggered by the first comparator circuit 814.

[0119] In some examples, the second comparator circuit 820 can compare incoming data from the second interface circuit 818 and the third communication channel 836 with the first breakpoint from the second breakpoint register(s) 822. In some such examples, the second comparator circuit 820 can indicate to the control circuit 824 that the first breakpoint is triggered based on the comparison (e.g., the incoming data matches the data associated with the first breakpoint). In some such examples, the control circuit 824 can determine that the second execution circuit thread 808 is executing the target workload based on the first breakpoint triggered by the second comparator circuit 820.

[0120] In example operation, in response to the first execution circuit thread 806 triggering a breakpoint for each workload, the control circuit 824 generates an example breakpoint hit signal 850 (identified by BP HIT) to the configuration register(s) 810. In some examples, the BP HIT signal 850 may indicate the triggering of a workload-specific breakpoint. For example, the BP HIT signal 850 may implement an example BREAKPOINT_ON_START signal, which may be used to indicate that a breakpoint has been triggered on the first data item of the workload. In some examples, the BP HIT signal 850 may implement an example BREAKPOINT_ON_DATA+DATA signal, which may be used to indicate that a breakpoint has been triggered on a specific data item (+DATA) in a memory transaction. In some examples, the BPHIT signal 850 may implement an example BREAKPOINT_ON_ADR+ADR+MASK signal, which may be used to indicate that a breakpoint has been triggered on a specific address (+ADR) in a memory transaction. In some such examples, a mask (+MASK) may be used to indicate which bit(s) of the address are to be compared and / or analyzed. Advantageously, the BREAKPOINT_ON_ADR+ADR+MASK signal can be used to instantiate breakpoints for entire address ranges (as well as specific addresses).

[0121] In an example operation, in response to generating the BP HIT signal 850, the control circuitry 824 may instruct the first interface circuitry 812 to pull down and / or otherwise disable the request signal 840 and the first ready signal 842. In response to pulling down the request signal 840 and the first ready signal 842 (e.g., by changing the request signal 840 and the first ready signal 842 to logic low signals (e.g., signals representing a digital “0”)), the first interface circuitry 812 stops execution of a portion of the accelerator pipeline implemented by the first execution circuitry thread 806. For example, in response to disabling the first ready signal 842, the first execution circuitry thread 806 may be unable to retrieve data from the memory 118.

[0122] In some examples, the debug application 114 can compile the executable binary object to trigger a second breakpoint on a per-core basis, which can be implemented when the second breakpoint corresponds to a specific or targeted core. In some such examples, the debug application 114 can load the second breakpoint into (one or more) configuration registers in the (one or more) configuration registers 810, and when triggered by the first execution circuit thread 806, the second breakpoint stops execution of the executable binary object by the first execution circuit thread 806. In some such examples, when the first execution circuit thread 806 is stopped and / or otherwise in a paused or standby execution state, a different thread can continue execution of the executable binary object. In example operation, in response to the first execution circuit thread 806 triggering the breakpoint for each core, the first comparator circuit can notify the control circuit 824 that the second breakpoint has been hit. In example operation, the control circuit 824 can generate a BP HIT signal 850 in response to receiving an indication from the first comparator circuit 814.

[0123] In some examples, the BP HIT signal 850 can implement one or more breakpoint configuration signals, which are generated in response to (one or more) triggering of core-specific breakpoints. For example, the BP HIT signal 850 can implement an example BREAKPOINT_ON_START signal, which can be used to indicate that a breakpoint has been triggered on the first data item of the workload on a specific core. In some examples, the BP HIT signal 850 can implement an example BREAKPOINT_ON_DATA+DATA signal, which can be used to indicate that a breakpoint has been triggered on a specific data item (+DATA) in a memory transaction by a specific core. In some examples, the BP HIT signal 850 can implement an example BREAKPOINT_ON_ADR+ADR+MASK signal, which can be used to indicate that a breakpoint has been triggered on a specific address (+ADR) in a memory transaction by a specific core. In some such examples, a mask (+MASK) can be used to indicate which bit(s) of the address are to be compared and / or analyzed. Advantageously, the BREAKPOINT_ON_ADR+ADR+MASK signal can be used to instantiate breakpoints for entire address ranges (as well as specific addresses).

[0124] In an example operation, in response to the BP hit signal 850 being generated, the control circuit 824 may direct the first interface circuit 812 to pull down and / or otherwise disable the request signal 840 and the first ready signal 842. In response to pulling down the request signal 840 and the first ready signal 842, the first interface circuit 812 causes the execution of the executable binary object by the first execution circuit thread 806 and / or the second execution circuit thread 808 to be stopped.

[0125] In some examples, the control circuitry 824 can provide an indication of what type of breakpoint was triggered. For example, the control circuitry 824 can provide a BP HIT signal 850 to the configuration register(s) 810, which can provide an indication to the debug application 114 that a first workload executed by the first execution circuit thread 806 triggered a breakpoint when the first workload was started (e.g., a start breakpoint indication). In some examples, the control circuitry 824 can provide a BP HIT signal 850 to the configuration register(s) 810, which can provide an indication to the debug application 114 that a second workload executed by the second execution circuit thread 808 triggered a breakpoint when a data value read as input or a data value generated as output matched a breakpoint value (e.g., a data breakpoint indication, a data value match breakpoint indication, etc.). In some examples, the control circuit 824 can generate a BP HIT signal 850, which can provide an indication to the debug application 114 that a third workload executed by the second execution circuit thread 808 has triggered a breakpoint (e.g., an address breakpoint indication, a memory address breakpoint indication, etc.) when the first memory address from which the data value is read matches the breakpoint value.

[0126] In an example operation, the control circuitry 824 may store in the configuration register(s) 810 at least one of the breakpoint(s) triggered by the first execution circuit thread 806 or the indication(s) of the completion progress of the executable binary object by the first execution circuit thread 806. For example, the debug application 114 may query the configuration register(s) 810 for the indication(s). In an example operation, the control circuitry 824 may store at least one of the machine learning input, the machine learning output, or the associated address(es) that triggered the breakpoint by the first execution circuit thread 806. In an example operation, the debug application 114 may query the configuration register(s) 810 for at least one of the machine learning input, the machine learning output, or the associated address(es). In an example operation, the debug application 114 may modify the configuration register(s) 810 to implement the change to the executable binary object and resume execution of the executable binary object for debugging purposes. In an example operation, the debug application 114 may modify the machine learning input stored in the memory 118 and / or the first execution circuit thread 806 and resume execution of the executable binary object for debugging purposes.

[0127] In some examples, debug application 114 can load one or more first breakpoints and / or one or more second breakpoints in one or more breakpoint registers of first breakpoint register(s) 816 and / or second breakpoint register(s) 822. For example, debug application 114 can store a first value in configuration register(s) 810, the first value representing a machine learning input, a memory address or memory address range at which the machine learning input is stored in memory 118, etc. In some such examples, control circuitry 824 can obtain the first value from configuration register(s) 810 and provide the first value to first breakpoint register(s) 816.

[0128] In example operation, in response to triggering of a breakpoint by the first execution circuit thread 806, the first interface circuit 812 may provide data from the first execution circuit thread 806 (e.g., an address, an address range, a machine learning input, etc. associated with a memory read operation) to the multiplexer circuit 826. The counter circuit 828 may increment the output value of the counter circuit 828 to instruct the multiplexer circuit 826 to cycle through the inputs of the multiplexer circuit 826 and / or more generally, to be output from the multiplexer circuit 826 by the execution circuit threads 806, 808, and their respective components. For example, the counter circuit 828 can output a first counter value of 0 to instruct the multiplexer circuit 826 to output a read request from the first execution circuit thread 806, output a second counter value of 1 to output a read response from the first execution circuit thread 806, output a third counter value of 2*N to output a read request from the second execution circuit thread 808, output a fourth counter value of (2*N)+1 to output a read response from the second execution circuit thread 808, and so on. For example, the counter circuit 828 can cause the multiplexer circuit 826 to output data in a round-robin distribution or pattern. Alternatively, the counter circuit 828 can output values ​​in any other order, distribution, or pattern. In some examples, the counter circuit 828 can skip inputs to the multiplexer circuit 826 that do not have data to be output from the multiplexer circuit 826.

[0129] In an example operation, multiplexer circuit 826 can output data associated with the machine learning input and / or associated address(es) to debug application 114 via configuration register(s) 810 as an example transaction 852 (identified by a debug transaction). For example, transaction 852 can implement a debug transaction (e.g., a debug data transaction) that includes at least one of the first value of the machine learning input that triggered a breakpoint, the address at which the machine learning input is stored in memory 118, and the like. In some examples, debug transaction 852 is generated via debug application 114 and configuration register(s) 810 in response to an example read transaction 854 from control circuit 824. For example, debug application 114 can write an example read transaction command 856 (identified by a read) to configuration register(s) 810. Control circuit 824 can obtain read transaction command 856 from configuration register(s) 810. Control circuit 824 can issue read transaction 854 in response to obtaining read transaction command 856. For example, read transaction 854 may implement commands, directions, instructions, etc. generated by debug application 114 that, when received by shift register 830 , cause shift register 830 to generate and / or otherwise output debug transaction 852 .

[0130] In some examples, the shift register 830 can read data on a single-bit basis to conserve resources. For example, to advance the shift register 830 by one bit, the debug application 114 can pulse and / or otherwise generate a read transaction 854. For example, the debug application 114 can write an example single-step command 858 to the configuration register(s) 810. The control circuit 824 can obtain the single-step command 858 and generate a read transaction 854 in response to obtaining the single-step command 858. Alternatively, the shift register 830 can read data on any other bit basis (e.g., a two-bit basis, a four-bit basis, a sixteen-bit basis, etc.). In some examples, a bit in the shift register 830 (e.g., a valid bit) can indicate whether a valid one of the debug transactions 852 has been captured. In some such examples, to reduce the number of read clock cycles, the valid bit can be the first bit shifted out of the shift register 830. In some such examples, in response to debug application 114 determining that a valid debug transaction in debug transactions 852 has not been captured, debug application 114 can terminate the reading of shift register 830 and proceed with another debug operation. In some examples, debug application 114 can instruct shift register 830 to read debug transactions 852 of interest rather than every debug transaction 852.

[0131] In an example operation, the debug application 114 may instruct the debug circuitry 804 and / or, more generally, the ninth accelerator circuitry 802 to perform one or more single-step operations. For example, the first interface circuitry 812 may pull down the request signal 840 and the first ready signal 842 in response to the invocation of a breakpoint. In some such examples, the debug application 114 may instruct the first interface circuitry 812 via the single-step command 858 to release the pull-down of the request signal 840 and the first ready signal 842 within the first clock cycle (or more if instructed by the debug application 114) to allow potential output from the first execution circuit thread 806 to be transferred to the memory 118. After the first clock cycle ends, the request signal 840 and the first ready signal 842 are pulled down to stop the execution of the executable binary object by the first execution circuit thread 806. The output is provided to the multiplexer circuitry 826, which may be provided to the shift register 830. A debug transaction 852 may be generated accordingly. Advantageously, the debug application 114 can cause the debug circuitry 804 to execute in discrete, individual accelerator operations, thereby identifying erroneous configuration, computation, or memory read / write operations with improved granularity, visibility, and accuracy compared to previous implementations.

[0132] In an example operation, the debug application 114 may instruct the debug circuit 804 to resume operation of the executable binary object via the first execution circuit thread 806 and the second execution circuit thread 808 in response to the breakpoint(s) being triggered by generating an example resume command 860. For example, the debug application 114 may write the resume command 860 into the configuration register(s) 810. The control circuit 824 may, in response to obtaining the resume command 860, instruct the first interface circuit 812 to release the pull-down force on the request signal 840 and the first ready signal 842 to resume data transfer between the first execution circuit thread 806 and the debug circuit 804.

[0133] In some examples, the debug application 114 can instruct the debug circuit 804 to be enabled or disabled. For example, the debug application 114 can enable the debug circuit 804 and thereby cause the debug circuit 804 to determine whether any breakpoints have been triggered. In some examples, the debug application 114 can disable the debug circuit 804 and thereby cause the debug circuit 804 to enter a bypass mode in which the debug circuit 804 does not halt execution of the executable binary object by the first execution circuit thread 806 and / or the second execution circuit thread 808.

[0134] In some examples, the debug application 114 writes breakpoint(s) to configuration register(s) in the configuration register(s) 810 to stop execution of the workload(s) based on comparison(s) of the breakpoint(s) to at least one of the input(s) or associated address(es) of the ninth accelerator circuit 802. For example, the debug application 114 may write a first breakpoint to the configuration register(s) 810, which may be based on a first machine learning input in the machine learning input(s) 204. The control circuitry 824 may obtain the first breakpoint from the configuration register(s) 810 and write the first breakpoint to the first breakpoint register(s) in the first breakpoint register(s) 816. In some such examples, the first comparator circuitry 814 may compare the first machine learning input from the memory 118 with the first breakpoint. In some such examples, in response to a match based on the comparison, the first comparator circuitry 814 may generate an indication and transmit the indication to the control circuitry 824. The control circuit 824 may cause the first interface circuit 812 to pull down the request signal 840 and the first ready signal 842 to stop the flow of data from the memory 118 and thereby stop the execution of the executable binary object by the first execution circuit thread 806 .

[0135] although Figure 8A The diagram shows the implementation Figure 1 Example manner of first accelerator circuit 108, second accelerator circuit 110 and / or debug circuit 112, but Figure 8A One or more of the elements, processes, and / or devices illustrated in the drawings may be combined, divided, rearranged, omitted, eliminated, and / or implemented in any other manner. Furthermore, debug circuitry 804, first execution circuit thread 806, second execution circuit thread 808, configuration register(s) 810, first interface circuitry 812, first comparator circuitry 814, first breakpoint register(s) 816, second interface circuitry 818, second comparator circuitry 820, second breakpoint register(s) 822, control circuitry 824, multiplexer circuitry 826, counter circuitry 828, example shift register 830, communication channels 832, 834, 836, 838, and / or more generally, Figure 1The first accelerator circuit 108 , the second accelerator circuit 110 , and / or the debug circuit 112 may be implemented by hardware, software, firmware, and / or any combination of hardware, software, and / or firmware. Thus, for example, the debug circuitry 804, the first execution circuit thread 806, the second execution circuit thread 808, the configuration register(s) 810, the first interface circuitry 812, the first comparator circuitry 814, the first breakpoint register(s) 816, the second interface circuitry 818, the second comparator circuitry 820, the second breakpoint register(s) 822, the control circuitry 824, the multiplexer circuitry 826, the counter circuitry 828, the example shift register 830, the communication channels 832, 834, 836, 838, and / or more generally, any of the first accelerator circuitry 108, the second accelerator circuitry 110, and / or the debug circuitry 112 may be implemented by a processor circuit, analog circuit(s), digital circuit(s), logic circuit(s), programmable processor(s), programmable microcontroller(s), GPU(s), DSP(s), ASIC(s), PLD(s), and / or FPLD(s) such as an FPGA. When any of the apparatus or system claims of this patent are read to cover a purely software and / or firmware implementation, at least one of the debug circuitry 804, the first execution circuitry thread 806, the second execution circuitry thread 808, the configuration register(s) 810, the first interface circuitry 812, the first comparator circuitry 814, the first breakpoint register(s) 816, the second interface circuitry 818, the second comparator circuitry 820, the second breakpoint register(s) 822, the control circuitry 824, the multiplexer circuitry 826, the counter circuitry 828, the example shift register 830, and / or the communication channels 832, 834, 836, 838 is thereby expressly defined as comprising a non-transitory computer-readable storage device or storage disk, such as a memory, a digital versatile disk (DVD), a compact disk (CD), a Blu-ray disk, etc., comprising the software and / or firmware. Still further, Figure 1 The first accelerator circuit 108, the second accelerator circuit 110 and / or the debug circuit 112 are Figure 8A One or more elements, processes and / or devices may be included in addition to or instead of those illustrated, and / or may include more than one of any or all of the illustrated elements, processes and devices.

[0136] Figure 8B yes Figure 8A 804 for debugging Figure 8AIn the illustrated example, debug circuitry 804 is instantiated at the end of the accelerator pipeline by intercepting and / or analyzing the output(s) from the first execution circuit thread 806 and / or the second execution circuit thread 808. The illustrated example includes Figure 8A 80, a debug circuit 804, a first execution circuit thread 806, a second execution circuit thread 808, configuration register(s) 810, a first interface circuit 812, a first comparator circuit 814, first breakpoint register(s) 816, a second interface circuit 818, a second comparator circuit 820, second breakpoint register(s) 822, a control circuit 824, a multiplexer circuit 826, a counter circuit 828, an example shift register 830, a BP CONFIG signal 848, a BP HIT signal 850, a debug transaction 852, a read transaction 854, a read transaction command 856, a single step command 858, and a restore command 860.

[0137] exist Figure 8B In the illustrated example, the communicative coupling(s) between the first execution circuit thread 806 and the first interface circuit 812 may implement a fifth example communication channel 862. In the illustrated example, the communicative coupling(s) between the first interface circuit 812, the memory 118, and / or the multiplexer circuit 826 may implement a sixth example communication channel 864. In the illustrated example, the communicative coupling(s) between the second execution circuit thread 808 and the second interface circuit 818 may implement a seventh example communication channel 866. In the illustrated example, the communicative coupling(s) between the second interface circuit 818, the memory 118, and / or the multiplexer circuit 826 may implement an eighth example communication channel 868.

[0138] In example operation, the communication channels 862, 864, 866, 868 facilitate debug write operations by the debug circuitry 804. For example, in response to execution of an executable binary object, the first execution circuitry thread 806 may execute a command based on Figure 2 The first machine learning input(s) in the machine learning input(s) 204 are used to generate Figure 2In response to generating the first machine learning output(s), the first execution circuit thread 806 generates an example request signal 870 (identified by REQ) to write the first machine learning output(s) to the memory 118. The first execution circuit thread 806 may generate an example address / data signal 872 (identified by ADR / DATA), which may include an address at which the first machine learning output(s) are to be written to the memory 118, the first machine learning output(s), the first machine learning output(s), the like, and / or a combination thereof. When the memory 118 is ready to receive the data to be written, the memory 118 may generate an example ready signal 874 (identified by RDY).

[0139] In an example operation, in response to the first execution circuit thread 806 not triggering a breakpoint, the first interface circuit 812 may provide the machine learning output from the first execution circuit thread 806 to the memory 118. In an example operation, the first interface circuit 812 may receive a first value representing one of the first machine learning output(s) generated by the first execution circuit thread 806 in response to execution of the executable binary object. In an example operation, the first comparator circuit 814 may compare the first value with a second value based on the breakpoint(s) in the first breakpoint register(s) 816. In response to a match, the first comparator circuit 814 may signal a match to the control circuit 824, which may instruct the first interface circuit 812 and the second interface circuit 818 to halt execution of the executable binary object by pulling down at least one of the request signal 870 or the ready signal 874. For example, the first comparator circuit 814 may pause execution of the executable binary object in response to the machine learning output from the first execution circuit thread 806, the associated address(es), etc., matching the machine learning output, the associated address(es), etc. of the breakpoint.

[0140] In an example operation, the multiplexer circuit 826 may output data associated with the machine learning output, associated address(es), etc. as a debug transaction 852. For example, the debug transaction 852 may include at least one of the machine learning output that triggered the breakpoint or the address at which the machine learning output is to be written to the memory 118.

[0141] Figure 8C yes Figure 8A and / or 8B is a block diagram of another example implementation of a debug circuit 804 for debugging Figure 8Aand / or read operations of the accelerator circuit 802 of 8B. The accelerator circuit 802 of the illustrated example includes a first execution circuit thread 806, (one or more) configuration registers 810 and a debug circuit 804, the debug circuit 804 including Figure 8A and / or the control circuit 824, multiplexer circuit 826, counter circuit 828, and shift register 830 of 8B. Figure 8C Further described in Figure 1 The debugging application 114 and the memory 118 are configured to be used for debugging. Figure 8C Also depicted in Figure 8A and 8B BP CONFIG signal 848, BPHIT signal 850, debug transaction 852, read transaction 854, read transaction command 856, single step command 858 and restore command 860.

[0142] The debug circuitry 804 of the illustrated example is included in the example execution circuitry 875. In some examples, the execution circuitry 875 may implement Figure 2-5 the first execution circuit 220 of the first core 212, Figure 6 The execution circuit 220 of the first core 604, etc. For example, the debug circuit 804 can be adapted to intercept signals inside the execution circuit 875. In the illustrated example, the execution circuit 875 includes and / or is otherwise implemented Figure 8A and 8B The first execution circuit thread 806 and the second execution circuit thread 808 are executed.

[0143] In the illustrated example, first example signals 876, 877, 878, 879 correspond to a first thread of execution circuitry 875, such as the first execution circuitry thread 806. For example, the first signals 876, 877, 878, 879 include a first example request signal 876 (identified by REQ_0), a first example address signal 877 (identified by ADR_0), a first example response signal 878 (identified by RSP_0), and a first example data signal 879 (identified by DATA_0) corresponding to the first execution circuitry thread 806.

[0144] In the illustrated example, the second example signals 880, 881, 882, 883 correspond to a second thread of the execution circuit 875, such as the second execution circuit thread 808. For example, the second signals 880, 881, 882, 883 include a second example request signal 880 (identified by REQ_N), a second example address signal 881 (identified by ADR_N), a second example response signal 882 (identified by RSP_N), and a second example data signal 883 (identified by DATA_N) corresponding to the second execution circuit thread 808.

[0145] In example operation, in response to a determination made by first execution circuit thread 806 to read data from memory 118, first execution circuit thread 806 generates a first request signal 876 to retrieve data from another portion of first execution circuit thread 806. First execution circuit thread 806 generates a first address signal 877 that indicates an address of data stored within first execution circuit thread 806 from which to read the data. Execution circuit 875 generates a first response signal 878 that indicates that the data is ready to be read from another portion of first execution circuit thread 806. Execution circuit 875 generates a first data signal 879 that includes the requested data.

[0146] In an exemplary operation, data associated with at least one of a first request signal 876, a first address signal 877, a first response signal 878, or a first data signal 879 is provided to the multiplexer circuit 826. The counter circuit 828 may select an input of the multiplexer circuit 826 corresponding to at least one of the first request signal 876, the first address signal 877, or the first data signal 879. The multiplexer circuit 826 outputs the selected data to the shift register 830. The shift register 830 outputs the selected data to the configuration register(s) 810 as a debug transaction 852. The debug application 114 may obtain the selected data from the configuration register(s) 810.

[0147] In example operation, in response to triggering a breakpoint based on at least one of an address, an address range, or a value of data retrieved from another portion of the execution circuitry 875, the control circuitry 824 may halt execution of the executable binary object by the execution circuitry 875 by generating an example stop signal 884. The stop signal 884 may be a signal that is a combination of the first request signal 876, the first response signal 878, and an accompanying ready signal (e.g., Figure 8A For example, in response to a determination that a breakpoint has been triggered based on the requested memory address, the control circuitry 824 may generate a corresponding one or more of the stop signals 884 to pull down at least one of the first request signal 876, the first response signal 878, the second request signal 880, the second response signal 882, and the accompanying ready signal, thereby stopping the flow of information through the execution circuitry 875.

[0148] In an example operation, the control circuitry 824 may single-step the executable binary object in response to the single-step command 858. For example, the control circuitry 824 may instruct the corresponding one or more stop signals 884 to release the pull-down force on at least one of the first request signal 876, the first response signal 878, the second request signal 880, the second response signal 882, and the accompanying ready signal within a single clock cycle, two or more clock cycles, etc. In an example operation, the control circuitry 824 may release the pull-down force on at least one of the first request signal 876, the first response signal 878, the second request signal 880, the second response signal 882, and the accompanying ready signal by generating the corresponding one or more stop signals 884, thereby not halting the execution of the executable binary object.

[0149] Figure 8D yes Figure 8A and / or 8B is a block diagram of another example implementation of a debug circuit 804 for debugging Figure 8A and / or write operations of the accelerator circuit 802 of 8B. The accelerator circuit 802 of the illustrated example includes a first execution circuit thread 806, (one or more) configuration registers 810 and a debug circuit 804, the debug circuit 804 including Figure 8A and / or the control circuit 824, multiplexer circuit 826, counter circuit 828, and shift register 830 of 8B. Figure 8D Further described in Figure 1 The debugging application 114 and the memory 118 are configured to be used for debugging. Figure 8D Also depicted in Figure 8A and 8B BP CONFIG signal 848, BPHIT signal 850, debug transaction 852, read transaction 854, read transaction command 856, single step command 858 and restore command 860. Figure 8D Further described in Figure 8C Execution circuit 875 and stop signal 884.

[0150] The debug circuit 804 of the illustrated example is included in Figure 8C In the execution circuit 875 of the embodiment of the present invention, first example signals 886, 888, 890 correspond to a first thread of the execution circuit 875, such as the first execution circuit thread 806. For example, the first signals 886, 888, 890 include a first example request signal 886 (identified by REQ_0), a first example address / data signal 888 (identified by ADR / DATA_0), and a first example ready signal 890 (identified by RDY_0) corresponding to the first execution circuit thread 806.

[0151] In the illustrated example, the second example signals 892, 894, 896 correspond to a second thread of the execution circuit 875, such as the second execution circuit thread 808. For example, the second signals 892, 894, 896 include a second example request signal 892 (identified by REQ_N), a second example address / data signal 894 (identified by ADR / DATA_N), and a second example ready signal 896 (identified by RDY_N) corresponding to the second execution circuit thread 808.

[0152] In example operation, in response to a determination made by first execution circuit thread 806 to write data to memory 118, first execution circuit thread 806 generates a first request signal 886 to write the data to memory 118. First execution circuit thread 806 generates a first address / data signal 888 indicating the address(es) and / or data to be written to memory 118. Memory 118 generates a first ready signal 890 indicating that the data is ready to be written to memory 118.

[0153] In an exemplary operation, data associated with at least one of a first request signal 886 or a first address / data signal 888 is provided to a multiplexer circuit 826. A counter circuit 828 may select an input of the multiplexer circuit 826 corresponding to at least one of the first request signal 886 or the first address / data signal 888. The multiplexer circuit 826 outputs the selected data to a shift register 830. The shift register 830 outputs the selected data to the configuration register(s) 810 as a debug transaction 852. The debug application 114 may obtain the selected data from the configuration register(s) 810.

[0154] In an example operation, in response to triggering a breakpoint based on at least one of an address, an address range, or a value of data to be written to the memory 118, the control circuitry 824 may halt execution of the executable binary object by the execution circuitry 875 by generating a halt signal 884. The halt signal 884 may pull down the first request signal 886, the first ready signal 890, the second request signal 892, and the second ready signal 896 from a logic high signal to a logic low signal. For example, in response to determining that a breakpoint has been triggered based on the requested memory address, the control circuitry 824 may generate a corresponding halt signal(s) in the halt signal 884 to pull down the first request signal 886, the first ready signal 890, the second request signal 892, and the second ready signal 896.

[0155] In an example operation, the control circuit 824 may single-step the executable binary object in response to the single-step command 858. For example, the control circuit 824 may instruct the corresponding (one or more) stop signals 884 to release the pull-down force on at least one of the first request signal 886, the first ready signal 890, the second request signal 892, and the second ready signal 896 within a single clock cycle, two or more clock cycles, etc. In an example operation, the control circuit 824 may release the pull-down force on at least one of the first request signal 886, the first ready signal 890, the second request signal 892, and the second ready signal 896 by instructing the corresponding (one or more) stop signals 884 to release the pull-down force on at least one of the first request signal 886, the first ready signal 890, the second request signal 892, and the second ready signal 896, thereby not stopping the execution of the executable binary object.

[0156] Figure 9 corresponds to Figure 1 The first accelerator circuit 108, Figure 1 The second accelerator circuit 110, Figure 2 The third accelerator circuit 202, Figure 3 The fourth accelerator circuit 302, Figure 4 The fifth accelerator circuit 402, Figure 5 The sixth accelerator circuit 502, Figure 6 The seventh accelerator circuit 602, Figure 7 The eighth accelerator circuit 702 and / or Figures 8A-8D A first example workflow 900 of example operations of the ninth accelerator circuit 802 is shown.

[0157] The first workflow 900 of the illustrated example can implement a sequence of example workloads 902, 904, 906, 908, 910 to generate an example output tensor 912 based on an example input tensor 914. For example, the workloads 902, 904, 906, 908, 910 can be based on Figure 1 The neural network computing workload is implemented using one or more machine learning models in the machine learning model(s) 124. The workloads 902, 904, 906, 908, 910 include a first example workload 902 (identified by workload 0), a second example workload 904 (identified by workload 1), a third example workload 906 (identified by workload 2), a fourth example workload 908 (identified by workload 3), and a fifth example workload 910 (identified by workload 4).

[0158] In the illustrated example, workloads 902, 904, 906, 908, 910 may be implemented by two cores of a hardware accelerator, such as Figure 2 The first core 212 and Figure 2In an example operation, the first core 212 may execute the first workload 902 on the input tensor 914, followed by the third workload 906. In an example operation, the second core 214 may execute the second workload 904 on the input tensor 914 in parallel with the first workload 902. In response to completing the second workload 904, the second core 214 may execute the fourth workload 908. In response to completing the third workload 906 and the fourth workload 908, the first core 212 and / or the second core 214 may execute the fifth workload 910 to generate the output tensor 912.

[0159] The illustrated example first workflow 900 can implement example accelerator circuit operations that include example breakpoints 916, 918 generated on a per-workload basis. For example, the illustrated example breakpoints 916, 918 include a first example breakpoint 916 corresponding to the execution of the second workload 904 and a second example breakpoint 918 corresponding to the execution of the third workload 906. In some such examples, the breakpoints 916, 918 are specific to the second workload 904 and the third workload 906 and can thus be activated on a core (e.g., the first core 212 or the second core 214) executing the respective second workload 904 and the third workload 906 for the duration of the second workload 904 and the third workload 906.

[0160] In some examples, second core 214 can trigger first breakpoint 916 in response to starting second workload 904. In response to satisfying the condition(s) associated with first breakpoint 916, second debug circuitry 210 can stop data flow at the input or output of second core 214 to stop execution of second workload 904. In this example, first core 212 is unaffected and can continue to execute first workload 902. In this example, first core 212 can complete first workload 902 and third workload 906 while second workload 904 is stopped by second debug circuitry 210. Advantageously, the initial state(s) of second core 214 can be read in response to a query by debug application 114 before executing second workload 904 to identify erroneous configurations, memory read / write operations, and the like. For example, the initial state(s) of second core 214 can include values ​​stored in configuration register(s) 222, values ​​of machine learning inputs 204 stored in execution circuitry 220, and the like.

[0161] In some examples, the first core 212 can trigger the second breakpoint 918 in response to an occurrence of a write operation of a machine learning output generated by the first core 212 matching the value 0x42. Advantageously, the state(s) of the first core 212 can be read out in response to a query by the debug application 114 to identify erroneous configurations, computations, memory read / write operations, etc. in response to execution of the third workload 906. For example, the state(s) of the first core 214 can include values ​​stored in the configuration register(s) 222, the value(s) of the machine learning input(s) 204 stored in the execution circuitry 220, the value(s) of the machine learning input(s) 204 stored in the memory 118, the value(s) of the machine learning output(s) 206 stored in the execution circuitry 220, etc.

[0162] Figure 10 is a second example workflow 1000 corresponding to example operations of the eleventh example accelerator circuit 1002. In some examples, the eleventh accelerator circuit 1002 can be composed of Figure 1 The first accelerator circuit 108, Figure 1 The second accelerator circuit 110, Figure 2 The third accelerator circuit 202, Figure 3 The fourth accelerator circuit 302, Figure 4 The fifth accelerator circuit 402, Figure 5 The sixth accelerator circuit 502, Figure 6 The seventh accelerator circuit 602, Figure 7 The eighth accelerator circuit 702 and / or Figures 8A-8D The ninth accelerator circuit 802 is implemented.

[0163] In the illustrated example, the eleventh accelerator circuit 1002 includes a first example core 1004 (identified by core 0), a second example core 1006 (identified by core 1), a third example core 1008 (identified by core 2), and a fourth example core 1010 (identified by core 3). For example, the first core 1004 may be Figure 2 The first core 212 is implemented by, and / or the second core 1006 can be implemented by Figure 2 In the second workflow 1000, the eleventh accelerator circuit 1002 reads data from a memory (e.g., Figure 1 The eleventh accelerator circuit 1002 may use a machine learning model (e.g., Figure 1The second workflow 1000 is executed in parallel in a loop over many iterations on one or more of the cores 1004, 1006, 1008, 1010 until the execution is complete and the desired output of the machine learning model is achieved. In this example, each of the cores 1004, 1006, 1008, 1010 includes an example debug circuit 1012, which can be provided by Figure 1 Debug circuit 112, Figure 2-7 Debug circuit 208, 210, Figure 4 and 7 Debug circuits 404, 406 and / or Figures 8A-8D This is achieved by the debugging circuit 804.

[0164] In some examples, if one or more of the cores 1004, 1006, 1008, 1010 and / or more generally the eleventh accelerator circuitry 1002 are misconfigured, execution of the machine learning model may never complete, or if it completes, the output may not be as expected. In some such examples, the debug circuitry 1012 may be invoked to perform debug operations to understand the cause of the unexpected output and, subsequently, perform corrective actions on the one or more of the cores 1004, 1006, 1008, 1010 and / or more generally the eleventh accelerator circuitry 1002 to correct the unexpected output.

[0165] In the illustrated example, debug circuitry 1012 can intercept and analyze transactions obtained from and / or transferred to memory to determine whether to trigger example breakpoint 1014. Breakpoint 1014 of the illustrated example is triggered in response to a write operation of data value 0x11 to memory by (one or more of) cores 1004, 1006, 1008, 1010. Alternatively, breakpoint 1014 can be triggered in response to a read operation of data value 0x11 from memory by (one or more of) cores 1004, 1006, 1008, 1010.

[0166] The second workflow 1000 can implement an example in which all cores 1004, 1006, 1008, and 1010 have been configured with the same data-driven breakpoint. In response to one or more of the cores 1004, 1006, 1008, and 1010 executing a workload to write the data value 0x11 to memory, a breakpoint 1014 is triggered. When the breakpoint 1014 is hit by the debug circuit 1012, the debug circuit 1012 stops the pipeline of the core intended to write the data value 0x11. For example, after one or more clock cycles, the entire core that triggered the breakpoint 1014 is stopped due to backpressure from the debug circuit 1012.

[0167] Advantageously, the debug circuitry 1012 can stop the core to enable analysis and extraction of transactions sent to memory, and / or trigger different breakpoints based on specific transactions. Advantageously, the debug circuitry 1012 enables improved visibility of actual transaction information to determine whether inputs from memory or outputs to memory are as expected. For example, if a transaction is not expected, the debug circuitry 1012 can gain visibility into an indication of a possible misconfiguration of the eleventh accelerator circuit 1002. In some such examples, the expected output from the (one or more) machine learning model(s) 124 can be compared with the actual output from the eleventh accelerator circuit 1002, which is instantiated using the same (one or more) machine learning model(s) 124. Advantageously, the debug circuitry 1012 can identify a mismatch based on the comparison.

[0168] In some examples, the first accelerator circuit 108, the second accelerator circuit 110, the third accelerator circuit 202, the fourth accelerator circuit 302, the fifth accelerator circuit 402, the sixth accelerator circuit 502, the seventh accelerator circuit 602, the eighth accelerator circuit 702, the ninth accelerator circuit 802, and / or the tenth accelerator circuit 1002 include components for executing an executable object to generate a data output based on a data input, and the executable object is based on a machine learning model (such as Figure 1 For example, the component for executing may be composed of Figure 2 20, the first execution circuit thread 806, the second execution circuit thread 808, the execution circuit 875, and / or more generally, the first core 212 and / or the second core 214. In some examples, the execution circuit 220, the first execution circuit thread 806, the second execution circuit thread 808, the execution circuit 875, and / or more generally, the first core 212 and / or the second core 214 may be implemented by, for example, Figure 14For example, the execution circuit 220, the first execution circuit thread 806, the second execution circuit thread 808, the execution circuit 875, and / or more generally, the first core 212 and / or the second core may be instantiated by the example processor circuit 1412 of FIG. Figure 15 The example general purpose processor circuit 1500 is instantiated to execute machine executable instructions such as those given by Figure 11 In some examples, the execution circuit 220, the first execution circuit thread 806, the second execution circuit thread 808, the execution circuit 875, and / or more generally, the first core 212 and / or the second core 214 can be instantiated by hardware logic circuitry that is configured to perform operations corresponding to the machine-readable instructions. Figure 16 ASIC or FPGA circuit 1600. Additionally or alternatively, execution circuit 220, first execution circuit thread 806, second execution circuit thread 808, execution circuit 875, and / or more generally, first core 212 and / or second core 214 may be instantiated by any other combination of hardware, software, and / or firmware. For example, execution circuit 220, first execution circuit thread 806, second execution circuit thread 808, execution circuit 875, and / or more generally, first core 212 and / or second core 214 may be implemented by at least one or more hardware circuits (e.g., processor circuits, discrete and / or integrated analog and / or digital circuits, FPGAs, ASICs, comparators, operational amplifiers (op-amps), logic circuits, etc.) configured to execute some or all of the machine-readable instructions and / or perform some or all of the operations corresponding to the machine-readable instructions without executing software or firmware, although other structures are equally suitable.

[0169] In some examples, first accelerator circuit 108, second accelerator circuit 110, third accelerator circuit 202, fourth accelerator circuit 302, fifth accelerator circuit 402, sixth accelerator circuit 502, seventh accelerator circuit 602, eighth accelerator circuit 702, ninth accelerator circuit 802, and / or tenth accelerator circuit 1002 include components for debugging the hardware accelerator. For example, the components for debugging can be implemented by debug circuit 112, first debug circuit 208, second debug circuit 210, debug circuit 404, debug circuit 406, debug circuit 804, and / or debug circuit 1012. In some examples, debug circuit 112, first debug circuit 208, second debug circuit 210, debug circuit 404, debug circuit 406, debug circuit 804, and / or debug circuit 1012 can be implemented by, for example, Figure 14For example, debug circuit 112, first debug circuit 208, second debug circuit 210, debug circuit 404, debug circuit 406, debug circuit 804, and / or debug circuit 1012 may be implemented by Figure 15 exemplified by an example general purpose processor circuit 1500 that executes machine-executable instructions, such as at least Figure 11 Blocks 1110, 1112, 1114, 1116, 1118 and / or Figure 13 In some examples, debug circuit 112, first debug circuit 208, second debug circuit 210, debug circuit 404, debug circuit 406, debug circuit 804, and / or debug circuit 1012 may be instantiated by hardware logic circuitry that is configured to perform operations corresponding to machine-readable instructions. Figure 16 ASIC or FPGA circuit 1600. Additionally or alternatively, debug circuit 112, first debug circuit 208, second debug circuit 210, debug circuit 404, debug circuit 406, debug circuit 804, and / or debug circuit 1012 may be instantiated by any other combination of hardware, software, and / or firmware. For example, debug circuit 112, first debug circuit 208, second debug circuit 210, debug circuit 404, debug circuit 406, debug circuit 804, and / or debug circuit 1012 may be implemented by at least one or more hardware circuits (e.g., processor circuits, discrete and / or integrated analog and / or digital circuits, FPGAs, ASICs, comparators, operational amplifiers (op-amps), logic circuits, etc.) configured to execute some or all of the machine-readable instructions and / or perform some or all of the operations corresponding to the machine-readable instructions without executing software or firmware, although other structures are equally suitable.

[0170] In some examples, the means for debugging includes means for receiving at least one of a data input or a data output. In some such examples, the means for receiving has at least one of an input coupled to an output of the means for executing or an output coupled to an input of the means for executing. For example, the means for receiving can be implemented by the first interface circuit 812 and / or the second interface circuit 818. In some such examples, the first means for receiving can be implemented by the first interface circuit 812, and the second means for receiving can be implemented by the second interface circuit 818.

[0171] In some examples, an input of the means for receiving is coupled to the means for storing, an output of the means for receiving is coupled to an input of the means for executing, and the means for receiving is configured to receive data input from the means for storing and, in response to a breakpoint not being triggered, provide the data input to the input of the means for executing. In some such examples, the means for executing is configured to provide data output from the output of the means for executing to the means for storing. In some examples, the first means for receiving can be implemented by the first interface circuit 812, and the second means for receiving can be implemented by the second interface circuit 818. In some examples, the first means for receiving can be implemented by the second interface circuit 818, and the second means for receiving can be implemented by the first interface circuit 812. In some examples, the first means for storing can be implemented by the memory 118. In some examples, the means for executing can be implemented by the first core 212, the execution circuit 220 of the first core 212, the second core 214, the execution circuit 220 of the second core 214, the first execution circuit thread 806, the second execution circuit thread 808, the execution circuit 875, and the like.

[0172] In some examples, an input of the means for executing is coupled to the means for storing, an input of the means for receiving is coupled to an output of the means for executing, an output of the means for receiving is coupled to the means for storing, and the means for executing is to receive data input from the means for storing and provide data output from the output of the means for executing to the input of the means for receiving. In some such examples, the means for receiving is to provide data output to the means for storing in response to a breakpoint not being triggered. In some examples, the first means for receiving can be implemented by the first interface circuit 812, and the second means for receiving can be implemented by the second interface circuit 818. In some examples, the first means for receiving can be implemented by the second interface circuit 818, and the second means for receiving can be implemented by the first interface circuit 812. In some examples, the first means for storing can be implemented by the memory 118. In some examples, means for executing may be implemented by first core 212, execution circuitry 220 of first core 212, second core 214, execution circuitry 220 of second core 214, first execution circuitry thread 806, second execution circuitry thread 808, execution circuitry 875, etc.

[0173] In some examples, the component for debugging is a first component for debugging, the component for receiving is a first component for receiving, an input of the first component for receiving is coupled to the component for storing, and the first component for receiving is to receive data input from the component for storing. In some such examples, the first component for receiving is to provide data input to the component for executing in response to a breakpoint not being triggered. In some examples, the second component for debugging the hardware accelerator includes a second component for receiving. In some such examples, an input of the second component for receiving is coupled to an output of the component for executing, an output of the second component for receiving is coupled to the component for storing, and the second component for debugging is to at least one of: receive data output from the component for executing, output at least one of the data input or the data output in response to a breakpoint being triggered, or output the data output to the component for storing in response to a breakpoint not being triggered.

[0174] In some examples, the means for debugging includes means for selecting the means for receiving. In some such examples, the means for selecting has an input coupled to an output of the means for receiving. For example, the means for selecting can be implemented by multiplexer circuit 826.

[0175] In some examples, the means for debugging includes means for outputting at least one of a data input or a data output in response to triggering a breakpoint associated with execution of the executable object. In some such examples, the means for outputting has an input coupled to an output of the means for selecting. For example, the means for outputting can be implemented by shift register 830.

[0176] In some examples where at least one of the data input or the data output includes a first value and the debugging component includes a component for controlling the debugging component, the controlling component is configured to obtain a second value corresponding to a breakpoint from the first storing component. In some such examples, the second storing component is coupled to the controlling component. In some such examples, the comparing component is coupled to the output of the receiving component. For example, a first input of the comparing component is coupled to the output of the receiving component, and a second input of the comparing component is coupled to the second storing component and the controlling component. In some examples, the controlling component controls the receiving component to provide at least one of the data input or the data output to the selecting component in response to a match between the first and second values ​​based on the comparison, in response to triggering a breakpoint, and the controlling component controls the receiving component to receive an indication of the match from the comparing component. In some such examples, the first debugging component may be implemented by first debug circuitry 208, debug circuitry 404, debug circuitry 804, and / or debug circuitry 1012. In some such examples, the second means for debugging may be implemented by second debug circuitry 210 and / or debug circuitry 406. In some such examples, the first means for storing may be implemented by configuration register(s) 810. In some such examples, the second means for storing may be implemented by first breakpoint register(s) 816 and / or second breakpoint register(s) 822. In some examples, the first means for receiving may be implemented by first interface circuitry 812, and the second means for receiving may be implemented by second interface circuitry 818. In some examples, the first means for receiving may be implemented by second interface circuitry 818, and the second means for receiving may be implemented by first interface circuitry 812. In some examples, at least one of the first means for debugging or the second means for debugging is included in the means for executing.

[0177] In some examples where the data input is a first data input and the data output is a first data output, the means for receiving is the first means for receiving, the means for executing is the first means for executing, the second means for receiving is to receive at least one of the second data input or the second data output, and the input of the second means for receiving is coupled to the second means for executing. In some such examples, the means for incrementing is to increment a counter, the output of the means for incrementing is coupled to a selection input of the means for selecting, and the means for incrementing is to output a first value of the counter to indicate to the means for selecting that the output of the first means for receiving the circuit is selected, and to output a second value of the counter to indicate to the means for selecting that the output of the second means for receiving the circuit is selected. In some such examples, the first means for executing may be configured by Figure 2 808, the execution circuit 875, and / or more generally, the first core 212 and / or the second core 214. In some such examples, the second means for executing may be implemented by Figure 2 808, execution circuitry 875, and / or more generally, first core 212 and / or second core 214. In some examples, the first means for receiving may be implemented by first interface circuitry 812, and the second means for receiving may be implemented by second interface circuitry 818. In some examples, the first means for receiving may be implemented by second interface circuitry 818, and the second means for receiving may be implemented by first interface circuitry 812. In some examples, the means for incrementing may be implemented by counter circuitry 828.

[0178] Figure 11-13 The following table shows a representative method for implementing Figure 1 The accelerator circuits 108, 110 (or any other accelerator circuits described herein, such as Figures 8A-8D accelerator circuit 802) and / or Figure 1 The debug circuit 112 (or any other debug circuit described herein, such as Figures 8A-8D 804) of an exemplary hardware logic circuit, machine readable instructions, a hardware implemented state machine, and / or any combination thereof. The machine readable instructions may be one or more executable programs or portions of executable programs for execution by a processor circuit, such as the following in conjunction with Figure 14 The processor circuit 1412 shown in the example processor platform 1400 discussed below and / or in conjunction with Figure 15and / or the example processor circuits discussed in 16. The program may be embodied in software stored on one or more non-transitory computer-readable storage media, such as a CD, floppy disk, hard disk drive (HDD), solid-state drive (SSD), DVD, Blu-ray disc, volatile memory (e.g., any type of random access memory (RAM), etc.), or non-volatile memory (e.g., electrically erasable programmable read-only memory (EEPROM), flash memory, HDD, SSD, etc.), associated with a processor circuit located in one or more hardware devices, but the entire program and / or portions thereof may alternatively be executed by one or more hardware devices other than the processor circuit and / or embodied in firmware or dedicated hardware. The machine-readable instructions may be distributed across multiple hardware devices and / or executed by two or more hardware devices (e.g., a server and a client hardware device). For example, the client hardware device may be implemented by an endpoint client hardware device (e.g., a hardware device associated with a user) or an intermediate client hardware device (e.g., a radio access network (RAN)) gateway that can facilitate communication between the server and the endpoint client hardware device. Similarly, non-transitory computer readable storage media may include one or more media located in one or more hardware devices. Figure 11-13 The flowcharts illustrated in the accompanying drawings describe example procedures, but many other methods of implementing the example accelerator circuits 108, 110 and / or the example debug circuit 112 may alternatively be used. For example, the order of execution of the blocks may be changed, and / or some of the blocks described may be changed, eliminated, or combined. Additionally or alternatively, any or all of the blocks may be implemented by one or more hardware circuits (e.g., processor circuits, discrete and / or integrated analog and / or digital circuits, FPGAs, ASICs, comparators, operational amplifiers (op-amps), logic circuits, etc.) configured to perform the corresponding operations without executing software or firmware. The processor circuits may be distributed in different network locations and / or locally on one or more hardware devices (e.g., a single-core processor (e.g., a single-core central processing unit (CPU)), a multi-core processor (e.g., a multi-core CPU), etc.) in a single machine, multiple processors distributed across multiple servers in a server rack, multiple processors distributed across one or more server racks, a CPU and / or FPGA located in the same package (e.g., the same integrated circuit (IC) package or two or more separate enclosures, etc.).

[0179] The machine-readable instructions described herein may be stored in one or more of a compressed format, an encrypted format, a segmented format, a compiled format, an executable format, a packaged format, and the like. The machine-readable instructions as described herein may be stored as data or data structures (e.g., as part of instructions, code, code representations, and the like) that can be utilized to create, manufacture, and / or generate machine-executable instructions. For example, the machine-readable instructions may be segmented and stored on one or more storage devices and / or computing devices (e.g., servers) located at the same or different locations on a network or a collection of networks (e.g., in the cloud, in an edge device, and the like). The machine-readable instructions may require one or more of installation, modification, adaptation, updating, combination, supplementation, configuration, decryption, decompression, unpacking, distribution, redistribution, compilation, and the like so that they are directly readable, interpretable, and / or executable by a computing device and / or other machine. For example, machine-readable instructions may be stored in multiple parts that are individually compressed, encrypted, and / or stored on separate computing devices, where the parts, when decrypted, decompressed, and / or combined, form a set of machine-executable instructions that implement one or more operations that together may form a program such as described herein.

[0180] In another example, the machine-readable instructions may be stored in a state in which they can be read by a processor circuit, but require the addition of a library (e.g., a dynamic link library (DLL)), a software development kit (SDK), an application programming interface (API), etc., in order to execute the machine-readable instructions on a particular computing device or other device. In another example, it may be necessary to configure the machine-readable instructions (e.g., stored settings, data inputs, recorded network addresses, etc.) before the machine-readable instructions and / or corresponding program(s) can be executed in whole or in part. Thus, as used herein, a machine-readable medium may include machine-readable instructions and / or program(s) regardless of the particular format or state of the machine-readable instructions and / or program(s) when stored or otherwise at rest or in transmission.

[0181] The machine-readable instructions described herein may be represented by any past, present, or future instruction language, scripting language, programming language, etc. For example, the machine-readable instructions may be represented using any of the following languages: C, C++, Java, C#, Perl, Python, JavaScript, Hypertext Markup Language (HTML), Structured Query Language (SQL), Swift, etc.

[0182] As mentioned above, Figure 11-13The example operations of may be implemented using executable instructions (e.g., computer and / or machine-readable instructions) stored on one or more non-transitory computer and / or machine-readable media, such as optical storage devices, magnetic storage devices, HDDs, flash memory, read-only memory (ROM), CDs, DVDs, caches, any type of RAM, registers, and / or any other storage device or storage disk in which information is stored for any duration (e.g., for an extended period of time, permanently, for temporary buffering, and / or for caching of information, for example). As used herein, the terms non-transitory computer-readable medium and non-transitory computer-readable storage medium are expressly defined to include any type of computer-readable storage device and / or storage disk, and to exclude propagating signals, and to exclude transmission media.

[0183] "Include" and "comprising" (and all their forms and tenses) are used herein as open-ended terms. Thus, whenever a claim employs any form of "include" or "comprising" (e.g., includes, comprises, contains, includes, has, etc.) as a preamble or within any type of claim recitation, it is to be understood that additional elements, terms, etc. may be present without falling outside the scope of the corresponding claim or recitation. As used herein, when the expression "at least" is used as a transition term, such as in the preamble of a claim, it is open-ended in the same way that the terms "include" and "comprising" are open-ended. When used in the form of, for example, A, B, and / or C, the term "and / or" refers to any combination or subset of A, B, C, such as (1) A alone, (2) B alone, (3) C alone, (4) A and B, (5) A and C, (6) B and C, or (7) A and B and C. As used herein in the context of describing structures, components, items, objects, and / or things, the expression “at least one of A and B” is intended to refer to implementations that include any of (1) at least one A, (2) at least one B, or (3) at least one A and at least one B. Similarly, as used herein in the context of describing structures, components, items, objects, and / or things, the expression “at least one of A or B” is intended to refer to implementations that include any of (1) at least one A, (2) at least one B, or (3) at least one A and at least one B. As used herein in the context of describing the performance or execution of processes, instructions, acts, activities, and / or steps, the expression “at least one of A and B” is intended to refer to implementations that include any of (1) at least one A, (2) at least one B, or (3) at least one A and at least one B. Similarly, as used herein in the context of describing the performance or execution of processes, instructions, actions, activities and / or steps, the expression "at least one of A or B" is intended to refer to an implementation that includes any of (1) at least one A, (2) at least one B, or (3) at least one A and at least one B.

[0184] As used herein, singular references (e.g., "a," "an," "first," "second," etc.) do not exclude the plural. As used herein, the term "a" or "an" object refers to one or more of the object. The terms "a" (or "an"), "one or more," and "at least one" are used interchangeably herein. Furthermore, although listed separately, multiple components, elements, or method actions can be implemented by, for example, the same entity or object. Additionally, although individual features may be included in different examples or claims, these features may be combined, and being included in different examples or claims does not imply that the combination of features is infeasible and / or disadvantageous.

[0185] Figure 11is a flow diagram representative of example machine-readable instructions and / or example operations 1100 that may be executed and / or instantiated by a processor circuit to perform debug operation(s) on an accelerator circuit. Figure 11 The machine readable instructions and / or operations 1100 begin at block 1102 where, Figure 1 The debugging application 114 generates (one or more) breakpoints associated with the machine learning (ML) model. For example, the debugging application 114 can generate a breakpoint to be triggered in response to a first value of the (one or more) machine learning inputs 204, a second value of the address at which the first value is to be read from the memory 118, a third value of the (one or more) machine learning outputs 206, a fourth value of the address at which the third value is to be written to the memory 118, etc. and / or a combination thereof. Figure 12 Example machine-readable instructions and / or example operations that may be executed and / or instantiated by a processor circuit to implement block 1102 are described.

[0186] At block 1104, the debug application 114 compiles an executable object based on at least one of the breakpoint(s) or the ML model to be executed by the accelerator circuit. For example, the debug application 114 may compile an executable binary object based on the machine learning model(s) 124, and the executable binary object may include breakpoints.

[0187] At block 1106, debug circuitry 112 configures at least one of the debug circuitry or the accelerator circuitry based on at least one of the breakpoint(s) or the ML model(s). For example, debug circuitry 208 may store a breakpoint in debug register(s) 216 to configure debug circuitry 208 to stop execution of the executable binary object in response to the breakpoint being hit or triggered. In some examples, in response to execution of the executable binary object, first core 212 may store a value(s) in configuration register(s) 222 that may be utilized to configure execution circuitry 220 based on machine learning model(s) 124. In some such examples, in response to execution of the executable binary object, first core 212 may store a breakpoint in configuration register(s) 222.

[0188] At block 1108, the accelerator circuitry 108, 110 executes the executable object to generate output(s) based on the input(s). For example, the first execution circuitry thread 806 may obtain a first machine learning input of the machine learning input(s) 204 from the memory 118 and generate a first machine learning output of the machine learning output(s) 206 based on the first machine learning input.

[0189] At block 1110, debug circuitry 112 determines whether to trigger breakpoint(s) based on the input(s). For example, debug circuitry 208 may determine to trigger a breakpoint in response to a first value of the address at which the first machine learning input is read from memory 118 matching a second value of the breakpoint. In some examples, debug circuitry 208 may determine to trigger a breakpoint in response to a third value of the first machine learning input matching a fourth value of the breakpoint.

[0190] If, at block 1110, the debug circuitry 112 determines that the breakpoint(s) are triggered based on the input(s), then, at block 1112, the debug circuitry 112 halts execution of the executable object. For example, the first comparator circuitry 814 may generate an output to the control circuitry 824 indicating that a breakpoint has been triggered in conjunction with the first execution circuit thread 806. In some such examples, the control circuitry 824 may generate a BP HIT signal 850 and instruct the first interface circuitry 812 to pull down the request signal 840, the first ready signal 842, the response signal 846, etc. of the first execution circuit thread 806 and the second execution circuit thread 808 to halt the flow of data from the memory 118.

[0191] In response to stopping execution of the executable object at block 1112, control proceeds to block 1116 to perform (one or more) debug operations. For example, the shift register 830 may output a debug transaction 852 corresponding to a breakpoint trigger based on the first machine learning input. Figure 13 Example machine-readable instructions and / or example operations that may be executed and / or instantiated by a processor circuit to implement block 1116 are described.

[0192] If at block 1110, debug circuitry 112 determines not to trigger the breakpoint(s) based on the input(s), control proceeds to block 1114 to determine whether to trigger the breakpoint(s) based on the output(s). For example, debug circuitry 208 may trigger a different portion of the memory 118 or execution circuitry in response to the first machine learning output being written thereto (e.g., Figure 8D In some examples, debug circuitry 208 may determine that a breakpoint is triggered in response to a seventh value of the first machine learning output matching an eighth value of the breakpoint.

[0193] If, at block 1114, debug circuitry 112 determines that the breakpoint(s) are triggered based on the output(s), then, at block 1112, debug circuitry 112 halts execution of the executable object. For example, first comparator circuitry 814 may determine that the first machine learning output (or data associated therewith) triggered the breakpoint. In some such examples, first comparator circuitry 814 may generate an output to control circuitry 824 to notify control circuitry 824 that the breakpoint has been triggered. In some such examples, control circuitry 824 generates BP HIT signal 850 and instructs first interface circuitry 812 to pull down request signal 840, first ready signal 842, and response signal 846 to halt data flow from first execution circuitry thread 806.

[0194] In response to stopping execution of the executable object at block 1112, control proceeds to block 1116 to perform debug operation(s). For example, shift register 830 may output debug transaction 852 corresponding to the triggering of a breakpoint based on the first machine learning output. In some examples, shift register 830 may output debug transaction 852 corresponding to any other type of breakpoint. In some examples, after the breakpoint is hit, one or more subsequent debug transactions in debug transaction 852 may be read out to configuration register(s) 810. In response to performing debug operation(s) at block 1116, control proceeds to block 1118 to provide ML input(s) to the execution circuitry to generate ML output(s) or write ML output(s) to memory.

[0195] If, at block 1114, the debug circuitry 112 determines not to trigger the breakpoint(s) based on the output(s), control proceeds to block 1118 to provide the ML input(s) to the execution circuitry to generate the ML output(s) or to write the ML output(s) to the memory. For example, the first interface circuitry 812 may provide the first machine learning input read from the memory 118 to the first execution circuitry thread 806 to cause the first execution circuitry thread 806 to generate the first machine learning output. In some examples, the first interface circuitry 812 may provide the first machine learning output from the first execution circuitry thread 806 to the memory 118. In response to providing the ML input(s) to the execution circuitry to generate the ML output(s) or to write the ML output(s) to the memory at block 1118, control proceeds to block 1120 to determine whether execution of the executable object is complete.

[0196] If at block 1120, the first accelerator circuit 108 and / or the second accelerator circuit 110 determines that execution of the executable object is not complete, then control returns to block 1108 to execute the executable object to generate output(s) based on the input(s). Figure 11 The machine readable instructions and / or operations 1100 end.

[0197] Figure 12 is a flow diagram representing example machine-readable instructions and / or example operations 1200 that may be executed and / or instantiated by a processor circuit to generate breakpoint(s) associated with a machine learning (ML) model. In some examples, the machine-readable instructions and / or operations 1200 may implement Figure 11 Block 1102. Figure 12 The machine-readable instructions and / or operations 1200 begin at block 1202, where the debug application 114 determines whether to add a breakpoint. For example, the debug application 114 may determine whether to add one or more core-specific breakpoints (e.g., breakpoints to be triggered on a per-core basis), one or more workload-specific breakpoints (e.g., breakpoints to be triggered on a per-workload basis), the like, and / or a combination thereof.

[0198] If at block 1202 the debug application 114 does not determine to add a breakpoint, then Figure 12 The machine readable instructions and / or operations 1200 end. For example, Figure 12 The machine readable instructions and / or operations 1200 may return to Figure 11 The machine-readable instructions and / or block 1104 of operation 1100 are to compile an executable object based on at least one of the breakpoint(s) or the ML model to be executed by the accelerator circuit.

[0199] If at block 1202 the debugging application 114 determines to add a breakpoint, then at block 1204 the debugging application 114 determines the type of breakpoint to be added. For example, the debugging application 114 may determine to add an immediate breakpoint (e.g., a breakpoint to be triggered at the start of a workload, Figure 9 breakpoints (e.g., a first breakpoint to be triggered based on an address or address range), data breakpoints (e.g., a breakpoint to be triggered based on a data value, Figure 9 's second breakpoint 918, etc.), and so on.

[0200] At block 1206, the debugging application 114 determines whether the breakpoint to be added is a core-specific breakpoint or a workload-specific breakpoint. For example, the debugging application 114 may determine that the breakpoint to be added is a core-specific breakpoint, which may be determined by Figure 9 In some examples, the debugging application 114 may determine that the breakpoint to be added is a workload-specific breakpoint, which may be implemented by Figure 10 The breakpoint 1014 is implemented.

[0201] If at block 1206, the debug application 114 determines that the breakpoint to be added is a core-specific breakpoint, control proceeds to block 1208 to write the breakpoint to the configuration register(s) of the corresponding core(s). For example, the debug application 114 may write a core-specific breakpoint to the configuration register(s) corresponding to the particular core. Figures 8A-8D In response to writing the breakpoint to the configuration register(s) of the corresponding core(s) at block 1208, control returns to block 1202 to determine whether to add another breakpoint.

[0202] If, at block 1206, the debug application 114 determines that the breakpoint to be added is a workload-specific breakpoint, control proceeds to block 1210 to compile the breakpoint into a workload executable object that is written into the configuration register(s) once deployed to the core for execution. For example, the debug application 114 may write the workload-specific breakpoint into an executable binary object (e.g., a workload executable binary object, a workload executable binary file, etc.) that is written into the configuration register(s) when deployed to the core for execution. Figures 8A-8D (one or more) configuration registers 810. In response to compiling the breakpoint into the workload executable object to be written into the (one or more) configuration registers once deployed to the core for execution at block 1210, control returns to block 1202 to determine whether to add another breakpoint.

[0203] Figure 13 is a flow diagram representing example machine-readable instructions and / or example operations 1300 that may be executed and / or instantiated by a processor circuit to perform debug operation(s). In some examples, the machine-readable instructions and / or operations 1300 may implement Figure 11 Block 1116. Figure 13 The machine-readable instructions and / or operations 1300 begin at block 1302, where the debug application 114 queries the debug circuitry for the invoked breakpoint(s). For example, the debug application 114 may retrieve the invoked breakpoint(s) from the configuration register(s) 810.

[0204] At block 1304, debug circuitry 112 outputs at least one of the machine learning (ML) input(s), the ML output(s), or the associated memory address(es). For example, debug application 114 may write a read transaction command 856 to configuration register(s) 810. In some such examples, control circuitry 824 may retrieve the read transaction command 856 from configuration register(s) 810 and generate a read transaction 854. In response to read transaction 854, shift register 830 may output a first machine learning input of first machine learning input(s) 204, a first machine learning output of machine learning output(s) 206, a memory address associated with the first machine learning input, a memory address associated with the first machine learning output, etc. to configuration register(s) 810 as part of debug transaction 852. In some such examples, the debug application 114 can retrieve from the configuration register(s) 810 a first machine learning input in the first machine learning input(s) 204, a first machine learning output in the machine learning output(s) 206, a memory address associated with the first machine learning input, a memory address associated with the first machine learning output, and the like.

[0205] At block 1306, the debug circuitry 112 and / or debug application 114 determines the completion progress of the workload(s) executed by the core(s). For example, the debug application 114 may request the completion status or progress of the executable binary object, the workload(s) to be executed by the first execution circuit thread 806, etc. from the configuration register(s) 810.

[0206] At block 1308, debug circuitry 112 and / or debug application 114 determines whether to modify data associated with the configuration image of the acceleration circuit. For example, debug application 114 may determine whether to modify, adjust, etc. the data to be generated by the acceleration circuit by writing different values ​​to the configuration register(s) in configuration register(s) 810. Figures 8A-8D The accelerator circuit 802 implements the portion(s) of the configuration image.

[0207] If, at block 1308, the debug circuitry 112 and / or debug application 114 determines not to modify the data associated with the configuration image of the acceleration circuitry, control proceeds to block 1312 to determine whether to modify the data associated with the ML model. If, at block 1308, the debug circuitry 112 and / or debug application 114 determines to modify the data associated with the configuration image of the acceleration circuitry, then, at block 1310, the debug application 114 adjusts the value(s) of the configuration register(s) to modify the configuration image. For example, the debug application 114 may modify, adjust, etc. the data to be modified by the ML model by writing different value(s) to the configuration register(s) in the configuration register(s) 810. Figures 8A-8D The accelerator circuit 802 implements the portion(s) of the configuration image.

[0208] In response to adjusting the value(s) of the configuration register(s) to modify the configuration image at block 1310, the debug circuitry 112 and / or the debug application 114 determines whether to modify data associated with the ML model at block 1312. For example, the debug application 114 may determine whether to adjust the value(s) of the machine learning input(s) 204 in the memory 118, the value(s) of the machine learning input(s) 204 in the first execution circuitry thread 806, the like, and / or a combination thereof.

[0209] If, at block 1312, the debug circuitry 112 and / or debug application 114 determine not to modify the data associated with the ML model, control proceeds to block 1318. If, at block 1312, the debug circuitry 112 and / or debug application 114 determine to modify the data associated with the ML model, then, at block 1314, the debug circuitry 112 and / or debug application 114 adjust the value(s) of the ML input(s) in the accelerator circuitry and / or memory. For example, the debug application 114 may change, modify, or adjust the value(s) of the machine learning input(s) 204 in the memory 118, the value(s) of the machine learning input(s) 204 in the first execution circuitry thread 806, and / or a combination thereof.

[0210] At block 1316, the debug circuitry 112 and / or debug application 114 adjusts the value(s) of the breakpoint(s). For example, the debug application 114 may write different value(s) for the breakpoint(s) stored in the configuration register(s) 810, the first breakpoint register(s) 816, the second breakpoint register(s) 822, etc., and / or combinations thereof.

[0211] At block 1318, debug circuitry 112 and / or debug application 114 determines whether to instruct the accelerator circuitry to perform (one or more) increment operations on the executable object. For example, debug application 114 may instruct debug circuitry 112 to perform one or more read, write, or compute operations. In some such examples, debug application 114 may write a single-step command 858 to configuration register(s) 810, which may cause control circuitry 824 to perform a single-step operation.

[0212] If at block 1318, debug circuitry 112 and / or debug application 114 determines not to instruct accelerator circuitry to perform the increment operation(s) of the executable object, then Figure 13 The machine readable instructions and / or operations 1300 end. For example, Figure 13 The machine readable instructions and / or operations 1300 may return to Figure 11 Machine readable instructions and / or block 1118 of operation 1100 to provide ML input(s) to execution circuitry to generate ML output(s) or write ML output(s) to memory.

[0213] If, at block 1318, the debug circuitry 112 and / or debug application 114 determines to instruct the accelerator circuitry to perform (one or more) increment operations of the executable object, then at block 1320, the debug circuitry 112 and / or debug application 114 performs (one or more) increment operations, which include at least one of (one or more) read, write, or compute operations. For example, the debug application 114 may instruct the accelerator circuitry to perform (one or more) increment operations of the executable object by means of the single-step command 858 and the control circuitry 824. Figure 8A The first interface circuit 812 facilitates the read operation by releasing the force of the request signal 840, the first ready signal 842, and the response signal 846 during the first clock cycle, and then pulling down the request signal 840, the first ready signal 842, and the response signal 846 after the first clock cycle. In some examples, the debug application 114 may instruct the control circuit 824 to read the request signal 840, the first ready signal 842, and the response signal 846 by means of the single-step command 858. Figure 8B The first interface circuit 812 facilitates the write operation by pulling up the request signal 840, the first ready signal 842, and the response signal 846 during the first clock cycle, and then pulling down the request signal 840, the first ready signal 842, and the response signal 846 after the first clock cycle. In some examples, the debug application 114 can instruct the shift register 830 to output the debug transaction 852 in response to the read transaction 854 by means of a read transaction command 856 and the control circuit 824.

[0214] In response to performing the increment operation(s) comprising at least one of a read, write, or compute operation(s) at block 1318, Figure 13The machine readable instructions and / or operations 1300 end. For example, Figure 13 The machine readable instructions and / or operations 1300 may return to Figure 11 The machine readable instructions and / or block 1120 of operation 1100 are to determine whether execution of the executable object is complete.

[0215] Figure 14 is a block diagram of an example processor platform 1400 that is configured to execute and / or instantiate Figure 11-13 Machine-readable instructions and / or operations to implement Figure 1 The processor platform 1400 may be, for example, a server, a personal computer, a workstation, a self-learning machine (e.g., a neural network), a mobile device (e.g., a cellular phone, a smartphone, an iPad, etc.), or a processor platform 1401. TM tablet computer), personal digital assistant (PDA), internet appliance, digital video recorder, Blu-ray player, game console, personal video recorder, set-top box, headset (e.g., augmented reality (AR) headset, virtual reality (VR) headset, etc.) or other wearable device, or any other type of computing device.

[0216] The processor platform 1400 of the illustrated example includes a processor circuit 1412. The processor circuit 1412 of the illustrated example is hardware. For example, the processor circuit 1412 can be implemented by one or more integrated circuits, logic circuits, FPGA microprocessors, CPUs, GPUs, DSPs, and / or microcontrollers from any desired family or manufacturer. The processor circuit 1412 can be implemented by one or more semiconductor-based (e.g., silicon-based) devices. In this example, the processor circuit 1412 implements Figure 1 The debug circuit 112 and debug application 114. For example, the processor circuit 1412 may implement Figures 8A-8D The debug circuit 804 and / or Figure 10 Debug circuit 1012.

[0217] The processor circuit 1412 of the illustrated example includes a local memory 1413 (e.g., cache, registers, etc.). The processor circuit 1412 of the illustrated example communicates with a main memory including a volatile memory 1414 and a non-volatile memory 1416 via a bus 1418. In some examples, the bus 1418 implements Figure 1 The volatile memory 1414 may be a synchronous dynamic random access memory (SDRAM), a dynamic random access memory (DRAM), Dynamic Random Access Memory The non-volatile memory 1416 may be implemented by flash memory and / or any other desired type of memory device. Access to the main memory 1414, 1416 of the illustrated example is controlled by a memory controller 1417.

[0218] The processor platform 1400 of the illustrated example further includes an interface circuit 1420. The interface circuit 1420 may be implemented by hardware according to any type of interface standard, such as an Ethernet interface, a Universal Serial Bus (USB) interface, a Bluetooth interface, or a USB interface. interface, near field communication (NFC) interface, PCI interface and / or PCIe interface.

[0219] In the illustrated example, one or more input devices 1422 are connected to the interface circuitry 1420. The input device(s) 1422 permit a user to enter data and / or commands into the processor circuitry 1412. The input device(s) 1422 may be implemented by, for example, an audio sensor, a microphone, a camera (still or video), a keyboard, buttons, a mouse, a touch screen, a trackpad, a trackball, an isopoint device, and / or a voice recognition system.

[0220] One or more output devices 1424 are also connected to the interface circuit 1420 of the illustrated example. The output device(s) 1424 may be implemented, for example, by a display device (e.g., a light emitting diode (LED), an organic light emitting diode (OLED), a liquid crystal display (LCD), a cathode ray tube (CRT) display, an in-situ switching (IPS) display, a touch screen, etc.), a tactile output device, a printer, and / or a speaker. Thus, the interface circuit 1420 of the illustrated example typically includes a graphics driver card, a graphics driver chip, and / or a graphics processor circuit such as a GPU. In this example, the output device(s) 1424 implements Figure 1 The user interface 130 is configured as follows.

[0221] The interface circuitry 1420 of the illustrated example also includes communication devices, such as transmitters, receivers, transceivers, modems, residential gateways, wireless access points, and / or network interfaces, to facilitate data exchange with external machines (e.g., any kind of computing device) over a network 1426. This communication can be through, for example, an Ethernet connection, a digital subscriber line (DSL) connection, a telephone line connection, a coaxial cable system, a satellite system, a line-of-site wireless system, a cellular telephone system, an optical connection, etc.

[0222] The processor platform 1400 of the illustrated example also includes one or more mass storage devices 1428 for storing software and / or data. Examples of such mass storage devices 1428 include magnetic storage devices, optical storage devices, floppy disk drives, HDDs, CDs, Blu-ray disk drives, redundant arrays of independent disks (RAID) systems, solid-state storage devices such as flash memory devices and / or SSDs, and DVD drives.

[0223] Can be Figure 11-13 The machine-executable instructions 1432 implemented by the machine-readable instructions may be stored in the mass storage device 1428, in the volatile memory 1414, in the non-volatile memory 1416, and / or in a removable non-transitory computer-readable storage medium such as a CD or DVD.

[0224] Figure 14 The processor platform 1400 of the illustrated example includes Figure 1 The first accelerator circuit 108 and the second accelerator circuit 110 are configured to communicate with different hardware of the processor platform 1400 , such as a volatile memory 1414 and a non-volatile memory 1416 , via a bus 1418 .

[0225] Figure 15 yes Figure 14 1412. In this example, Figure 14 The processor circuit 1412 is implemented by a general-purpose microprocessor 1500. The general-purpose microprocessor circuit 1500 executes Figure 11-13 some or all of the machine-readable instructions of the flowchart to effectively Figure 1 The first accelerator circuit 108, the second accelerator circuit 110 and / or the debug circuit 112 are instantiated as logic circuits to perform operations corresponding to those machine-readable instructions. For example, the microprocessor 1500 can implement a multi-core hardware circuit, such as a CPU, a DSP, a GPU, an XPU, etc. Although it can include any number of example cores 1502 (for example, 1 core), the microprocessor 1500 of this example is a multi-core semiconductor device including N cores. The cores 1502 of the microprocessor 1500 can operate independently or can cooperate to execute machine-readable instructions. For example, the machine code corresponding to a firmware program, an embedded software program, or a software program can be executed by one of the cores 1502, or can be executed by multiple cores in the cores 1502 at the same or different times. In some examples, the machine code corresponding to the firmware program, the embedded software program, or the software program is divided into threads and executed in parallel by two or more of the cores 1502. The software program may correspond to a program executed by Figure 11-13The flowcharts represent a portion or all of machine-readable instructions and / or operations.

[0226] The cores 1502 can communicate via a first example bus 1504. In some examples, the first bus 1504 can implement a communication bus to facilitate communications associated with one or more cores in the cores 1502. For example, the first bus 1504 can implement at least one of an Inter-Integrated Circuit (I2C) bus, a Serial Peripheral Interface (SPI) bus, a PCI bus, or a PCIe bus. Additionally or alternatively, the first bus 1504 can implement any other type of computing or electrical bus. The cores 1502 can obtain data, instructions, and / or signals from one or more external devices via example interface circuitry 1506. The cores 1502 can output data, instructions, and / or signals to one or more external devices via interface circuitry 1506. While the cores 1502 of this example include example local memory 1520 (e.g., a level 1 (L1) cache that may be partitioned into an L1 data cache and an L1 instruction cache), the microprocessor 1500 also includes example shared memory 1510 that may be shared by the cores (e.g., a level 2 (L2 cache)) for high-speed access to data and / or instructions. Data and / or instructions may be transferred (e.g., shared) by writing to and / or reading from the shared memory 1510. The local memory 1520 and the shared memory 1510 of each of the cores 1502 may be a memory comprising multiple levels of cache memory and main memory (e.g., Figure 14 Cache memory is part of a storage device hierarchy (e.g., main memory 1414, 1416). Typically, higher-level memory in the hierarchy exhibits shorter access times and has smaller storage capacity than lower-level memory. Changes in the various levels of the cache hierarchy are managed (e.g., coordinated) by a cache coherence policy.

[0227] Each core 1502 may be referred to as a CPU, DSP, GPU, or any other type of hardware circuit. Each core 1502 includes a control unit circuit 1514, an arithmetic and logic (AL) circuit (sometimes referred to as an ALU) 1516, a plurality of registers 1518, an L1 cache 1520, and a second example bus 1522. Other configurations are possible. For example, each core 1502 may include a vector unit circuit, a single instruction multiple data (SIMD) unit circuit, a load / store unit (LSU) circuit, a branch / jump unit circuit, a floating point unit (FPU) circuit, and the like. The control unit circuit 1514 includes semiconductor-based circuitry configured to control (e.g., coordinate) data movement within the corresponding core 1502. The AL circuit 1516 includes semiconductor-based circuitry configured to perform one or more mathematical and / or logical operations on data within the corresponding core 1502. Some examples of the AL circuit 1516 perform integer-based operations. In other examples, the AL circuit 1516 also performs floating-point operations. In yet other examples, the AL circuit 1516 may include a first AL circuit that performs integer-based operations and a second AL circuit that performs floating-point operations. In some examples, the AL circuit 1516 may be referred to as an arithmetic logic unit (ALU). The register 1518 is a semiconductor-based structure to store data and / or instructions, such as one or more results of operations performed by the AL circuit 1516 of the corresponding core 1502. For example, the register 1518 may include (one or more) vector registers, (one or more) SIMD registers, (one or more) general registers, (one or more) flag registers, (one or more) segment registers, (one or more) machine-specific registers, (one or more) instruction pointer registers, (one or more) control registers, (one or more) debug registers, (one or more) memory management registers, (one or more) machine check registers, etc. As Figure 15 As shown in , registers 1518 can be arranged into groups. Alternatively, registers 1518 can be organized in any other arrangement, format, or structure, including being distributed throughout core 1502 to reduce access time. Second bus 1522 can implement at least one of an I2C bus, an SPI bus, a PCI bus, or a PCIe bus.

[0228] Each core 1502 and / or more generally the microprocessor 1500 may include additional and / or alternative structures to those shown and described above. For example, there may be one or more clock circuits, one or more power supplies, one or more power gates, one or more cache home agents (CHAs), one or more converged / common mesh stops (CMSs), one or more shifters (e.g., (one or more) barrel shifters), and / or other circuits. The microprocessor 1500 is a semiconductor device that is manufactured to include many interconnected transistors to implement the above structure in one or more integrated circuits (ICs) contained in one or more packages. The processor circuit may include one or more accelerators and / or collaborate with one or more accelerators. In some examples, the accelerator is implemented by a logic circuit to perform certain tasks faster and / or more efficiently than can be performed by a general-purpose processor. Examples of accelerators include ASICs and FPGAs, such as those discussed herein. A GPU or other programmable device may also be an accelerator. The accelerator may be on the processor circuit, in the same chip package as the processor circuit, and / or in one or more packages separate from the processor circuit.

[0229] Figure 16 yes Figure 14 14. In this example, the processor circuit 1412 is implemented by an FPGA circuit 1600. The FPGA circuit 1600 can be used, for example, to execute a program that would otherwise be executed by executing corresponding machine-readable instructions. Figure 15 However, once configured, FPGA circuit 1600 instantiates machine-readable instructions in hardware and, therefore, can often perform operations faster than they could be performed by a general-purpose microprocessor executing corresponding software.

[0230] More specifically, Figure 15 The microprocessor 1500 (which is a general purpose device that can be programmed to perform the Figure 11-13 Compared to a process where some or all of the machine-readable instructions are represented by a flowchart, but whose interconnections and logic circuits are fixed once manufactured, Figure 16 The example FPGA circuit 1600 includes interconnects and logic circuits that can be configured and / or interconnected in different ways after fabrication to instantiate, for example, Figure 11-13In particular, FPGA circuit 1600 can be thought of as an array of logic gates, interconnects, and switches. The switches can be programmed to change how the logic gates are interconnected via the interconnects, effectively forming one or more dedicated logic circuits (unless and until FPGA circuit 1600 is reprogrammed). The configured logic circuits enable the logic gates to work together in different ways to perform different operations on data received by the input circuits. Those operations may correspond to the operations performed by Figure 11-13 As such, FPGA circuit 1600 can be constructed to effectively Figure 11-13 Some or all of the machine-readable instructions of the flowchart of FIG are instantiated into dedicated logic circuits to perform operations corresponding to those software instructions in a dedicated manner similar to an ASIC. Therefore, the FPGA circuit 1600 can perform operations corresponding to the software instructions more efficiently than a general-purpose microprocessor can. Figure 11-13 Some or all of the operations in the machine-readable instructions may be performed more quickly.

[0231] exist Figure 16 In the example of FPGA circuit 1600, FPGA circuit 1600 is configured to be programmed (and / or reprogrammed one or more times) by an end user via a hardware description language (HDL) such as Verilog. Figure 16 FPGA circuit 1600 includes example input / output (I / O) circuit 1602 to obtain data from and / or output data to example configuration circuit 1604 and / or external hardware (e.g., external hardware circuit) 1606. For example, configuration circuit 1604 can implement interface circuitry that can obtain machine-readable instructions to configure FPGA circuit 1600 or portion(s) thereof. In some such examples, configuration circuit 1604 can obtain machine-readable instructions from a user, a machine (e.g., a hardware circuit (e.g., programmed or dedicated circuit) that can implement an artificial intelligence / machine learning (AI / ML) model to generate instructions), etc. In some examples, external hardware 1606 can implement Figure 15 The FPGA circuit 1600 also includes an array of example logic gate circuits 1608, a plurality of example configurable interconnects 1610, and example storage circuits 1612. The logic gate circuits 1608 and the interconnects 1610 are configurable to instantiate one or more operations that may correspond to Figure 11-13 At least some of the machine-readable instructions and / or other desired operations. Figure 16The logic gate circuits 1608 shown in FIG are manufactured in groups or blocks. Each block includes a semiconductor-based electrical structure that can be configured into a logic circuit. In some examples, the electrical structure includes logic gates (e.g., AND gates, OR gates, NOR gates, etc.) that provide the basic building blocks of the logic circuit. Electrically controllable switches (e.g., transistors) are present in each of the logic gate circuits 1608 to enable configuration of the electrical structure and / or logic gates to form a circuit to perform the desired operation. The logic gate circuits 1608 may include other electrical structures, such as a lookup table (LUT), a register (e.g., a flip-flop or latch), a multiplexer, etc.

[0232] The interconnect 1610 of the illustrated example is a conductive path, trace, via, etc., which may include an electrically controllable switch (e.g., a transistor) whose state can be changed by programming (e.g., using an HDL instruction language) to activate or deactivate one or more connections between one or more of the logic gate circuits 1608, thereby programming the desired logic circuit.

[0233] The storage circuit 1612 of the illustrated example is configured to store one or more results of one or more operations performed by the corresponding logic gates. The storage circuit 1612 may be implemented by registers, etc. In the illustrated example, the storage circuit 1612 is distributed among the logic gate circuits 1608 to facilitate access and increase execution speed.

[0234] Figure 16 The example FPGA circuit 1600 also includes an example dedicated operation circuit 1614. In this example, the dedicated operation circuit 1614 includes a dedicated circuit 1616, which can be called to implement common functions, thereby avoiding the need to program those functions in the field. Examples of such dedicated circuits 1616 include memory (e.g., DRAM) controller circuits, PCIe controller circuits, clock circuits, transceiver circuits, memories, and multiplier-accumulator circuits. Other types of dedicated circuits may exist. In some examples, the FPGA circuit 1600 may also include an example general-purpose programmable circuit 1618, such as an example CPU 1620 and / or an example DSP 1622. Other general-purpose programmable circuits 1618 may additionally or alternatively be present, such as a GPU, XPU, etc. that can be programmed to perform other operations.

[0235] Although Figure 15 and Figure 16 Pictured Figure 14 These are two example implementations of the processor circuit 1412, but many other approaches are also contemplated. For example, as mentioned above, the modem FPGA circuit may include an onboard CPU such as Figure 16 One or more of the example CPUs 1620. Thus, Figure 14 The processor circuit 1412 may additionally be combined with Figure 15 An example microprocessor 1500 and Figure 16 In some such hybrid examples, the Figure 11-13 The first portion of the machine-readable instructions represented by the flowchart may be represented by Figure 15 1502 and is executed by one or more of the cores 1502 and by Figure 11-13 The second portion of the machine-readable instructions represented by the flowchart may be represented by Figure 16 FPGA circuit 1600 executes.

[0236] In some examples, Figure 14 The processor circuit 1412 may be in one or more packages. For example, Figure 15 The processor circuit 1500 and / or Figure 16 The FPGA circuit 1600 can be in one or more packages. In some examples, the XPU can be composed of Figure 14 The XPU may be implemented as a processor circuit 1412, which may be in one or more packages. For example, the XPU may include a CPU in one package, a DSP in another package, a GPU in yet another package, and an FPGA in yet another package.

[0237] Figure 17 is a block diagram illustrating an example software distribution platform 1705 for distributing software such as Figure 14 The example machine-readable instructions 1432 of software are provided for distributing software to hardware devices owned and / or operated by a third party. The example software distribution platform 1705 can be implemented by any computer server, data facility, cloud service, etc. capable of storing software and transferring it to other computing devices. Examples of third parties can include client devices associated with end users and / or customers (e.g., for licensing, sale, and / or use), retailers (e.g., for sale, resale, licensing, and / or sublicensing), and / or original equipment manufacturers (OEMs) (e.g., for inclusion in products to be distributed to, for example, retailers and / or other end users such as direct purchasing customers). In some examples, the third party can be a customer of the entity that owns and / or operates the software distribution platform 1705. For example, the entity that owns and / or operates the software distribution platform 1705 can be a company such as Figure 14The third party may be a customer, user, retailer, OEM, etc. who purchases and / or licenses the software for use and / or resale and / or sublicense. In the illustrated example, the software distribution platform 1705 includes one or more servers and one or more storage devices. The storage device stores the machine readable instructions 1432, which may correspond to the instructions described above. Figure 11-13 Example machine readable instructions and / or operations 1100, 1200, 1300 of the example software distribution platform 1705. One or more servers of the example software distribution platform 1705 communicate with a network 1710, which may correspond to any one or more of the Internet and / or any of the example networks 132, 1426 described above. In some examples, as part of a commercial transaction, one or more servers respond to a request to transfer software to a requesting party. Payment for the delivery, sale, and / or license of the software may be handled by one or more servers of the software distribution platform and / or by a third-party payment entity. The server enables a purchaser and / or licensee to download machine readable instructions 1432 from the software distribution platform 1705. For example, it may correspond to Figure 11-13 The example machine readable instructions and / or software for the operations 1100, 1200, 1300 may be downloaded to the example processor platform 1400, which is to execute the machine readable instructions 1432 to implement Figure 1 The first accelerator circuit 108, the second accelerator circuit 110, the debugging circuit 112, and / or the debugging application 114. In some examples, one or more servers of the software distribution platform 1705 periodically supplies, transmits, and / or forces the software (e.g., Figure 14 Example machine readable instructions 1432) to ensure that improvements, patches, updates, etc. are distributed and applied to the software at the end user device.

[0238] Based on the foregoing, it will be appreciated that example systems, methods, devices, and articles of manufacture for debugging accelerator hardware have been disclosed. The disclosed systems, methods, devices, and articles of manufacture allow a unified software approach during debugging by having dedicated debug circuitry for debugging. For example, any accelerator core can be stopped at a given time and incrementally executed (e.g., single-stepped) using existing hardware with breakpoints (e.g., breakpoint instructions) disclosed herein. The disclosed systems, methods, devices, and articles of manufacture can allow any executable binary object (e.g., the output of a compiler dispatched to a hardware accelerator) to be used and debugged. The disclosed systems, methods, devices, and articles of manufacture implement the execution of a stopped core to incrementally execute a workload using a core with read and debug transactions, thereby detecting (one or more) transactions that fall outside of expected behavior (e.g., expected value, expected address, etc.). The disclosed systems, methods, devices, and articles of manufacture utilize controlled techniques to implement the output of debug transactions, which allow detection of memory write operations that mistakenly overwrite each other to improve visibility throughout the hardware accelerator pipeline.

[0239] The disclosed systems, methods, apparatus, and articles of manufacture implement automatic detection of pre-programmed data in a generated data stream to identify at which point in execution an unexpected occurrence of a data segment was generated, and which workload on which core was responsible for the unexpected occurrence. The disclosed systems, methods, apparatus, and articles of manufacture implement the ability to set breakpoints on specific memory transaction addresses or address ranges to enable improved identification of unexpected operations.

[0240] The disclosed systems, methods, apparatus, and articles of manufacture improve the efficiency of using a computing device by improving and / or otherwise optimizing the execution of a hardware accelerator in response to the identification and correction of an erroneous accelerator configuration. Thus, the disclosed systems, methods, apparatus, and articles of manufacture are directed to one or more improvements in the operation of a machine, such as a computer or other electronic and / or mechanical device.

[0241] Disclosed herein are example methods, apparatuses, systems, and articles of manufacture for debugging accelerator hardware. Further examples and combinations thereof include the following:

[0242] Example 1 includes an apparatus for debugging a hardware accelerator, the apparatus comprising: a core having a core input and a core output, the core to execute executable code to generate a data output based on a data input, the executable code being based on a machine learning model; and debug circuitry coupled to at least one of the core input or the core output, the debug circuitry comprising: an interface circuit having at least one of an interface input coupled to the core output or an interface output coupled to the core input, the interface circuit to receive at least one of the data input or the data output; a multiplexer circuit having a multiplexer input and a multiplexer output, the multiplexer input coupled to the interface output; and a shift register having a shift register input coupled to the multiplexer output, the shift register to output at least one of the data input or the data output in response to triggering of a breakpoint associated with execution of the executable code.

[0243] Example 2 includes the apparatus of Example 1, wherein the interface input is coupled to a memory, the interface output is coupled to a core input, and wherein the interface circuit is to receive a data input from the memory and provide the data input to the core input in response to the breakpoint not being triggered, and the core is to provide a data output from the core output to the memory.

[0244] Example 3 includes the apparatus of Example 1, wherein the core input is coupled to a memory, the interface input is coupled to a core output, the interface output is coupled to the memory, and wherein the core is to receive a data input from the memory and to provide a data output from the core output to the interface input, and the interface circuit is to provide the data output to the memory in response to the breakpoint not being triggered.

[0245] Example 4 includes the apparatus of Example 1, wherein the debug circuit is a first debug circuit, the interface circuit is a first interface circuit, the interface input is a first interface input, the interface output is a first interface output, the first interface input is coupled to a memory, the first interface circuit is to receive data input from the memory, and further includes: a first interface circuit to provide data input to the core in response to a breakpoint not being triggered; a second debug circuit including a second interface circuit having a second interface input and a second interface output, the second interface input being coupled to the core output, the second interface output being coupled to the memory, the second debug circuit being to at least one of: receive data output from the core, output at least one of the data input or the data output in response to a breakpoint being triggered, or output the data output to the memory in response to a breakpoint not being triggered.

[0246] Example 5 includes the apparatus of Example 1, wherein the debug circuit is included in the core, the hardware accelerator is a neural network accelerator, the machine learning model is a neural network, and wherein the core is to execute executable code to generate data output based on data input, the executable code includes a breakpoint, the executable code is based on at least one of the neural network or the breakpoint, and the debug circuit is to trigger the breakpoint to stop execution of the executable code and output at least one of the data input, the data output, or the breakpoint.

[0247] Example 6 includes the apparatus of Example 1, wherein at least one of the data input or the data output includes a first value, and the debug circuit includes: a control circuit to obtain a second value corresponding to a breakpoint from a configuration register of the core; a breakpoint register coupled to the control circuit, the breakpoint register to store the second value; a comparator circuit having a first comparator input and a second comparator input, the first comparator input coupled to the interface output, the second comparator input coupled to the breakpoint register and the control circuit, the comparator circuit to compare the first value and the second value, and the control circuit to instruct the interface circuit to provide at least one of the data input or the data output to the multiplexer circuit in response to a match based on the compared first value and the second value, the breakpoint being triggered in response to the match, and the control circuit to instruct the interface circuit to receive an indication of the match from the comparator circuit.

[0248] Example 7 includes the apparatus of Example 1, wherein the interface circuit is a first interface circuit, the interface input is a first interface input, the interface output is a first interface output, the core includes a first thread and a second thread, the first thread is coupled to the first interface input, and further includes: a second interface circuit having a second interface input and a second interface output, the second interface input is coupled to the second thread; and a counter circuit having a counter output coupled to a selection input of a multiplexer circuit, the counter circuit being configured to output a first value to instruct the multiplexer circuit to select the output of the first interface circuit, and to output a second value to instruct the multiplexer circuit to select the output of the second interface circuit.

[0249] Example 8 includes an apparatus for debugging a hardware accelerator, the apparatus comprising: a component for executing executable code to generate a data output based on a data input, the executable code being based on a machine learning model; and a component for debugging the hardware accelerator, the component for debugging being coupled to the component for executing, the component for debugging comprising: a component for receiving at least one of a data input or a data output, the component for receiving having at least one of an input coupled to an output of the component for executing or an output coupled to an input of the component for executing; a component for selecting the component for receiving, the component for selecting having an input coupled to the output of the component for receiving; and a component for outputting at least one of the data input or the data output in response to triggering of a breakpoint associated with execution of the executable code, the component for outputting having an input coupled to the output of the component for selecting.

[0250] Example 9 includes the apparatus of Example 8, wherein the input of the means for receiving is coupled to the means for storing, the output of the means for receiving is coupled to the input of the means for executing, and wherein: the means for receiving is to receive a data input from the means for storing and, in response to a breakpoint not being triggered, provide the data input to the input of the means for executing; and the means for executing is to provide a data output from the output of the means for executing to the means for storing.

[0251] Example 10 includes the apparatus of Example 8, wherein an input of the means for executing is coupled to the means for storing, an input of the means for receiving is coupled to an output of the means for executing, the output of the means for receiving is coupled to the means for storing, and wherein: the means for executing is to receive a data input from the means for storing and provide a data output from the output of the means for executing to the input of the means for receiving; and the means for receiving is to provide the data output to the means for storing in response to a breakpoint not being triggered.

[0252] Example 11 includes the apparatus of Example 8, wherein the means for debugging is a first means for debugging, the means for receiving is a first means for receiving, an input of the first means for receiving being coupled to the means for storing, the first means for receiving being configured to receive data input from the means for storing, and further comprising: the first means for receiving being configured to provide data input to the means for executing in response to a breakpoint not being triggered; and a second means for debugging a hardware accelerator, the second means for debugging comprising a second means for receiving, the input of the second means for receiving being coupled to the output of the means for executing, the output of the second means for receiving being coupled to the means for storing, the second means for debugging being configured to perform at least one of the following: receiving data output from the means for executing, outputting at least one of the data input or the data output in response to triggering a breakpoint, or outputting the data output to the means for storing in response to a breakpoint not being triggered.

[0253] Example 12 includes the apparatus of Example 8, wherein the means for debugging is included in the means for executing, the hardware accelerator is a neural network accelerator, the machine learning model is a neural network, and wherein: the means for executing is to execute executable code to generate data output based on data input, the executable code includes a breakpoint, the executable code is based on at least one of the neural network or the breakpoint; and the means for debugging is to trigger the breakpoint to stop execution of the executable code and output at least one of the data input, the data output, or the breakpoint.

[0254] Example 13 includes the apparatus of Example 8, wherein at least one of the data input or the data output comprises a first value, and the means for debugging comprises: means for controlling the means for debugging, the means for controlling being configured to obtain a second value corresponding to a breakpoint from the first means for storing; second means for storing the second value, the second means for storing being coupled to the means for controlling; means for comparing being configured to compare the first value and a second value, a first input of the means for comparing being coupled to the output of the means for receiving, a second input of the means for comparing being coupled to the second means for storing and the means for controlling; and the means for controlling being configured to control the means for receiving to provide at least one of the data input or the data output to the means for selecting in response to a match between the first value and the second value based on the comparison, the breakpoint being triggered in response to the match, the means for controlling being configured to control the means for receiving to receive an indication of the match from the means for comparing.

[0255] Example 14 includes the apparatus of Example 8, wherein the data input is a first data input, the data output is a first data output, the component for receiving is the first component for receiving, the component for executing is the first component for executing, and further comprising: a second component for receiving at least one of a second data input or a second data output, the input of the second component for receiving being coupled to the second component for executing; and a component for incrementing a counter, the output of the component for incrementing being coupled to the selection input of the component for selecting, the component for incrementing being operable to output a first value of the counter to indicate that the component for selecting selects the output of the first component for receiving, and output a second value of the counter to indicate that the component for selecting selects the output of the second component for receiving.

[0256] Example 15 includes an apparatus for debugging a hardware accelerator, the apparatus comprising: at least one memory; instructions in the apparatus; and a processor circuit to execute or instantiate at least one of the instructions to generate a breakpoint associated with a machine learning model, compile executable code based on at least one of the machine learning model or the breakpoint, the executable code to be executed by the processor circuit to generate a data output based on a data input, trigger the breakpoint in response to execution of the executable code to stop execution of the executable code, and output at least one of the data input, the data output, or the breakpoint using debugging circuitry included in the processor circuit.

[0257] Example 16 includes the apparatus of Example 15, wherein the processor circuit is configured to identify a breakpoint to be triggered on a per-workload basis, insert the breakpoint into executable code to be called on a per-workload basis, and in response to execution of the executable code by a first core of the processor circuit, stop execution of the executable code by the first core when the breakpoint is triggered by the first core, and in response to execution of the executable code by a second core of the processor circuit, stop execution of the executable code by the second core when the breakpoint is triggered by the second core.

[0258] Example 17 includes the apparatus of Example 15, wherein the processor circuit is configured to identify a breakpoint to be triggered on a per-core basis, identify the breakpoint to be written into a first configuration register of a first core of the processor circuit rather than a second configuration register of a second core of the processor circuit, and write the breakpoint into the first configuration register, wherein triggering of the breakpoint is configured to stop execution of the executable code by the first core while the second core continues execution of the executable code.

[0259] Example 18 includes the apparatus of Example 15, wherein the data input includes first data, the data output includes second data, and the processor circuit is configured to identify a breakpoint to be triggered based on third data, write the third data into a configuration register of a core of the processor circuit, perform a first comparison of the first data and the third data, the breakpoint being triggered in response to a first match between the first data and the third data based on the first comparison, and perform a second comparison of the second data and the third data, the breakpoint being triggered in response to a second match between the second data and the third data based on the second comparison.

[0260] Example 19 includes the apparatus of Example 15, wherein the processor circuit is to identify a breakpoint to be triggered based on a first address in a memory associated with the data output, write the first address to a configuration register of a core of the processor circuit, identify a second address in the memory at which the data output is written in response to executing the executable code, and perform a comparison of the first address and the second address, the triggering of the breakpoint being responsive to a match between the first address and the second address.

[0261] Example 20 includes the apparatus of Example 15, wherein the data input is a first data input, and the processor circuit is to: in response to triggering of a breakpoint, obtain a control signal instructing execution of an increment operation of the executable code, the increment operation comprising at least one of a read operation to read a first value, a write operation to write a second value, or a calculation operation to determine a third value based on the second data input; and output at least one of the first value, the second value, or the third value.

[0262] Example 21 includes the apparatus of Example 15, wherein the processor circuit is operable to, in response to triggering of a breakpoint, perform at least one of: adjusting a first value of a data input, adjusting a second value of a first register of a core of the processor circuit, or adjusting a third value of a second register of a debug circuit; and resuming execution of the executable code based on at least one of the first value, the second value, or the third value.

[0263] Example 22 includes at least one non-transitory computer-readable medium comprising instructions that, when executed, cause a first processor circuit to at least: generate a breakpoint associated with a machine learning model; compile executable code based on the machine learning model or at least one of the breakpoints, the executable code to be executed by the first processor circuit or the second processor circuit to generate a data output based on a data input; and in response to execution of the executable code: trigger a breakpoint to halt execution of the executable code, and output at least one of the data input, the data output, or the breakpoint using debugging circuitry included in the first processor circuit or the second processor circuit.

[0264] Example 23 includes at least one non-transitory computer-readable medium of Example 22, wherein the instructions, when executed, cause the first processor circuit to: identify a breakpoint to be triggered on a per-workload basis, insert the breakpoint into executable code to be called on a per-workload basis, in response to execution of the executable code by a first core of the first processor circuit or the second processor circuit, stop execution of the executable code by the first core when the breakpoint is triggered by the first core, and in response to execution of the executable code by a second core of the first processor circuit or the second processor circuit, stop execution of the executable code by the second core when the breakpoint is triggered by the second core.

[0265] Example 24 includes at least one non-transitory computer-readable medium of Example 22, wherein the instructions, when executed, cause the first processor circuit to identify a breakpoint to be triggered on a per-core basis, identify a breakpoint to be written into a first configuration register of a first core of the first processor circuit or the second processor circuit but not into a second configuration register of a second core of the first processor circuit or the second processor circuit, and write the breakpoint into the first configuration register, wherein triggering of the breakpoint is used to stop execution of the executable code by the first core while the second core is to continue execution of the executable code.

[0266] Example 25 includes the at least one non-transitory computer-readable medium of Example 22, wherein the data input includes first data, the data output includes second data, and the instructions, when executed, cause the first processor circuit to: identify a breakpoint to be triggered based on third data, write the third data to a configuration register of a core of the first processor circuit or the second processor circuit, perform a first comparison of the first data and the third data, the breakpoint being triggered in response to a first match between the first data and the third data based on the first comparison, and perform a second comparison of the second data and the third data, the breakpoint being triggered in response to a second match between the second data and the third data based on the second comparison.

[0267] Example 26 includes the at least one non-transitory computer-readable medium of Example 22, wherein the instructions, when executed, cause the first processor circuit to: identify a breakpoint to be triggered based on a first address in a memory associated with the data output, write the first address to a configuration register of a core of the first processor circuit or the second processor circuit, identify a second address in the memory at which the data output is written in response to executing the executable code, and perform a comparison of the first address and the second address, the triggering of the breakpoint being responsive to a match between the first address and the second address.

[0268] Example 27 includes at least one non-transitory computer-readable medium of Example 22, wherein the data input is a first data input, and the instructions, when executed, cause the first processor circuit to: in response to triggering of a breakpoint, obtain a control signal indicating execution of an increment operation of the executable code, the increment operation comprising at least one of a read operation to read a first value, a write operation to write a second value, or a calculation operation to determine a third value based on the second data input; and output at least one of the first value, the second value, or the third value.

[0269] Example 28 includes at least one non-transitory computer-readable medium of Example 22, wherein the instructions, when executed, cause the first processor circuit to: in response to triggering of a breakpoint, perform at least one of: adjust a first value of a data input, adjust a second value of a first register of a core of the first processor circuit or the second processor circuit, or adjust a third value of a second register of a debug circuit; and resume execution of the executable code based on at least one of the first value, the second value, or the third value.

[0270] Example 29 includes an apparatus for debugging a hardware accelerator, the apparatus comprising: a first interface circuit for obtaining a machine learning model; and a processor circuit comprising one or more of the following: at least one of a central processing unit, a graphics processing unit, or a digital signal processor, the at least one of the central processing unit, the graphics processing unit, or the digital signal processor having a control circuit for controlling data movement within the processor circuit, arithmetic and logic circuits for performing one or more first operations corresponding to an instruction, and one or more registers for storing results of the one or more first operations, instructions in the apparatus; a field programmable gate array (FPGA), the FPGA comprising logic gate circuits, a plurality of configurable interconnects, and a memory circuit, the logic gate circuit and the interconnection perform one or more second operations, the storage circuit is used to store results of the one or more second operations; or an application-specific integrated circuit (ASIC) including a logic gate circuit to perform one or more third operations; the processor circuit is used to perform at least one of the first operation, the second operation or the third operation to instantiate: a core circuit to execute executable code to generate a data output based on a data input, the executable code being based on a machine learning model; a second interface circuit to receive at least one of the data input or the data output; a multiplexer circuit to select the second interface circuit; and a shift register to output at least one of the data input or the data output in response to triggering of a breakpoint associated with the execution of the executable code.

[0271] Example 30 includes the apparatus of Example 29, wherein the second interface circuit is to receive data input from the memory, and the processor circuit is to perform at least one of the first operation, the second operation, or the third operation to instantiate: a buffer, in response to the breakpoint not being triggered, receiving the data input from the second interface circuit, and outputting the data input to the core circuit.

[0272] Example 31 includes the apparatus of Example 29, wherein the second interface circuit is to receive the data output from the core circuit, and the processor circuit is to perform at least one of the first operation, the second operation, or the third operation to instantiate: a buffer, in response to the breakpoint not being triggered, receiving the data output from the second interface circuit, and outputting the data output to the memory.

[0273] Example 32 includes the apparatus of Example 29, wherein the second interface circuit is to receive data input from a memory, and the processor circuit is to perform at least one of the first operation, the second operation, or the third operation to instantiate: a buffer to receive data input from the second interface circuit and output the data input to the core circuit in response to the breakpoint not being triggered; and a debug circuit to output at least one of the data input or the data output in response to the breakpoint being triggered, or to output the data output to the memory in response to the breakpoint not being triggered.

[0274] Example 33 includes the apparatus of Example 32, wherein the debug circuit is included in the core circuit.

[0275] Example 34 includes the apparatus of Example 29, wherein at least one of the data input or the data output comprises a first value, and the processor circuit is to perform at least one of a first operation, a second operation, or a third operation to instantiate: a configuration register to store a second value corresponding to the breakpoint; and a comparator circuit to compare the first value and a second value and, in response to a match between the first value and the second value based on the comparison, to instruct the second interface circuit to provide at least one of the data input or the data output to the multiplexer circuit, the breakpoint being triggered in response to the match.

[0276] Example 35 includes the apparatus of Example 29, wherein the core circuit is a first core circuit, and the processor circuit is to perform at least one of a first operation, a second operation, or a third operation to instantiate: a counter circuit to output a first value to instruct the multiplexer circuit to select an output of the second interface circuit, and to output a second value to instruct the multiplexer circuit to select an output of a third interface circuit associated with the second core circuit.

[0277] Example 36 includes a method for debugging a hardware accelerator, the method comprising generating breakpoints associated with a machine learning model, compiling executable code based on the machine learning model or at least one of the breakpoints, the executable code to be executed by an accelerator circuit to generate a data output based on a data input, triggering a breakpoint in response to execution of the executable code to halt execution of the executable code, and outputting at least one of the data input, the data output, or the breakpoint using debug circuitry included in the accelerator circuit.

[0278] Example 37 includes the method of Example 36, further comprising identifying a breakpoint to be triggered on a per-workload basis, inserting the breakpoint into executable code to be called on a per-workload basis, in response to execution of the executable code by a first core of the accelerator circuit, stopping execution of the executable code by the first core when the breakpoint is triggered by the first core, and in response to execution of the executable code by a second core of the accelerator circuit, stopping execution of the executable code by the second core when the breakpoint is triggered by the second core.

[0279] Example 38 includes the method of Example 36, further comprising identifying a breakpoint to be triggered on a per-core basis, identifying a breakpoint to be written to a first configuration register of a first core of the accelerator circuit rather than a second configuration register of a second core of the accelerator circuit, and writing the breakpoint to the first configuration register, wherein triggering of the breakpoint is used to stop execution of the executable code by the first core while the second core continues execution of the executable code.

[0280] Example 39 includes the method of Example 36, wherein the data input includes first data, the data output includes second data, and further includes identifying a breakpoint to be triggered based on third data, writing the third data into a configuration register of a core of the accelerator circuit, performing a first comparison of the first data and the third data, the breakpoint being triggered in response to a first match between the first data and the third data based on the first comparison, and performing a second comparison of the second data and the third data, the breakpoint being triggered in response to a second match between the second data and the third data based on the second comparison.

[0281] Example 40 includes the method of Example 36, further comprising identifying a breakpoint to be triggered based on a first address in a memory associated with the data output, writing the first address to a configuration register of a core of the accelerator circuit, identifying a second address in the memory at which the data output is written in response to executing the executable code, and performing a comparison of the first address and the second address, the triggering of the breakpoint being responsive to a match between the first address and the second address.

[0282] Example 41 includes the method of Example 36, wherein the data input is a first data input, and further includes: in response to triggering of a breakpoint, obtaining a control signal indicating execution of an increment operation of the executable code, the increment operation including at least one of a read operation for reading a first value, a write operation for writing a second value, or a calculation operation for determining a third value based on the second data input, and outputting at least one of the first value, the second value, or the third value.

[0283] Example 42 includes the method of Example 36, further comprising: in response to triggering of a breakpoint, performing at least one of the following: adjusting a first value of a data input, adjusting a second value of a first register of a core of an accelerator circuit, or adjusting a third value of a second register of a debug circuit; and resuming execution of the executable code based on at least one of the first value, the second value, or the third value.

[0284] The following claims are hereby incorporated by reference into this detailed description. Although certain example systems, methods, apparatus, and articles of manufacture have been disclosed herein, the scope of coverage of this patent is not limited thereto. Rather, this patent covers all systems, methods, apparatus, and articles of manufacture that fully fall within the scope of the claims of this patent.

Claims

1. A computing system comprising: A core configured to execute one or more workloads during execution of a neural network; and Debug module, which is used to: receiving a breakpoint configuration signal, the breakpoint configuration signal indicating a debug event associated with execution of the neural network, compiling the neural network based on the breakpoint configuration signal to generate a compiled neural network, providing the compiled neural network to the core, receiving output tensors generated by executing the one or more workloads by the core using the compiled neural network, and An error associated with the neural network is detected based on the output tensor.

2. The computing system of claim 1 , further comprising one or more other cores, wherein the core and the one or more other cores are configured to execute a plurality of workloads in parallel including the one or more workloads in the execution of the neural network. 3 . The computing system of claim 1 , wherein the debugging module is further configured to transfer the output tensor to a memory. 4 . The computing system of claim 3 , wherein the output tensor is generated by the core using input data, wherein the debug module is further configured to transfer the input data to the memory. The computing system of claim 4 , wherein the debugging module is further configured to detect the error based on the input data.

6. The computing system of claim 1 , wherein the debug module is further configured to stop execution of the neural network after detecting the error.

7. The computing system of claim 6, wherein the debug event is specific to a workload during execution of the neural network, and the debug module is configured to stop execution of the neural network by stopping the workload.

8. A method comprising: receiving a breakpoint configuration signal, the breakpoint configuration signal indicating a debug event associated with execution of the neural network; compiling the neural network based on the breakpoint configuration signal to generate a compiled neural network; generating an output tensor by executing one or more workloads during execution of the neural network using a core of the compiled neural network; as well as An error associated with the neural network is detected based on the output tensor.

9. The method of claim 8, wherein a plurality of workloads including the one or more workloads in the execution of the neural network are executed in parallel by the core and one or more other cores.

10. The method of claim 8, further comprising transferring the output tensor to a memory.

11. The method of claim 10, wherein the output tensor is generated from input data, wherein the method further comprises transferring the input data to the memory. The method of claim 11 , wherein detecting the error comprises detecting the error based on the input data.

13. The method of claim 8, further comprising stopping execution of the neural network after detecting the error.

14. The method of claim 13, wherein the debug event is specific to a workload during execution of the neural network, and stopping execution of the neural network comprises stopping the workload.

15. One or more computer-readable media storing instructions that, when executed by one or more processors, cause the one or more processors to perform the method of any one of claims 8 to 14.

16. A computing device comprising means for performing the method according to any one of claims 8 to 14.

17. A computer program product comprising instructions which, when executed by one or more processors, cause the one or more processors to perform the method according to any one of claims 8 to 14.