Method and apparatus for performing translation operations using direct memory access
By performing conversion operations in memory, using the Direct Memory Access (DMA) platform and memory conversion circuit module, the problem of neural network dependence on CPU in conversion operations is solved, improving efficiency and freeing up computing resources.
Patent Information
- Application Number
- CN202280101765.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-12
- Publication Date
- 2025-06-20
AI Technical Summary
When performing conversion operations, neural networks need to frequently pass information with the CPU, resulting in additional clock cycles, reducing efficiency and occupying computing resources.
Through the Direct Memory Access (DMA) platform, conversion operations are performed directly in memory, and data type conversion of source tensors is used to convert data to eliminate dependence on the CPU.
Reduces the dependence on the CPU in the conversion operation, reduces the additional clock cycle, improves the efficiency of the neural network, and frees up the use of computing resources.
Smart Images

Figure CN120188152A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to memory operations in computing devices and, more particularly, to methods and apparatus for performing conversion operations using direct memory access. Background Art
[0002] Conversion operations are applied to neural networks (e.g., machine learning models, artificial intelligence, etc.) to perform element-wise conversions on input tensors to convert them to a defined destination data type. The conversion operations are generally implemented at runtime (e.g., during the computer program phase of running or executing a computer program on a computer system) via instructions issued by a central processing unit (CPU). Brief Description of the Drawings
[0003] Figure 1 is a schematic illustration of an example direct memory access platform.
[0004] Figure 2 is a block diagram of an example memory conversion circuit module.
[0005] Figure 3 is a flow diagram representing example machine-readable instructions and / or example operations executable by an example processor circuit module to implement Figure 1 the example direct memory access platform.
[0006] Figure 4 is a flow diagram representing example machine-readable instructions and / or example operations executable by an example processor circuit module to implement Figure 2 the example memory conversion circuit module.
[0007] Figure 5 is a flow diagram representing example machine-readable instructions and / or example operations executable by an example processor circuit module to implement an example data type conversion circuit module as represented in Figure 2 the example data type conversion circuit module.
[0008] Figure 6 is a block diagram of an example processing platform including a processor circuit module configured to execute Figure 3 the example machine-readable instructions and / or example operations to implement Figure 1 the example direct memory access platform and Figure 2 the example memory conversion circuit module.
[0009] Figure 7 is Figure 6 a block diagram of an example implementation of the processor circuit module.
[0010] Figure 8 is Figure 6 a block diagram of another example implementation of the processor circuit module.
[0011] Figure 9 is a block diagram of an example software distribution platform (e.g., one or more servers) for distributing software (e.g., software corresponding to Figure 3 , Figure 4 and / or Figure 5 example machine-readable instructions) to client devices associated with end users and / or consumers (e.g., for licensing, selling, and / or using), retailers (e.g., for selling, reselling, licensing, and / or sublicensing), and / or original equipment manufacturers (OEMs) (e.g., for inclusion in products to be distributed to, e.g., retailers and / or other end users such as direct purchase customers).
[0012] In general, the same reference numerals will be used throughout the (one or more) figures and the accompanying written description to refer to the same or like parts. The figures are not drawn to scale.
[0013] Unless otherwise specifically stated, descriptors such as “first,” “second,” “third,” etc., used herein do not imply any order of precedence, physical order, arrangement in a list, and / or ranking in any way, but are merely used as labels and / or arbitrary names to distinguish elements for ease of understanding the disclosed examples. In some examples, the descriptor “first” may be used in the detailed description to refer to an element, while a different descriptor such as “second” or “third” may be used in the claims to refer to the same element. In such cases, it should be understood that such descriptors are merely used to unambiguously identify those elements that may, for example, otherwise share the same name.
[0014] As used herein, the phrase “in communication” (including its variants) encompasses direct communication and / or indirect communication through one or more intermediate components, and does not require direct physical (e.g., wired) communication and / or continuous communication, but additionally includes selective communication at periodic intervals, scheduled intervals, aperiodic intervals, and / or one-time events.
[0015] As used herein, "processor circuitry" is defined to include: (i) one or more special-purpose electrical circuits configured to perform one or more specific operations and including one or more semiconductor-based logic devices (e.g., electrical hardware implemented by one or more transistors); and / or (ii) one or more general-purpose semiconductor-based electrical circuits programmable with instructions to perform specific operations and including one or more semiconductor-based logic devices (e.g., electrical hardware implemented by one or more transistors). Examples of processor circuitry include programmable microprocessors, field programmable gate arrays (FPGAs) that can instantiate instructions, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), XPUs, or microcontrollers and integrated circuits such as application specific integrated circuits (ASICs). For example, an XPU can be implemented by a heterogeneous computing system that includes multiple types of processor circuitry (e.g., one or more FPGAs, one or more CPUs, one or more GPUs, one or more DSPs, etc., and / or combinations thereof) and one or more application programming interfaces (APIs) that can assign one or more computing tasks to whichever or which of the multiple types of processor circuitry is most suitable for performing the one or more computing tasks. Detailed Description
[0016] Neural networks typically perform transformation operations by issuing instructions from a CPU to retrieve a source tensor (e.g., the tensor on which the transformation operation is to be performed) from a memory (e.g., a connection matrix X SRAM memory (CMX)) and sending the source tensor to a kernel (e.g., the main software layer between an operating system (OS) and underlying computer hardware such as a CPU, memory, etc.) to perform the transformation operation on the source tensor. Once the source tensor has been transformed into a destination tensor by performing the transformation operation on the source tensor, the destination tensor is sent back to the memory.
[0017] During the use of a separate kernel to perform transformation operations, neural networks suffer from additional clock cycles required to perform the information transfer to and from the kernel. During large-scale transformation operations (e.g., transforming large amounts of data), these clock cycles escalate and pose a problem for the efficiency of the neural network and utilize computing resources that could otherwise be used elsewhere in the neural network. Accordingly, there is a need to optimize transformation operations in neural networks to eliminate the need for additional kernel operations.
[0018] Figure 1FIG. 0 is a schematic illustration of an example direct memory access platform 100 that performs a conversion operation via direct memory access (DMA). The example direct memory access platform 100 includes a memory 110, a source memory circuit module 120, a memory conversion circuit module 130, and a destination memory circuit module 140.
[0019] Figure 1 The source memory circuit module 120, the memory conversion circuit module 130, and the destination memory circuit module 140 of FIG. 0 may be instantiated (e.g., created, made to exist for any length of time, embodied, implemented, etc.) by a processor circuit module (such as a central processing unit that executes instructions). Additionally or alternatively, Figure 1 The source memory circuit module 120, the memory conversion circuit module 130, and the destination memory circuit module 140 of FIG. 0 may be instantiated (e.g., created, made to exist for any length of time, embodied, implemented, etc.) by an ASIC or FPGA configured to perform operations corresponding to the instructions. It should be understood that Figure 1 Some or all of the circuit modules of FIG. 0 may thus be instantiated at the same or different times. Some or all of the circuit modules may be instantiated, for example, in one or more threads that execute in parallel on hardware and / or serially on hardware. Additionally, in some examples, Figure 1 Some or all of the circuit modules of FIG. 0 may be implemented by a microprocessor circuit module that executes instructions to implement one or more virtual machines and / or containers.
[0020] The memory 110 communicates directly with the source memory circuit module 120 to send a source tensor (e.g., source data points) to the source memory circuit module 120. Additionally, the memory 110 communicates with the destination memory circuit module 140 to receive a destination tensor (e.g., the source tensor that has been converted to a different data type) from the destination memory circuit module 140. In the examples disclosed herein, the memory 110 may be any one or more of a volatile memory (e.g., any type of random access memory (RAM), etc.) or a non-volatile memory (e.g., electrically erasable programmable read-only memory (EEPROM), flash memory, HDD, SSD, etc.). In some examples, the direct memory access platform 100 is placed directly on a memory device that stores the memory 110 (e.g., on a processor placed on the memory device).
[0021] The source memory circuit module 120 obtains a source tensor from the memory 110. In some examples, the source memory circuit module 120 is instantiated by a processor circuit module that executes source memory instructions and / or is configured to perform operations (such as Figure 3 those represented by the flowchart of FIG. 0).
[0022] In some examples, direct memory access platform 100 includes means for obtaining source tensors from memory 110. For example, the means for obtaining may be implemented by source memory circuit module 120. In some examples, source memory circuit module 120 may be implemented by, for example, Figure 6 For example, source memory circuit module 120 may be implemented by executing machine executable instructions (such as at least Figure 3 Those instructions implemented by block 310 of Figure 7 In some examples, the source memory circuit module 120 may be instantiated by a hardware logic circuit module that may be configured to perform operations corresponding to machine-readable instructions. Figure 8 The source memory circuit module 120 may be implemented by an FPGA circuit module 800, an ASIC, or an XPU. Additionally or alternatively, the source memory circuit module 120 may be instantiated by any other combination of hardware, software, and / or firmware. For example, the source memory circuit module 120 may be implemented by at least one or more hardware circuits (e.g., a processor circuit module, a discrete and / or integrated analog and / or digital circuit module, an FPGA, an ASIC, an XPU, a comparator, an operational amplifier (op-amp), a logic circuit, etc.), and the hardware circuit is configured to execute some or all of the machine-readable instructions and / or execute some or all of the operations corresponding to the machine-readable instructions without executing software or firmware, but other structures are equally appropriate.
[0023] Memory conversion circuit module 130 performs data conversion on the source tensor to obtain the destination tensor. In some examples, memory conversion circuit module 130 is configured to execute memory conversion instructions and / or to perform operations such as Figure 3 , Figure 4 and / or Figure 5 The processor circuit module that instantiates the operations (those operations represented by the flowcharts of FIG. 1 ) is used to execute the operations.
[0024] In some examples, the direct memory access platform 100 includes a component for performing data conversion on the source tensor to obtain the destination tensor. For example, the component for performing can be implemented by the memory conversion circuit module 130. In some examples, the memory conversion circuit module 130 can be implemented by, for example, Figure 6 For example, the memory conversion circuit module 130 may be implemented by executing machine executable instructions (such as at least Figure 3 Frame 320, Figure 4 410, 420, 430 and 440 and Figure 5those instructions implemented by the frames 510, 520, 530, 540, and 550) of Figure 7 instantiated by the example microprocessor 700. In some examples, the memory translation circuit module 130 may be instantiated by a hardware logic circuit module, which may be configured to perform operations corresponding to machine-readable instructions Figure 8 implemented by the FPGA circuit 800, ASIC, or XPU. Additionally or alternatively, the memory translation circuit module 130 may be instantiated by any other combination of hardware, software, and / or firmware. For example, the memory translation circuit module 130 may be implemented by at least one or more hardware circuits (e.g., processor circuit modules, discrete and / or integrated analog and / or digital circuit modules, FPGAs, ASICs, XPUs, comparators, operational amplifiers (op-amps), logic circuits, etc.), which are configured to execute some or all of the machine-readable instructions and / or perform some or all of the operations corresponding to the machine-readable instructions without executing software or firmware, but other architectures are equally suitable.
[0025] The destination memory circuit module 140 applies the destination tensor to the memory 110. In some examples, the destination memory circuit module 140 is instantiated by a processor circuit module that executes destination memory instructions and / or is configured to perform operations (such as Figure 3 those operations represented by the flowchart of
[0026] In some examples, the direct memory access platform 100 includes components for applying the destination tensor to the memory 110. For example, the components for application may be implemented by the destination memory circuit module 140. In some examples, the destination memory circuit module 140 may be instantiated by a processor circuit module such as Figure 6 the example processor circuit module 612. For example, the destination memory circuit module 140 may be instantiated by a microprocessor that executes machine-executable instructions (such as at least Figure 3 those instructions implemented by the frame 330 of Figure 7 the example microprocessor 700. In some examples, the destination memory circuit module 140 may be instantiated by a hardware logic circuit module, which may be configured to perform operations corresponding to machine-readable instructions Figure 8implemented by the FPGA circuit 800, ASIC, or XPU. Additionally or alternatively, the destination memory circuit module 140 may be instantiated by any other combination of hardware, software, and / or firmware. For example, the destination memory circuit module 140 may be implemented by at least one or more hardware circuits (e.g., a processor circuit module, discrete and / or integrated analog and / or digital circuit modules, FPGA, ASIC, XPU, comparator, operational amplifier (op-amp), logic circuit, etc.), which are configured to execute some or all of the machine-readable instructions and / or perform some or all of the operations corresponding to the machine-readable instructions without executing software or firmware, but other architectures are equally suitable.
[0027] Although Figure 1 shows an example manner of implementing Figure 1 the direct memory access platform 100, Figure 1 one or more of the elements, processes, and / or devices shown in Figure 1 may be combined, split, rearranged, omitted, eliminated, and / or implemented in any other way. Additionally, Figure 1 the example source memory circuit module 120, example memory conversion circuit module 130, example destination memory circuit module 140, and / or (more generally) the example direct memory access platform 100 may be implemented by hardware only, or by hardware in combination with software and / or firmware. Thus, for example, any one of the example source memory circuit module 120, example memory conversion circuit module 130, example destination memory circuit module 140, and / or (more generally) the example direct memory access platform 100 may be implemented by a processor circuit module, (one or more) analog circuits, (one or more) digital circuits, (one or more) logic circuits, (one or more) programmable processors, (one or more) programmable microcontrollers, (one or more) graphics processing units (GPUs), (one or more) digital signal processors (DSPs), (one or more) application-specific integrated circuits (ASICs), (one or more) programmable logic devices (PLDs), and / or (one or more) field-programmable logic devices (FPLDs) such as field-programmable gate arrays (FPGAs). Additionally, Figure 1 the example direct memory access platform 100 may include one or more elements, processes, and / or devices that supplement or replace Figure 1 those shown in
[0028] Figure 2 is a block diagram of an example memory conversion circuit module 130 that uses direct memory access (DMA) to implement the conversion operation. Figure 1 ofFigure 2 The memory conversion circuit module 130 of may be instantiated (e.g., created, made to exist for any length of time, embodied, implemented, etc.) by a processor circuit module such as a central processing unit that executes instructions. Additionally or alternatively, Figure 2 The example memory conversion circuit module 130 of may be instantiated (e.g., created, made to exist for any length of time, embodied, implemented, etc.) by an ASIC or FPGA configured to perform operations corresponding to the instructions. It should be understood that Figure 2 Some or all of the circuit modules of may thus be instantiated at the same or different times. Some or all of the circuit modules may be instantiated, for example, in one or more threads that execute in parallel on hardware and / or serially on hardware. Additionally, in some examples, Figure 2 Some or all of the circuit modules of may be implemented by a microprocessor circuit module that executes instructions to implement one or more virtual machines and / or containers.
[0029] Figure 2 The memory conversion circuit module 130 of includes an input recognition circuit module 220, an output recognition circuit module 230, a data type conversion circuit module 240, and an output circuit module 250.
[0030] The input recognition circuit module 220 recognizes the source tensor data type. In the examples disclosed herein, the input recognition circuit module 220 receives the source tensor by communicating with the source memory circuit module 120 via the I / O interface 210. In some examples, the input recognition circuit module 220 is instantiated by a processor circuit module that executes input recognition instructions and / or is configured to perform operations (such as Figure 4 those represented by the flowchart of ).
[0031] In some examples, the memory conversion circuit module 130 includes components for recognizing the data type of the source tensor. For example, the components for recognition may be implemented by the input recognition circuit module 220. In some examples, the input recognition circuit module 220 may be instantiated by a processor circuit module such as Figure 6 the example processor circuit module 612 of. For example, the input recognition circuit module 220 may be instantiated by a microprocessor that executes machine-executable instructions (such as at least Figure 4 those implemented by block 410 of ). Figure 7 the example microprocessor 700 of. In some examples, the input recognition circuit module 220 may be instantiated by a hardware logic circuit module that may be configured to perform operations corresponding to machine-readable instructions. Figure 8implemented by an FPGA circuit 800, ASIC, or XPU. Additionally or alternatively, the input recognition circuit module 220 may be instantiated by any other combination of hardware, software, and / or firmware. For example, the input recognition circuit module 220 may be implemented by at least one or more hardware circuits (e.g., a processor circuit module, discrete and / or integrated analog and / or digital circuit modules, FPGA, ASIC, XPU, comparator, operational amplifier (op-amp), logic circuit, etc.), the hardware circuits being configured to execute some or all of the machine-readable instructions and / or perform some or all of the operations corresponding to the machine-readable instructions without executing software or firmware, but other configurations are equally suitable.
[0032] In some examples, the component for identifying the data type (of the source tensor) further identifies the shape of the source tensor.
[0033] The output recognition circuit module 230 determines the destination tensor data type to which the source tensor is to be converted. In the examples disclosed herein, the output recognition circuit module 230 determines the destination tensor data type by communicating with the source memory circuit module 120 via the I / O interface 210 to receive instructions regarding the data type to which the source tensor is to be converted. In some examples, the output recognition circuit module 230 is instantiated by a processor circuit module that executes output recognition instructions and / or is configured to perform operations (such as Figure 4 those represented by the flowchart of
[0034] In some examples, the memory conversion circuit module 130 includes a component for determining the destination tensor data type to which the source tensor is to be converted. For example, the component for determining may be implemented by the output recognition circuit module 230. In some examples, the output recognition circuit module 230 may be instantiated by a processor circuit module such as Figure 6 the example processor circuit module 612 of Figure 4 For example, the output recognition circuit module 230 may be instantiated by a Figure 7 example microprocessor 700 that executes machine-executable instructions (such as those implemented by at least Figure 8implemented by the FPGA circuit 800, ASIC, or XPU. Additionally or alternatively, the output recognition circuit module 230 may be instantiated by any other combination of hardware, software, and / or firmware. For example, the output recognition circuit module 230 may be implemented by at least one or more hardware circuits (such as, for example, a processor circuit module, discrete and / or integrated analog and / or digital circuit modules, FPGA, ASIC, XPU, comparator, operational amplifier (op-amp), logic circuits, etc.), the hardware circuits being configured to execute some or all of the machine-readable instructions and / or perform some or all of the operations corresponding to the machine-readable instructions without executing software or firmware, but other configurations are equally suitable.
[0035] The data type conversion circuit module 240 converts the source tensor to a destination tensor data type to obtain a destination tensor. In some examples, the data type conversion circuit module 240 is instantiated by a processor circuit module that executes data type conversion instructions and / or is configured to perform operations (such as Figure 4 and / or Figure 5 the operations represented by the flowchart of
[0036] In some examples, the memory conversion circuit module 130 includes components for converting the source tensor to a destination tensor data type. For example, the components for conversion may be implemented by the data type conversion circuit module 240. In some examples, the data type conversion circuit module 240 may be instantiated by a processor circuit module such as Figure 6 the example processor circuit module 612 of Figure 4 For example, the data type conversion circuit module 240 may be instantiated by an example microprocessor 700 that executes machine-executable instructions (such as at least Figure 5 the instructions implemented by block 430 of Figure 7 and Figure 8implemented by the FPGA circuit 800, ASIC, or XPU. Additionally or alternatively, the data type conversion circuit module 240 may be instantiated by any other combination of hardware, software, and / or firmware. For example, the data type conversion circuit module 240 may be implemented by at least one or more hardware circuits (e.g., a processor circuit module, discrete and / or integrated analog and / or digital circuit modules, FPGA, ASIC, XPU, comparator, operational amplifier (op-amp), logic circuit, etc.), the hardware circuits being configured to execute some or all of the machine-readable instructions and / or perform some or all of the operations corresponding to the machine-readable instructions without executing software or firmware, but other configurations are equally suitable.
[0037] In some examples, when the destination tensor data type is greater than the source tensor data type, the component for converting the source tensor (to the destination tensor data type) further selects values to fill in the blank spaces in the destination tensor. In other examples, when the destination tensor data type is less than the data type of the source tensor, the component for converting the source tensor further selects the data to be removed from the source tensor.
[0038] The output circuit module 250 outputs the destination tensor. In the examples disclosed herein, the output circuit module 250 outputs the destination tensor by communicating with the destination memory circuit module 140 via an I / O interface. In some examples, the output circuit module 250 is instantiated by a processor circuit module that executes output instructions and / or is configured to perform operations (such as Figure 4 those operations represented by the flowchart).
[0039] In some examples, the memory conversion circuit module 130 includes a component for outputting the destination tensor. For example, the component for outputting may be implemented by the output circuit module 250. In some examples, the output circuit module 250 may be instantiated by a processor circuit module such as Figure 6 the example processor circuit module 612. For example, the output circuit module 250 may be instantiated by a microprocessor that executes machine-executable instructions (such as at least Figure 4 those instructions implemented by block 440). Figure 7 the example microprocessor 700. In some examples, the output circuit module 250 may be instantiated by a hardware logic circuit module that may be configured to perform operations corresponding to the machine-readable instructions Figure 8implemented by the FPGA circuit 800, ASIC, or XPU. Additionally or alternatively, the output circuit module 250 may be instantiated by any other combination of hardware, software, and / or firmware. For example, the output circuit module 250 may be implemented by at least one or more hardware circuits (e.g., processor circuit modules, discrete and / or integrated analog and / or digital circuit modules, FPGAs, ASICs, XPU, comparators, operational amplifiers (op-amps), logic circuits, etc.), which are configured to execute some or all of the machine-readable instructions and / or perform some or all of the operations corresponding to the machine-readable instructions without executing software or firmware, but other configurations are equally suitable.
[0040] Although Figure 2 shows an example manner of implementing Figure 1 the memory conversion circuit module 130, Figure 2 one or more of the elements, processes, and / or devices shown in Figure 1 may be combined, split, rearranged, omitted, eliminated, and / or implemented in any other way. Additionally, Figure 1 the example memory conversion circuit module 130 of Figure 2 may include one or more elements, processes, and / or devices that supplement or replace
[0041] Figure 3 , Figure 4 and / or Figure 5FIG. shows a flow diagram representing example machine-readable instructions that, when executed, configure a processor circuit module to implement Figure 1 the direct memory access platform 100. The machine-readable instructions may be for a processor circuit module (such as the processor circuit module 612 shown in the example processor platform 600 discussed below in connection with Figure 6 and / or the example processor circuit modules discussed below in connection with Figure 7 and / or Figure 8 ). The program may be implemented in software stored on one or more non-transitory computer-readable storage media associated with the processor circuit module located in one or more hardware devices, such as a compact disc (CD), floppy disk, hard disk drive (HDD), solid state drive (SSD), digital versatile disc (DVD), Blu-ray disc, volatile memory (e.g., any type of random access memory (RAM), etc.), or non-volatile memory (e.g., electrically erasable programmable read-only memory (EEPROM), flash memory, HDD, SSD, etc.), but the entire program and / or portions thereof may alternatively be executed by one or more hardware devices different from the processor circuit module and / or implemented in firmware or dedicated hardware. The machine-readable instructions may be distributed across multiple hardware devices and / or executed by two or more hardware devices (e.g., server and client hardware devices). For example, the client hardware device may be implemented by an endpoint client hardware device (e.g., a hardware device associated with a user) or an intermediate client hardware device (e.g., a radio access network (RAN) gateway that may facilitate communication between the server and the endpoint client hardware device). Similarly, the non-transitory computer-readable storage media may include one or more media located in one or more hardware devices. Additionally, although reference is made to Figure 3 , Figure 4 and / or Figure 5The flowchart shown in depicts an example program, but many other methods of implementing the example direct memory access platform 100 may alternatively be used. For example, the order of execution of the blocks may be changed, and / or some of the blocks described may be changed, eliminated, or combined. Additionally or alternatively, any or all of the blocks may be implemented by one or more hardware circuits (e.g., processor circuit modules, discrete and / or integrated analog and / or digital circuit modules, FPGAs, ASICs, comparators, operational amplifiers (op-amps), logic circuits, etc.), which are configured to perform the corresponding operations without executing software or firmware. The processor circuit modules may be distributed at different network locations and / or be local to one or more hardware devices (e.g., a single-core processor (e.g., a single-core central processing unit (CPU)) in a single machine, a multi-core processor (e.g., a multi-core CPU, XPU, etc.), multiple processors distributed across multiple servers of a server rack, multiple processors distributed across one or more server racks, CPUs and / or FPGAs located in the same package (e.g., the same integrated circuit (IC) package) or in two or more separate enclosures, etc.).
[0042] The machine-readable instructions described herein may be stored in one or more formats such as a compressed format, an encrypted format, a fragmented format, a compiled format, an executable format, a packaged format, etc. The machine-readable instructions described herein may be stored as data or data structures (e.g., as instruction portions, code, code representations, etc.) that can be used to create, manufacture, and / or generate machine-executable instructions. For example, the machine-readable instructions may be fragmented and stored on one or more storage devices and / or computing devices (e.g., servers) located at the same or different locations of a network or network collection (e.g., in the cloud, in edge devices, etc.). The machine-readable instructions may need to be installed, modified, adapted, updated, combined, supplemented, configured, decrypted, decompressed, unpacked, distributed, reassigned, compiled, etc. in order to be directly readable, interpretable, and / or executable by a computing device and / or other machine. For example, the machine-readable instructions may be stored as multiple parts, each of which is compressed, encrypted, and / or stored on a separate computing device, where the parts form a set of machine-executable instructions that implement one or more operations when decrypted, decompressed, and / or combined, and these instructions may together form a program, such as the program described herein.
[0043] In another example, the machine-readable instructions may be stored in a state in which they are readable by a processor circuit, but require the addition of a library (e.g., a dynamic link library (DLL)), a software development kit (SDK), an application programming interface (API), etc., in order to execute these machine-readable instructions on a particular computing device or other device. In another example, the machine-readable instructions may need to be configured (e.g., store settings, input data, record network addresses, etc.) before all or part of the machine-readable instructions and / or the corresponding program(s) can be executed. Thus, as used herein, a machine-readable medium may include machine-readable instructions and / or program(s), regardless of the particular format or state of the machine-readable instructions and / or program(s) when they are stored or otherwise at rest or in transit.
[0044] The machine-readable instructions described herein may be represented by any past, present, or future instruction language, scripting language, programming language, etc. For example, the machine-readable instructions may be represented using any of the following languages: C, C++, Java, C#, Perl, Python, JavaScript, HyperText Markup Language (HTML), Structured Query Language (SQL), Swift, etc.
[0045] As mentioned above, Figure 3 、 Figure 4 and / or Figure 5Example operations may be implemented using executable instructions (e.g., computer and / or machine-readable instructions) stored on one or more non-transitory computer and / or machine-readable media, the readable media being such as optical storage devices, magnetic storage devices, HDDs, flash memories, read-only memories (ROMs), CDs, DVDs, caches, any type of RAM, registers, and / or any other storage device or storage disk, where information is stored for any duration (e.g., storing for an extended time period, storing permanently, storing for short instances, buffering temporarily, and / or caching information). As used herein, the terms non-transitory computer-readable media, non-transitory computer-readable storage media, non-transitory machine-readable media, and non-transitory machine-readable storage media are expressly defined to include any type of computer-readable storage device and / or storage disk, while excluding propagated signals and excluding transmission media. As used herein, the terms "computer-readable storage device" and "machine-readable storage device" are defined to include any physical (mechanical and / or electrical) structure that stores information, but excluding propagated signals and excluding transmission media. Examples of computer-readable storage devices and machine-readable storage devices include any type of random access memory, any type of read-only memory, solid-state memory, flash memory, optical disks, magnetic disks, disk drives, and / or redundant array of independent disks (RAID) systems. As used herein, the term "device" refers to a physical structure such as a mechanical and / or electrical device, hardware, and / or circuit module, which may or may not be configured by computer-readable instructions, machine-readable instructions, etc., and / or is manufactured to execute computer-readable instructions, machine-readable instructions, etc.
[0046] "Comprising" and "including" (and all of their forms and tenses) are used herein as open-ended terms. Thus, whenever a claim uses any form of "comprising" or "including" (e.g., comprises, includes, comprising, including, having, etc.) as a preamble or within any kind of claim recitation, it is to be understood that additional elements, terms, etc. may exist without falling outside the scope of the corresponding claim or recitation. As used herein, when the phrase "at least" is used as a transitional term in, for example, the preamble of a claim, it is open-ended in the same way that the terms "including" and "comprising" are open-ended. The term "and / or" when used in the form such as A, B, and / or C, refers to any combination or subset of A, B, C, such as (1) A alone, (2) B alone, (3) C alone, (4) A and B, (5) A and C, (6) B and C, or (7) A and B and C. As used herein in the context of describing a structure, component, item, object, and / or thing, the phrase "at least one of A and B" is intended to refer to an implementation that includes any one of the following: (1) at least one A, (2) at least one B, or (3) at least one A and at least one B. Similarly, as used herein in the context of describing a structure, component, item, object, and / or thing, the phrase "at least one of A or B" is intended to refer to an implementation that includes any one of the following: (1) at least one A, (2) at least one B, or (3) at least one A and at least one B. As used herein in the context of describing the performance or execution of a process, instruction, action, activity, and / or step, the phrase "at least one of A and B" is intended to refer to an implementation that includes any one of the following: (1) at least one A, (2) at least one B, or (3) at least one A and at least one B. Similarly, as used herein in the context of describing the performance or execution of a process, instruction, action, activity, and / or step, the phrase "at least one of A or B" is intended to refer to an implementation that includes any one of the following: (1) at least one A, (2) at least one B, or (3) at least one A and at least one B.
[0047] As used herein, references to the singular (e.g., "a", "an", "first", "second", etc.) do not exclude the plural. As used herein, the term "a" or "an" object refers to one or more of that object. The terms "a" (or "an"), "one or more", and "at least one" may be used interchangeably herein. In addition, although listed separately, multiple components, elements, or method acts may be implemented by, for example, the same entity or object. Further, although individual features may be included in different examples or claims, these may be combinable, and inclusion in different examples or claims does not imply that a combination of features is not feasible and / or advantageous.
[0048] Figure 3 is a flowchart of example machine-readable instructions and / or example operations that may be executed and / or instantiated by a processor circuit module to implement Figure 1 the direct memory access platform 100. Figure 3 An example direct memory access (DMA) process 300 of starts at block 310, where the source memory circuit module 120 obtains a source tensor from the memory 110.
[0049] Then, the memory conversion circuit module 130 performs a data conversion on the source tensor to obtain a destination tensor (block 320). In some examples, the data conversion performed by the memory conversion circuit module 130 converts a larger source tensor (e.g., unsigned 16-bit integer) into a smaller destination tensor (e.g., unsigned 8-bit integer). In other examples, the memory conversion circuit module 130 converts a smaller source tensor (e.g., unsigned 8-bit integer) into a larger destination tensor (e.g., unsigned 16-bit integer). In some examples, the memory conversion circuit module 130 may convert the source tensor into a destination tensor of the same bit value (e.g., unsigned integer to unsigned integer, signed integer to signed integer, etc.).
[0050] Once the memory conversion circuit module 130 has performed the data conversion, the destination memory circuit module 140 then applies the destination tensor to the memory 110 (block 330). In some examples, the destination memory circuit module 140 applies the destination tensor to the same space in the memory 110 where the source memory circuit module 120 obtained the source tensor. In other examples, the destination memory circuit module 140 applies the destination tensor to a different space in the memory 110 or to a completely different memory (e.g., a different memory source). Once the destination memory circuit module 140 has applied the destination tensor to the appropriate memory, the example DMA process 300 ends. The example DMA process 300 may be executed the number of times required to convert the number of source tensors required by a neural network.
[0051] Figure 4 is a flowchart of example machine-readable instructions and / or example operations that may be executed and / or instantiated by a processor circuit module to implement Figure 2 the example memory conversion circuit module 130. Figure 4An example data conversion operation 320 begins at block 410, where the input recognition circuit module 220 recognizes the source tensor data type and shape (e.g., the attributes of the source tensor). The data type can be any logical data type, such as unsigned 8-bit integer, etc. In some examples, the input recognition circuit module 220 recognizes the attributes of the source tensor via the I / O interface 210 to retrieve the source tensor from the source memory circuit module 120. In some examples, the shape of the source tensor represents the number of elements contained in the source tensor. The shape can be a matrix representation of the data within the source tensor, such as matrices of 2×1, 3×1, 4×1, etc. that logically represent a single row of data in memory.
[0052] Once the input recognition circuit module 220 has recognized the attributes of the source tensor, the output recognition circuit module 230 determines the destination data type to which the source tensor is to be converted (block 420). In some examples, the destination data type is a larger data type than the source tensor (e.g., the source tensor is expanded to fit the larger data type). In other examples, the destination data type is a smaller data type than the source tensor (e.g., the source tensor is contracted to fit the smaller data type).
[0053] Once the output recognition circuit module 230 has determined the destination data type, the data type conversion circuit module 240 converts the source tensor into the destination tensor data type to obtain the destination tensor (block 430). Further information regarding an example process for converting a source tensor into a destination tensor is provided below with reference to Figure 5 discloses.
[0054] Once the data type conversion circuit module 240 obtains the destination tensor, the output circuit module 250 outputs the destination tensor (block 440). In some examples, the output circuit module 250 outputs the destination tensor by passing the destination tensor to the destination memory circuit module 140 via the I / O interface 210. Once the output circuit module 250 has output the destination tensor, the example data conversion operation 320 ends. The example data conversion operation 320 can be repeated the required number of times to perform the conversion operation on the required amount of data.
[0055] Figure 5 is a flowchart representing example machine-readable instructions and / or example operations that can be executed and / or instantiated by a processor circuit module to implement Figure 2 the example data type conversion circuit module 240. Figure 5 An example destination tensor creation process 430 begins at block 510, where the data type conversion circuit module 240 recognizes the data type (or width) of the destination tensor as determined by the output recognition circuit module 230.
[0056] Once the data type conversion circuit module 240 identifies the data type of the destination tensor, the data type conversion circuit module 240 determines whether the source tensor can fit the data type of the destination tensor (block 520). In the examples disclosed herein, the data type of the destination tensor is different from the data type of the source tensor.
[0057] When the data type conversion circuit module 240 determines that the source tensor can fit the data type of the destination tensor (e.g., the result of block 520 is "yes"), the data type conversion circuit module 240 fills the blank spaces in the destination tensor data type with a predefined value (e.g., the extra space in the destination tensor data type once the source tensor has been converted) (block 530). In such an example, the data type of the destination tensor is any data type greater than the data type of the source tensor. In the examples disclosed herein, the predefined value can be logical 0 so as not to change the value of the source tensor. However, any other predefined value can be included in the blank spaces of the destination tensor.
[0058] When the data type conversion circuit module 240 determines that the source tensor cannot fit the data type of the destination tensor (e.g., the result of block 520 is "no"), the data type conversion circuit module 240 ignores / deletes the least significant bits of the source tensor (block 540). In such an example, the data type of the destination tensor is any data type less than the data type of the source tensor. In the examples disclosed herein, ignoring / deleting the least significant bits of the source tensor results in a loss of precision of the source tensor, but reduces the memory allocation for the source tensor on the memory 110. Such a reduction in memory allocation may be desirable for a given neural network in order to be able to allocate computing resources and memory elsewhere.
[0059] When the data type conversion circuit module 240 determines the appropriate way to manage the source tensor based on the data type of the destination tensor (e.g., block 530 or 540), the data type conversion circuit module 240 creates the destination tensor by mapping the source tensor to the data type of the destination tensor (block 550). Once the source tensor has been mapped to the data type of the destination tensor, the example destination tensor creation process 430 ends.
[0060] Figure 6 is a block diagram of an example processor platform 600 that is configured to execute and / or instantiate Figure 3 、 Figure 4 and / or Figure 5 machine-readable instructions and / or operations to implement Figure 1The direct memory access platform 100. The processor platform 600 can be, for example, a server, a personal computer, a workstation, a self-learning machine (e.g., a neural network), a mobile device (e.g., a cellular phone, a smart phone, a tablet such as an iPad TM ), a personal digital assistant (PDA), an Internet appliance, a game console, a personal video recorder, a set-top box, headphones (e.g., augmented reality (AR) headphones, virtual reality (VR) headphones, etc.) or other wearable devices, or any other type of computing device.
[0061] The illustrated example of the processor platform 600 includes processor circuitry 612. The illustrated example of the processor circuitry 612 is hardware. For example, the processor circuitry 612 can be implemented by one or more integrated circuits, logic circuits, FPGAs, microprocessors, CPUs, GPUs, DSPs, and / or microcontrollers from any desired family or manufacturer. The processor circuitry 612 can be implemented by one or more semiconductor (e.g., silicon-based) devices. In this example, the processor circuitry 612 implements the source memory circuitry 120, the memory conversion circuitry 130, the destination memory circuitry 140, the input recognition circuitry 220, the output recognition circuitry 230, the data type conversion circuitry 240, and the output circuitry 250.
[0062] The illustrated example of the processor circuitry 612 includes local memory 613 (e.g., a cache, registers, etc.). The illustrated example of the processor circuitry 612 communicates with the main memory (including volatile memory 614 and non-volatile memory 616) via a bus 618. The volatile memory 614 can be implemented by synchronous dynamic random access memory (SDRAM), dynamic random access memory (DRAM), dynamic random access memory and / or any other type of RAM device. The non-volatile memory 616 can be implemented by flash memory and / or any other desired type of memory device. Access to the illustrated example of the main memory 614, 616 is controlled by a memory controller 617.
[0063] The illustrated example of the processor platform 600 also includes interface circuitry 620. The interface circuitry 620 can be implemented by hardware in accordance with any type of interface standard, such as an Ethernet interface, a universal serial bus (USB) interface, interface, a near field communication (NFC) interface, a peripheral component interconnect (PCI) interface, and / or a rapid peripheral component interconnect (PCIe) interface.
[0064] In the illustrated example, one or more input devices 622 are connected to interface circuit module 620. The (one or more) input devices 622 allow a user to input data and / or commands to processor circuit module 612. The (one or more) input devices 622 may be implemented by, for example, audio sensors, microphones, cameras (still or video), keyboards, buttons, mice, touchscreens, trackpads, trackballs, and other pointing devices and / or voice recognition systems.
[0065] One or more output devices 624 are also connected to interface circuit module 620 of the illustrated example. The (one or more) output devices 624 may be implemented by, for example, display devices (such as light-emitting diodes (LEDs), organic light-emitting diodes (OLEDs), liquid crystal displays (LCDs), cathode ray tube (CRT) displays, in-plane switching (IPS) displays, touchscreens, etc.), printers, and / or speakers. Thus, interface circuit module 620 of the illustrated example generally includes a graphics driver card, a graphics driver chip, and / or a graphics processor circuit module (such as a GPU).
[0066] Interface circuit module 620 of the illustrated example also includes communication devices (such as transmitters, receivers, transceivers, modems, residential gateways, wireless access points, and / or network interfaces) to facilitate the exchange of data with external machines (such as any kind of computing device) via network 626. Communication may be carried out, for example, via Ethernet connections, digital subscriber line (DSL) connections, telephone line connections, coaxial cable systems, satellite systems, line-of-sight wireless systems, cellular telephone systems, optical connections, etc.
[0067] Processor platform 600 of the illustrated example also includes one or more mass storage devices 628 for storing software and / or data. Examples of such mass storage devices 628 include magnetic storage devices, optical storage devices, floppy disk drives, HDDs, CDs, Blu-ray disc drives, redundant arrays of independent disks (RAID) systems, solid-state storage devices (such as flash memory devices and / or SSDs), and DVD drives.
[0068] May be implemented by Figure 3 , Figure 4 and / or Figure 5 Machine-readable instructions 632 may be stored in mass storage device 628, volatile memory 614, non-volatile memory 616, and / or on a removable non-transitory computer-readable storage medium such as a CD or DVD.
[0069] Figure 7 is Figure 6 A block diagram of an example implementation of processor circuit module 612 of. In this example, Figure 6The processor circuit module 612 is implemented by the microprocessor 700. For example, the microprocessor 700 can be a general-purpose microprocessor (e.g., a general-purpose microprocessor circuit module). The microprocessor 700 executes Figure 3 , Figure 4 and / or Figure 5 of some or all of the machine-readable instructions of the flowchart to effectively instantiate the source memory circuit module 120, the memory conversion circuit module 130, and / or the destination memory circuit module 140 into logic circuits to perform operations corresponding to those machine-readable instructions. In some such examples, the source memory circuit module 120, the memory conversion circuit module 130, and / or the destination memory circuit module 140 are instantiated by the hardware circuits of the microprocessor 700 in combination with instructions. For example, the microprocessor 700 can be implemented by a multi-core hardware circuit module (such as a CPU, a DSP, a GPU, an XPU, etc.). Although it can include any number of example cores 702 (e.g., 1 core), the microprocessor 700 in this example is a multi-core semiconductor device including N cores. The cores 702 of the microprocessor 700 can operate independently or can cooperate to execute machine-readable instructions. For example, the machine code corresponding to a firmware program, an embedded software program, or a software program can be executed by one of these cores 702, or can be executed by multiple ones of these cores 702 at the same or different times. In some examples, the machine code corresponding to a firmware program, an embedded software program, or a software program is split into threads and executed in parallel by two or more of these cores 702. The software program can correspond to Figure 3 , Figure 4 and / or Figure 5 of some or all of the machine-readable instructions and / or operations represented by the flowchart.
[0070] The core 702 can communicate via the first example bus 704. In some examples, the first bus 704 can be implemented by a communication bus to effect communication associated with one or more of these cores 702. For example, the first bus 704 can be implemented by at least one of an Inter-Integrated Circuit (I2C) bus, a Serial Peripheral Interface (SPI) bus, a PCI bus, or a PCIe bus. Additionally or alternatively, the first bus 704 can be implemented by any other type of computing or electrical bus. The core 702 can obtain data, instructions, and / or signals from one or more external devices via the example interface circuit module 706. The core 702 can output data, instructions, and / or signals to one or more external devices via the interface circuit module 706. Although the core 702 of this example includes an example local memory 720 (e.g., a level 1 (L1) cache, which can be split into an L1 data cache and an L1 instruction cache), the microprocessor 700 also includes an example shared memory 710 (e.g., a level 2 (L2) cache) that can be shared by these cores for high-speed access to data and / or instructions. Data and / or instructions can be transferred (e.g., shared) by writing to and / or reading from the shared memory 710. The local memory 720 and the shared memory 710 of each of these cores 702 can be part of a hierarchy of storage devices including multiple levels of cache memory and main memory (e.g., Figure 6 the main memories 614, 616). Generally, higher-level memories in the hierarchy exhibit less access time and have smaller storage capacities compared to lower-level memories. Changes in the levels of the cache hierarchy are managed (e.g., coordinated) via a cache coherence policy.
[0071] Each core 702 may be referred to as a CPU, DSP, GPU, etc. or any other type of hardware circuit module. Each core 702 includes a control unit circuit module 714, an arithmetic and logic (AL) circuit module (sometimes referred to as an ALU) 716, a plurality of registers 718, a local memory 720, and a second example bus 722. Other structures may exist. For example, each core 702 may include a vector unit circuit module, a single instruction multiple data (SIMD) unit circuit module, a load / store unit (LSU) circuit module, a branch / jump unit circuit module, a floating point unit (FPU) circuit module, etc. The control unit circuit module 714 includes semiconductor-based circuitry configured to control (e.g., coordinate) data movement within the corresponding core 702. The AL circuit module 716 includes semiconductor-based circuitry configured to perform one or more mathematical and / or logical operations on data within the corresponding core 702. Some example AL circuit modules 716 perform integer-based operations. In other examples, the AL circuit module 716 also performs floating point operations. In still other examples, the AL circuit module 716 may include a first AL circuit module that performs integer-based operations and a second AL circuit module that performs floating point operations. In some examples, the AL circuit module 716 may be referred to as an arithmetic logic unit (ALU). The registers 718 are semiconductor-based structures for storing data and / or instructions such as the results of one or more operations performed by the AL circuit module 716 of the corresponding core 702. For example, the registers 718 may include one or more vector registers, one or more SIMD registers, one or more general purpose registers, one or more flag registers, one or more segment registers, one or more machine specific registers, one or more instruction pointer registers, one or more control registers, one or more debug registers, one or more memory management registers, one or more machine check registers, etc. The registers 718 may be arranged in banks as shown in Figure 7 . Alternatively, the registers 718 may be organized in any other arrangement, format, or structure, including being distributed throughout the core 702 to reduce access time. The second bus 722 may be implemented by at least one of an I2C bus, an SPI bus, a PCI bus, or a PCIe bus.
[0072] Each core 702 and / or (more generally) microprocessor 700 may include additional and / or alternative structures of those shown and described above. For example, there may be one or more clock circuits, one or more power supplies, one or more power gating, one or more cache coherence agents (CHA), one or more convergence / general mesh access points (CMS), one or more shifters (e.g., (one or more) barrel shifters), and / or other circuit modules. The microprocessor 700 is a semiconductor device in one or more integrated circuits (ICs) contained in one or more packages, which are manufactured to include many transistors interconnected to implement the above structures. The processor circuit module may include one or more accelerators and / or cooperate with one or more accelerators. In some examples, the accelerator is implemented by a logic circuit module to perform certain tasks more quickly and / or efficiently than a general-purpose processor can. Examples of accelerators include ASICs and FPGAs (such as those discussed herein). A GPU or other programmable device may also be an accelerator. The accelerator may be on-board the processor circuit module, in the same chip package as the processor circuit module, and / or in one or more packages separate from the processor circuit module.
[0073] Figure 8 is Figure 6 Another example implementation of the processor circuit module 612. In this example, the processor circuit module 612 is implemented by the FPGA circuit module 800. For example, the FPGA circuit module 800 may be implemented by an FPGA. The FPGA circuit module 800 can be used to perform operations that might otherwise be performed by Figure 7 the example microprocessor 700 executing corresponding machine-readable instructions. However, once configured, the FPGA circuit module 800 instantiates these machine-readable instructions in hardware and, therefore, can often perform these operations faster than can be achieved by operating on corresponding software executed by a general-purpose microprocessor.
[0074] More specifically, compared to the Figure 7 microprocessor 700 described above (which is programmable to execute Figure 3 , Figure 4 and / or Figure 5 some or all of the machine-readable instructions represented by the flowchart, but whose interconnections and logic circuits are fixed once manufactured), Figure 8 the example FPGA circuit module 800 of Figure 3 , Figure 4 and / or Figure 5Some or all of the machine-readable instructions represented by the flowchart instantiate an interconnected and logic circuit module. In particular, the FPGA circuit module 800 can be regarded as an array of logic gates, interconnections, and switches. The switches can be programmed to change the way in which the logic gates are interconnected via the interconnections, effectively forming one or more dedicated logic circuits (unless and until the FPGA circuit module 800 is reprogrammed). The configured logic circuit enables the logic gates to cooperate in different ways, thereby performing different operations on the data received by the input circuit module. Those operations can correspond to Figure 3 , Figure 4 and / or Figure 5 Some or all of the software represented by the flowchart. In this way, the FPGA circuit module 800 can be configured to Figure 3 , Figure 4 and / or Figure 5 Some or all of the machine-readable instructions of the flowchart are effectively instantiated as a dedicated logic circuit so as to perform operations corresponding to those software instructions in a dedicated manner similar to an ASIC. Therefore, the FPGA circuit module 800 can perform the same operations faster than a general-purpose microprocessor can perform operations corresponding to Figure 3 , Figure 4 and / or Figure 5 Some or all of the machine-readable instructions.
[0075] In the example of Figure 8 , the FPGA circuit module 800 is configured to be programmed (and / or reprogrammed one or more times) by an end user via a hardware description language (HDL) such as Verilog. Figure 8 The FPGA circuit module 800 of Figure 7implemented by the microprocessor 700. The FPGA circuit module 800 also includes an example logic gate circuit module 808, an array of multiple example configurable interconnections 810, and an example storage circuit module 812. The logic gate circuit module 808 and the configurable interconnection 810 are configurable to instantiate one or more operations corresponding to at least some of the machine-readable instructions corresponding to Figure 3 , Figure 4 and / or Figure 5 and / or other desired operation instances of the machine-readable instructions. Figure 8 The logic gate circuit module 808 shown in is fabricated in groups or blocks. Each block includes a semiconductor-based electrical structure that can be configured into a logic circuit. In some examples, the electrical structure includes logic gates (e.g., "AND" gates, "OR" gates, "NOR" gates, etc.) that provide basic building blocks for the logic circuit. There are electrically controlled switches (e.g., transistors) within each logic gate circuit module 808 to enable the configuration of the electrical structure and / or the logic gates to form a circuit that performs the desired operation. The logic gate circuit module 808 may include other electrical structures, such as look-up tables (LUTs), registers (e.g., flip-flops or latches), multiplexers, etc.
[0076] The example configurable interconnection 810 shown in is a conductive path, trace, via, etc., which may include electrically controlled switches (e.g., transistors), and the state of the electrically controlled switches can be changed by programming (e.g., using an HDL instruction language) to activate or deactivate one or more connections between one or more of the logic gate circuit modules 808, thereby programming the desired logic circuit.
[0077] The example storage circuit module 812 shown in is configured to store the result(s) of one or more of the operations performed by the corresponding logic gates. The storage circuit module 812 can be implemented by registers, etc. In the example shown, the storage circuit module 812 is distributed among the logic gate circuit modules 808 for easy access and to improve execution speed.
[0078] Figure 8The example FPGA circuit module 800 also includes an example dedicated operation circuit module 814. In this example, the dedicated operation circuit module 814 includes a special-purpose circuit module 816, and the special-purpose circuit module can be called to implement common functions to avoid the need to program those functions on-site. Examples of such special-purpose circuit modules 816 include memory (e.g., DRAM) controller circuit modules, PCIe controller circuit modules, clock circuit modules, transceiver circuit modules, memories, and multiplier-accumulator circuit modules. Other types of special-purpose circuit modules may exist. In some examples, the FPGA circuit module 800 may also include an example general-purpose programmable circuit module 818, such as an example CPU 820 and / or an example DSP 822. Additionally or alternatively, other general-purpose programmable circuit modules 818 may exist, such as GPUs, XPUs, etc., which can be programmed to perform other operations.
[0079] Although Figure 7 and Figure 8 show Figure 6 two example implementations of the processor circuit module 612, many other ways are contemplated. For example, as mentioned above, modern FPGA circuit modules may include on-board CPUs, such as one or more Figure 8 example CPUs 820. Thus, Figure 6 the processor circuit module 612 may alternatively be implemented by combining Figure 7 an example microprocessor 700 and Figure 8 an example FPGA circuit module 800. In some such hybrid examples, Figure 3 , Figure 4 and / or Figure 5 the first part of the machine-readable instructions represented by the flowchart of Figure 7 may be executed by one or more cores in the Figure 3 , Figure 4 and / or Figure 5 core 702, Figure 8 the second part of the machine-readable instructions represented by the flowchart of Figure 3 , Figure 4 and / or Figure 5 may be executed by the FPGA circuit module 800, and / or Figure 1 and / or Figure 2 the third part of the machine-readable instructions represented by the flowchart of Figure 3 , Figure 4 and / or Figure 5 may be executed by the ASIC. It should be understood that Figure 1 and / or Figure 2 some or all of the circuit modules in Figure 1 and / orFigure 2 some or all of the circuit modules in the circuit module.
[0080] In some examples, Figure 6 the processor circuit module 612 of can be in one or more packages. For example, Figure 7 the microprocessor 700 of and / or Figure 8 the FPGA circuit module 800 of can be in one or more packages. In some examples, the XPU can be implemented by the Figure 6 processor circuit module 612 that can be in one or more packages. For example, the XPU can include a CPU in one package, a DSP in another package, a GPU in yet another package, and an FPGA in still another package.
[0081] In Figure 9 shows a block diagram that shows an example software distribution platform 905 for distributing software (such as Figure 6 the example machine-readable instructions 632 of ) to hardware devices owned and / or operated by third parties. The example software distribution platform 905 can be implemented by any computer server, data facility, cloud service, etc. that can store software and transfer the software to other computing devices. The third party can be a customer of the entity that owns and / or operates the software distribution platform 905. For example, the entity that owns and / or operates the software distribution platform 905 can be a developer, seller, and / or licensor of software such as Figure 6 the example machine-readable instructions 632 of. The third party can be a consumer, user, retailer, OEM, etc. that purchases and / or licenses the software for use and / or resale and / or sublicense. In the example shown, the software distribution platform 905 includes one or more servers and one or more storage devices. As described above, the storage device stores the machine-readable instructions 632, and the machine-readable instructions 632 can correspond to Figure 3 , Figure 4 and / or Figure 5 the example machine-readable instructions of. One or more servers of the example software distribution platform 905 communicate with the example network 910, and the example network 910 can correspond to the Internet and / or any one or more of the example networks 626 described above. In some examples, as part of a commercial transaction, one or more servers respond to a request to transfer software to a requester. Payment for the delivery, sale, and / or license of the software can be handled by one or more servers of the software distribution platform and / or by a third-party payment entity. The server enables the purchaser and / or licensee to download the machine-readable instructions 632 from the software distribution platform 905. For example, it can correspond to Figure 3 , Figure 4 and / or Figure 5Software of example machine-readable instructions can be downloaded to example processor platform 600, which is to execute machine-readable instructions 632 to implement direct memory access platform 100. In some examples, one or more servers of software distribution platform 905 periodically supply, transmit, and / or enforce updates to software (e.g., Figure 6 example machine-readable instructions 632) to ensure that improvements, patches, updates, etc. are distributed and applied to the software at the end-user device.
[0082] From the foregoing, it will be appreciated that example systems, methods, devices, and articles for performing transformation operations using direct memory access have been disclosed. The disclosed systems, methods, devices, and articles improve the efficiency of using a computing device by eliminating the need for an external kernel to perform memory transformation operations. Accordingly, the disclosed systems, methods, devices, and articles are directed to one or more improvements in the operation of machines such as computers or other electronic and / or mechanical devices.
[0083] Example methods, devices, systems, and articles for performing transformation operations using direct memory access are disclosed herein. Further examples and combinations thereof include the following:
[0084] Example 1 includes a device that includes at least one memory, machine-readable instructions, and a processor circuit module for instantiating or executing at least one of the machine-readable instructions to: obtain a source tensor from the memory, the source tensor containing data and being obtained from the memory by direct memory access; perform a data transformation on the source tensor to obtain a destination tensor, the performance of the data transformation occurring by direct memory access; and apply the destination tensor to the memory, the destination tensor being applied to the memory by direct memory access.
[0085] Example 2 includes the device of Example 1, wherein the processor circuit module is for identifying a first data type corresponding to the source tensor.
[0086] Example 3 includes the device of Example 2, wherein the processor circuit module is for identifying the shape of the source tensor.
[0087] Example 4 includes the device of Example 2, wherein the processor circuit module is for determining a second data type corresponding to the destination tensor.
[0088] Example 5 includes the device of Example 4, wherein, to perform the data transformation, the processor circuit module is for converting the source tensor from the first data type to the second data type.
[0089] Example 6 includes the device of Example 5, wherein the processor circuit module is for selecting values to fill in blank spaces in the destination tensor.
[0090] Example 7 includes the apparatus of Example 5, wherein the processor circuit module is configured to select data to be removed from the source tensor.
[0091] Example 8 includes the apparatus of Example 1, wherein the processor circuit module is configured to output the destination tensor before applying the destination tensor to the memory.
[0092] Example 9 includes an apparatus for performing a conversion operation, the apparatus comprising: means for obtaining a source tensor from a memory, the source tensor containing data and being obtained from the memory by direct memory access; means for performing a data conversion on the source tensor to obtain a destination tensor; and means for applying the destination tensor to the memory, the destination tensor being applied to the memory by direct memory access; wherein the obtaining means, the performing means, and the applying means are implemented using direct memory access.
[0093] Example 10 includes the apparatus of Example 9, further comprising means for identifying a first data type corresponding to the source tensor.
[0094] Example 11 includes the apparatus of Example 10, wherein the means for identifying is configured to identify the shape of the source tensor.
[0095] Example 12 includes the apparatus of Example 10, further comprising means for determining a second data type corresponding to the destination tensor.
[0096] Example 13 includes the apparatus of Example 12, further comprising means for converting the source tensor from the first data type to the second data type.
[0097] Example 14 includes the apparatus of Example 13, wherein the means for converting is configured to select values to fill in blank spaces in the destination tensor.
[0098] Example 15 includes the apparatus of Example 13, wherein the means for converting is configured to select data to be removed from the source tensor.
[0099] Example 16 includes the apparatus of Example 9, further comprising means for outputting the destination tensor before applying the destination tensor to the memory.
[0100] Example 17 includes a non-transitory machine-readable storage medium containing instructions which, when executed, cause a processor circuit module to at least: obtain a source tensor from a memory, the source tensor containing data and being obtained from the memory by direct memory access; perform a data conversion on the source tensor to obtain a destination tensor, the data conversion being performed by direct memory access; and apply the destination tensor to the memory, the destination tensor being applied to the memory by direct memory access.
[0101] Example 18 includes the non-transitory machine-readable storage medium of Example 17, wherein, when executed, the instructions further cause the processor circuitry to identify a first data type corresponding to a source tensor.
[0102] Example 19 includes the non-transitory machine-readable storage medium of Example 18, wherein, when executed, the instructions further cause the processor circuitry to identify the shape of the source tensor.
[0103] Example 20 includes the non-transitory machine-readable storage medium of Example 18, wherein, when executed, the instructions further cause the processor circuitry to determine a second data type corresponding to a destination tensor.
[0104] Example 21 includes the non-transitory machine-readable storage medium of Example 20, wherein, when executed, the instructions further cause the processor circuitry to perform a data conversion by converting the source tensor from the first data type to the second data type.
[0105] Example 22 includes the non-transitory machine-readable storage medium of Example 21, wherein, when executed, the instructions further cause the processor circuitry to select values to fill in blank spaces in the destination tensor.
[0106] Example 23 includes the non-transitory machine-readable storage medium of Example 21, wherein, when executed, the instructions further cause the processor circuitry to select data to be removed from the source tensor.
[0107] Example 24 includes the non-transitory machine-readable storage medium of Example 17, wherein, when executed, the instructions further cause the processor circuitry to output the destination tensor before applying the destination tensor to a memory platform storage device.
[0108] Example 25 includes a method for performing a conversion operation, the method comprising: obtaining a source tensor from a memory, the source tensor containing data; performing a data conversion on the source tensor to obtain a destination tensor; and applying the destination tensor to the memory; wherein, obtaining, performing, and applying are all performed using direct memory access.
[0109] Example 26 includes the method of Example 25, further comprising: identifying a first data type corresponding to the source tensor.
[0110] Example 27 includes the method of Example 26, further comprising: identifying the shape of the source tensor.
[0111] Example 28 includes the method of Example 26, further comprising: determining a second data type corresponding to the destination tensor.
[0112] Example 29 includes the method of Example 28, further comprising: converting the source tensor from a first data type to a second data type.
[0113] Example 30 includes the method of Example 29, further comprising: selecting values to fill in the blank spaces in the destination tensor.
[0114] Example 31 includes the method of Example 29, further comprising: selecting data to be removed from the source tensor.
[0115] Example 32 includes the method of Example 25, further comprising: outputting the destination tensor before applying the destination tensor to the memory.
[0116] The following claims are hereby incorporated by reference into this detailed description. Although certain example systems, methods, devices, and articles have been disclosed herein, the scope of this patent is not limited thereto. Instead, this patent covers all systems, methods, devices, and articles that fall entirely within the scope of the claims of this patent.
Claims
1. An apparatus, comprising: At least one memory; Machine-readable instructions; And A processor circuit module for instantiating or executing at least one of the machine-readable instructions for: Obtaining a source tensor from the memory, the source tensor containing data and being obtained from the memory via direct memory access; Performing a data transformation on the source tensor to obtain a destination tensor, the execution of the data transformation occurring via direct memory access; And Applying the destination tensor to the memory, the destination tensor being applied to the memory via direct memory access.
2. The apparatus according to claim 1, wherein, The processor circuit module is for identifying a first data type corresponding to the source tensor.
3. The apparatus according to claim 2, wherein, The processor circuit module is for identifying the shape of the source tensor.
4. The apparatus according to claim 2, wherein, The processor circuit module is for determining a second data type corresponding to the destination tensor.
5. The apparatus according to claim 4, wherein, To perform the data transformation, the processor circuit module is for converting the source tensor from the first data type to the second data type.
6. The apparatus according to claim 5, wherein, The processor circuit module is for selecting values to fill in blank spaces in the destination tensor.
7. The apparatus according to claim 5, wherein, The processor circuit module is for selecting data to be removed from the source tensor.
8. The apparatus according to claim 1, wherein, The processor circuit module is for outputting the destination tensor before applying the destination tensor to the memory.
9. An apparatus for performing a conversion operation, comprising: A component for obtaining a source tensor from a memory, the source tensor containing data and being obtained from the memory via direct memory access; A component for performing a data transformation on the source tensor to obtain a destination tensor; And A component for applying the destination tensor to the memory, the destination tensor being applied to the memory via direct memory access; Wherein, the obtaining component, the performing component, and the applying component are implemented using direct memory access.
10. The apparatus according to claim 9, further comprising: A component for identifying a first data type corresponding to the source tensor.
11. The apparatus according to claim 10, wherein, The component for identifying is for identifying the shape of the source tensor.
12. The apparatus according to claim 10, further comprising: A component for determining a second data type corresponding to the destination tensor.
13. The apparatus according to claim 12, further comprising: A component for converting the source tensor from the first data type to the second data type.
14. The apparatus according to claim 13, wherein, The component for converting is for selecting values to fill in blank spaces in the destination tensor.
15. The apparatus according to claim 13, wherein, The component for converting is for selecting data to be removed from the source tensor.
16. The apparatus according to claim 9, further comprising: A component for outputting the destination tensor before applying the destination tensor to the memory.
17. A non - transitory machine - readable storage medium containing instructions which, when executed, cause a processor circuit module to at least: Obtain a source tensor from a memory, the source tensor containing data and being obtained from the memory by direct memory access; Perform a data transformation on the source tensor to obtain a destination tensor, the execution of the data transformation being achieved by direct memory access; and Apply the destination tensor to the memory, the destination tensor being applied to the memory by direct memory access.
18. The non - transitory machine - readable storage medium according to claim 17, wherein, When executed, the instructions further cause the processor circuit module to identify a first data type corresponding to the source tensor.
19. The non - transitory machine - readable storage medium according to claim 18, wherein, When executed, the instructions further cause the processor circuit module to identify the shape of the source tensor.
20. The non - transitory machine - readable storage medium according to claim 18, wherein, When executed, the instructions further cause the processor circuit module to determine a second data type corresponding to the destination tensor.
21. The non - transitory machine - readable storage medium according to claim 20, wherein, When executed, the instructions further cause the processor circuit module to perform the data transformation by converting the source tensor from the first data type to the second data type.
22. The non - transitory machine - readable storage medium according to claim 21, wherein, When the instruction is executed, it further causes the processor circuit module to select values to fill in the blank spaces in the destination tensor.
23. The non - transitory machine - readable storage medium according to claim 21, wherein, When the instruction is executed, it further causes the processor circuit module to select data to be removed from the source tensor.
24. The non - transitory machine - readable storage medium according to claim 17, wherein, When the instruction is executed, it further causes the processor circuit module to output the destination tensor before applying the destination tensor to the memory platform storage device.
25. A method for performing a transformation operation, comprising: Obtain a source tensor from memory, the source tensor containing data; Perform a data transformation on the source tensor to obtain a destination tensor; and Apply the destination tensor to the memory; wherein, the obtaining, executing, and applying are all performed using direct memory access.