Techniques to learn and offload common modes of memory access and computing
By introducing memory execution circuits and machine learning detector software into the computing system, and automatically routing memory access and computing operations, the delay and energy consumption problems of traditional systems when dealing with irregular memory access modes are solved, and a more efficient and flexible computing system is achieved.
Patent Information
- Application Number
- CN202510158082.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2020-06-27
- Filing Date
- 2020-12-10
- Publication Date
- 2025-05-16
AI Technical Summary
When traditional computing systems deal with irregular memory access modes, they lead to unpredictable delays and high energy consumption, and existing solutions are difficult to flexibly configure and adapt to dynamic application needs.
By introducing in-memory execution circuitry and detector software into the computing system, machine learning is used to detect the benefits of offloading operations into the memory execution circuitry and automatically routing memory access and computing operations, reducing data movement between the CPU and off-chip memory.
It realizes the reduction of memory access latency, reduces energy consumption, and improves the flexibility and adaptability of the computing system, and is suitable for a variety of application scenarios.
Smart Images

Figure CN120010921A_ABST
Abstract
Description
[0001] This application is a divisional application of the invention patent application with application date of December 10, 2020, application number 202011432868.0, and name “Technology for learning and offloading common patterns of memory access and calculation”. Technical Field
[0002] Embodiments generally relate to techniques for computing systems. More specifically, embodiments relate to techniques for automatically routing memory accesses and computing operations to be performed by secondary computing devices. Background Art
[0003] Traditional memory architectures assume that most programs will repeatedly access the same set of memories in short time periods. That is, they follow the rules of spatial and temporal locality. However, many applications in graph analysis, machine learning, and artificial intelligence (AI) exhibit irregular memory access patterns that do not follow the traditional rules of spatial and temporal locality. Irregular memory access patterns are poorly addressed by traditional CPU and GPU architectures, resulting in unpredictable delays when performing memory operations. One reason for this is that irregular memory access requires repeated data movement between the CPU and off-chip memory storage.
[0004] Moving data between the CPU core and off-chip memory incurs approximately 100 times more energy than floating-point operations inside the CPU core. Traditionally, compute-centric von-Neumann architectures are increasingly constrained by memory bandwidth and energy consumption. Hardware devices known as in-memory compute (IMC) or compute near memory (CNM) devices place the computing power within or near the memory array itself. These devices can eliminate or significantly reduce the data movement required to execute a program.
[0005] In the absence of widespread standards to specify how to embed IMC or CNM devices in computing systems, most current solutions require users to manually map desired computing kernels to IMC or CNM memory arrays. This solution is quite inflexible, making it difficult to configure these devices for a variety of different applications. Furthermore, because this solution relies on static compilation of applications, they cannot adapt to the dynamic aspects of real-world application execution (e.g., dynamic resource usage, workload characteristics, memory access patterns, etc.). Summary of the invention
[0006] According to one aspect of the present disclosure, a system is provided, comprising: a memory; a central processing unit (CPU) for performing a first operation; an in-memory execution circuit in the memory; and detector software for offloading the second operation to the in-memory execution circuit after detecting a benefit of offloading the second operation to the in-memory execution circuit based on machine learning, the in-memory execution circuit being used to perform the second operation in a first execution path, and the CPU being used to perform the first operation in a second execution path.
[0007] According to another aspect of the present disclosure, at least one computer-readable random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), flash memory, dynamic random access memory (DRAM), static random access memory (SRAM), solid-state drive (SSD), hard disk drive (HDD) or optical disk is provided, including instructions for causing at least one processor to at least: perform a first operation; and after detecting the benefit of offloading the second operation to an in-memory execution circuit based on machine learning, causing the second operation to be offloaded to the in-memory execution circuit, the in-memory execution circuit being used to perform the second operation in a first execution path, and at least one processor being used to perform the first operation in a second execution path.
[0008] According to another aspect of the present disclosure, a method is provided, comprising: executing a first instruction via a CPU; and after detecting the benefit of offloading the second instruction to an in-memory execution circuit based on machine learning, unloading the second instruction to the in-memory execution circuit, the in-memory execution circuit is used to execute the second instruction in a first execution path, and the CPU is used to execute the first instruction in a second execution path. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Various advantages of the embodiments will become apparent to those skilled in the art upon reading the following description and appended claims, and by referring to the following drawings, in which:
[0010] Figure 1 is a block diagram illustrating an example of a system for offloading a sequence of machine instructions and memory accesses according to one or more embodiments;
[0011] Figure 2 is a diagram illustrating a diagram of an example of a sequence detector according to one or more embodiments;
[0012] Figures 3 to 5 Diagrams are provided that illustrate some aspects of an example application for offloading a sequence of machine instructions and memory accesses according to one or more embodiments;
[0013] FIG. 6A to FIG. 6BA flow chart illustrating the operation of an example of a system for offloading a sequence of machine instructions and memory accesses according to one or more embodiments is provided;
[0014] Figure 7 is a block diagram illustrating an example of a performance-enhanced computing system according to one or more embodiments;
[0015] Figure 8 is a block diagram illustrating an example semiconductor device according to one or more embodiments;
[0016] Fig. 9 is a block diagram illustrating an example of a processor according to one or more embodiments; and
[0017] Fig.10 is a block diagram illustrating an example of a multi-processor based computing system according to one or more embodiments. DETAILED DESCRIPTION
[0018] In summary, embodiments provide a computing system that automatically offloads sequences of machine instructions and memory accesses to secondary computing devices, such as in-memory computing (IMC) devices or near-memory computing (CNM) devices. Embodiments also support determining which memory access and compute operations to route to IMC / CNM hardware based on identifying relationships between memory accesses and computations. In addition, embodiments include techniques to exploit dependencies that exist between memory operations and other instructions to learn and predict sequences with these dependencies.
[0019] More specifically, an embodiment of a computing system provides a memory system that uses an auxiliary computing device and a trainable machine learning component to intelligently route computing and memory operations between a CPU and an auxiliary device. The computing system automatically learns according to the embodiment to identify common patterns of memory access and computing instructions, determines whether it is useful to offload a sequence from the CPU to the auxiliary computing device, and maps machine instructions from the CPU to the auxiliary computing device for execution. In addition, an embodiment provides the following technology, which will preferentially unload sequences that will result in high latency from memory operations, such as cache misses from irregular memory accesses, or memory operations that will typically require extensive cache consistency updates. Thus, the embodiment will speed up program execution times and alleviate memory bottlenecks, including in multi-threaded applications where there may be many ongoing competing demands for CPU resources.
[0020] Figure 11 is a block diagram illustrating an example of a computing system 100 for offloading a sequence of machine instructions and memory accesses according to one or more embodiments with reference to the components and features described herein (including but not limited to the drawings and associated descriptions). The system may include a pre-fetch unit, an instruction decoder, a sequence detector, a decision engine, a central processing unit (CPU) with a program counter (PC), an auxiliary computing device (in Figure 1 IMC), and instruction converter. Figure 1 An in-memory computing (IMC) device is illustrated as the secondary computing device, but the system may equally use near-memory computing (CNM) as the secondary computing device.
[0021] The computing system 100 may be operated by loading the binary code of an application into memory. The binary code may then be provided as input to the CPU, and to a hardware prefetcher, which may include a large look-ahead window. The prefetcher may collect an ordered sequence of binary code that needs to be executed. The CPU may begin executing these instructions using its normal cycles, while the prefetched binary code may be decoded into machine instructions using an instruction decoder (which may be a copy of the CPU decoder). The decoded machine instructions may be provided as input to a sequence detector, which may detect instructions that may be offloaded to an auxiliary computing device (IMC, such as Figure 1 The decision engine (in Figure 1 The CPU instructions that are offloaded can be converted into instructions for in-memory computing (IMC) hardware via an instruction converter, which can then execute these instructions in parallel with the CPU.
[0022] In order to maintain the correct order of program execution, the first address of the unloaded instruction can be specially marked in the program counter (PC) with a special value indicating that it is being processed elsewhere. Once the IMC has completed executing the unloaded instruction, the IMC can store the calculated results back to the main memory. The IMC can also increment the corresponding program counter to the address of the next instruction to be obtained by the CPU. The CPU can then continue to process the application at the next instruction until the next unload decision is made.
[0023] In some embodiments, the IMC may act as main memory during normal CPU execution. In this way, memory load and store operations from the CPU may be handled as usual. When the system decides via the sequence detector to offload certain computations to the IMC, the converted code may be sent to the IMC controller. The IMC controller may queue the execution of this converted code and schedule execution appropriately whenever resources are available. Because no memory copies are involved in this process (that is, updates to memory are in situ), the memory addresses in the code being executed by the IMC remain the same as the memory addresses of the CPU code. The results calculated in the IMC may be transmitted back to the CPU and stored in CPU registers; in this case, this occurs before the CPU continues processing.
[0024] In other embodiments, the IMC and main memory (e.g., DRAM) may be separate. In this scenario, the system must copy data back and forth between the IMC and main memory to maintain consistency between the copies. Address mapping between the IMC and main memory addresses will also be required. In this alternative scenario, additional hardware logic will be required, and the throughput of the system will be much higher for multi-threaded read-only workloads.
[0025] An analogy can be made between the operation of system 100 and the multiple memory systems believed to exist in the brain. It has been proposed that the brain includes a procedural memory system that automatically learns to detect sequences of frequently used operations and offloads them to a separate neural system that is protected from interference from the main memory system. This is claimed to free up the main memory system for use and to make operations in procedural memory execute more quickly so that they are often performed automatically (i.e., without conscious thought). Similarly, system 100 operates via a procedural memory system (IMC / CNM) with trainable machine learning components (i.e., sequence detector, offload decision) to intelligently route computation and memory operations between the CPU and IMC / CNM.
[0026] Adaptive Sequence Detector
[0027] System 100( Figure 1, already discussed) can be an adaptive algorithm that learns to recognize common sequences of memory accesses and computations. In some embodiments, the sequence detector can be trained to recognize only sequences that can be offloaded if there is an execution time benefit. Machine instructions from the CPU (or from a CPU instruction decoder) that contain both computational operations (e.g., addition) and memory operations (e.g., load a memory address) can be provided as input to the adaptive algorithm. The algorithm can then learn the transition probabilities, or sequence dependencies, between instructions to automatically detect recurring sequences involving computations and memory operations.
[0028] The adaptive algorithm for sequence detection in system 100 may be implemented via a trained neural network. In some embodiments, the neural network may be implemented in a field programmable gate array (FPGA) accelerator. In an embodiment, the neural network may be implemented in a combination of a processor and an FPGA accelerator. In some embodiments, a recurrent neural network (RNN) may be used to implement the adaptive sequence detection algorithm. The recurrent neural network may take a sequence of machine instructions and output a sequence of decisions about whether to offload the instructions. Figure 2 2 is a diagram illustrating an example of a sequence detector 200 according to one or more embodiments with reference to the components and features described herein (including but not limited to the figures and associated descriptions). Figure 2 As shown in , the sequence detector is an RNN. The RNN can receive a sequence of machine instructions {i 1 ,i 2 ,i 3 ,…i n} as input (labeled as element 202), and may output a sequence of decisions regarding whether to offload instructions {o 1 ,o 2 ,o 3 ,…o n} (labeled as element 204). Given a machine instruction i (e.g., i 2 ) and the previous RNN state (box), the RNN output is about the instruction (e.g., i 2 ) True or false value that is offloaded to the IMC hardware (for example, o 2 ).
[0029] In some embodiments, the prefetcher may have a large look-ahead window, and the RNN may process instructions that are yet to be executed on the CPU. In this case, if the RNN makes a decision to offload instructions, additional logic may be used to mark the offloaded instructions in the program counter on the CPU.
[0030] The feed-forward pass of RNN can be implemented in hardware, such as Intel Gaussian Neural Accelerator (GNA). In some embodiments, another neural network model may be used as a sequence detector, and in this case, the other neural network model may be similarly used via, for example, Gaussian Neural Accelerator (GNA) is implemented in hardware.
[0031] In some embodiments, the RNN can be trained in hardware simulation using benchmarks that are known to increase cache misses and cause memory latency issues. When running the training simulation, the training data will include information about whether it is appropriate to offload a given instruction to the IMC (called "prophecy data"). This prophecy data must include the overall execution time, but may also include other data, such as markers for the start and end of a repeated sequence. The prophecy data can be used to construct an appropriate error signal for training using error back propagation. The goal of training the RNN is to have the RNN remember the sequence that caused the cache miss.
[0032] In other embodiments, the RNN can be trained to directly predict cache misses. In this scenario, the RNN can be directly trained to detect sequences that result in long latency for memory accesses. Logic (e.g., an algorithm) can be added to convert the RNN output (i.e., the sequence of predicted cache misses) into an offload decision, either by adding another neural network layer or by creating a set of static rules for offloading.
[0033] In another embodiment, the RNN can be trained by embedding it in a reinforcement learning agent. The state of the environment of the reinforcement learning agent is a sequence of instructions, and the action it takes is to decide whether to offload the instruction. The reward obtained by the agent is proportional to the execution time of the instruction. Therefore, the sequence detector can run independently without user intervention.
[0034] After training an RNN in simulation, you can use the "fine-tuning" mode to continue modifying the RNN weights to optimize for a specific application. Fine-tuning mode can be run until the average execution time decreases. Once suitable performance is achieved after fine-tuning, the RNN can be run in inference mode (with weights static or frozen).
[0035] In-Memory Computing (IMC) Hardware
[0036] Hardware devices called in-memory computing (IMC) devices can have multiple processors attached to them. Examples of IMC hardware include Non-Volatile Memory (NVM) such as Optane memory, and resistive memory (ReRAM). IMC devices support in-memory computing by repurposing their memory structure to have in-situ computing capabilities. For example, ReRAM stores data in the form of resistance of titanium oxide; by sensing the current on the bit line, Ohm's law and Kirchhoff's law can be used to calculate the dot product of the input voltage and the cell conductance.
[0037] Embodiments may use IMC hardware with distributed memory arrays interleaved with small bit logic that can be programmed to perform simple functions on data in parallel "in memory." For example, these memory arrays (up to several thousand distributed memory arrays) plus tiny computational units can be programmed as single instruction multiple data (SIMD) processing units that can compute concurrently, extending the memory arrays to support in-place operations like dot products, additions, element-wise multiplications, and subtractions.
[0038] According to an embodiment, the IMC hardware may be used in a SIMD execution model where in each cycle, instructions issued to the IMC are multicast to multiple memory arrays and executed in lock step. The IMC hardware may also have a queue of pending instruction "requests" and a scheduler that may enable instruction level parallelism. Performance may be directed, for example, based on overall capacity and based on speed of access to memory.
[0039] In an embodiment, the unloaded sequence may be created as a state machine to be executed on the IMC. If the IMC exceeds its capacity, an eviction policy may determine which state machines should be kept on the IMC and which should be evicted back to be executed via the host CPU in order to achieve optimal instruction throughput. In some embodiments, a recency policy may be used to simply evict the least recently used state machine. In some embodiments, in addition to the recency policy, the IMC may also store information about the efficiency benefits of each state machine and balance recency with the overall benefits of keeping the sequence in the IMC. The scheduler may also consider memory media properties, such as media wear level measurements and thermal conditions, before creating and executing state machine schedules to ensure optimal use of the in-memory computing hardware.
[0040] Instruction Converter
[0041] The instruction converter may be implemented as a hardware table indicating a direct mapping between a CPU machine instruction and an IMC instruction according to some embodiments. In most cases, the CPU machine instruction will have an equivalent instruction with a 1:1 mapping on the IMC. In a few cases, there may be a one-to-many mapping - for example, a fuse-multiply-add (FMA) instruction on the CPU may be executed as three separate instructions on the IMC device. In the case where the CPU contains an instruction that has no equivalent on the IMC, the instruction should not be unloaded. Any decision to unload such an instruction (or a sequence containing such an instruction) will be an incorrect decision that results in a longer processing delay. Therefore, those cases where the CPU instruction is not mapped on the IMC device may be included in the training data set of the sequence detector and the unloading engine to ensure that the system avoids unloading sequences with such instructions.
[0042] Execution decision engine
[0043] The execution decision engine determines when the instruction sequence should be offloaded to execute on the secondary computing device (e.g., IMC). In some embodiments, the decision engine can be implemented in logic that is merged with the sequence detector. For example, in a computer using a recurrent neural network ( Figure 2 , already discussed), the RNN can provide a decision as to whether to offload a particular instruction sequence as an output, thereby operating as a decision engine. An embodiment in which a sequence detector is implemented using another neural network structure can similarly provide an execution decision as an output of the neural network.
[0044] In some cases, the overall throughput of the IMC hardware may be lower than that of the CPU, especially if the IMC executes several state machines at once while the CPU is idle. Therefore, embodiments may include additional logic to check whether the CPU is idle. If the CPU is idle, the execution engine may delegate the task of executing the sequence to both the CPU and the IMC, achieving faster results, and killing (i.e., terminating or stopping execution of) the remaining processes. To ensure consistency in processing, the pipeline for the slower process may be flushed.
[0045] Sample Application: Scatter-Gather Programming Model
[0046] Embodiments for offloading sequences of machine instructions and memory accesses may be applied to algorithms using the scatter-gather programming model. The scatter-gather programming model may be used to compute many graph algorithms such as breadth-first search, connected component labeling, and PageRank. Figure 3 A diagram 300 illustrating a scatter-gather programming model is provided. Figure 3 As shown in , the scatter-gather model can have three key levels: gather, apply, and scatter. Figure 3The left panel of the diagram in shows that at the aggregation level, a node (V) receives information from its incoming neighbors (U 1 and U 2 ) collects information. The middle panel of the figure shows that in the apply level, the node (V) performs some computation on the information received in the gather step. The right panel of the figure shows that in the scatter operation, the node (V) broadcasts some information (which usually contains the result of the apply step) to its outgoing neighbors (U 3 and U 4 ).exist Figure 4 Pseudo code 400 describing the scatter-gather programming model is illustrated in FIG. Figure 4 In the pseudocode in , these key levels (the application step is omitted for brevity, but it occurs after each call to vertex_gather) are applied in a loop until some stopping condition is met (for example, no nodes have updates).
[0047] When executing an algorithm using the scatter-gather programming model, a computation cycle occurs, where each iteration consists of a scatter stage and a gather / apply stage. In a given iteration, the set of "active" vertices that need to be scattered and updated is called the compute frontier. After these vertex scatter updates, the updated vertices need to be collected to gather all inputs and use these inputs to apply the update function. The computation cycle terminates when the compute frontier becomes empty. By correctly defining vertex_scatter() and vertex_gather() functions, a large set of graphics algorithms can be computed. Figure 5 5 illustrates the computational frontier 500 of a breadth-first search (BFS) algorithm while exploring the second level of the BFS tree. Figure 5 As shown in , all nodes in the shaded boxes are marked as active vertices that require iterations of the scatter-gather cycle.
[0048] For a given graph in the scatter-gather programming model, the same computation is performed repeatedly in the application level. Since updating each vertex depends on information from other connected vertices, however, executing this loop without offloading can result in pointer chasing, causing cache misses and unpredictable delays in execution. By offloading the sequence of machine instructions and memory accesses according to an embodiment, the execution of this programming model can be accelerated - by learning the most common sequence of computations (sequences of operations in the update function at the application level) and memory accesses (memory locations that receive updated vertices during the scatter level) required for a certain (certain) iteration of the loop. With the offloading technique as described herein, the data in each vertex will occupy a static set of memory addresses that need to be loaded and stored, and the scatter and gather functions consist of a set of instructions that operate on the data in each vertex.
[0049] Figure 6A-6B Flowcharts are provided that illustrate examples of processes 600 and 650 for operating a system for offloading a sequence of machine instructions and memory accesses according to one or more embodiments with reference to components and features described herein (including but not limited to the figures and associated descriptions). Processes 600 and 650 may be implemented in accordance with the embodiments discussed herein with reference to Figure 1-Figure 2 More specifically, processes 600 and 650 may be implemented in one or more modules as a collection of logic instructions stored in a machine or computer readable storage medium such as a random access memory (RAM), a read only memory (ROM), a programmable ROM (PROM), firmware, a flash memory, etc.; may be implemented in configurable logic such as a programmable logic array (PLA), a field programmable gate array (FPGA), a complex programmable logic device (CPLD), etc.; may be implemented in fixed function logic hardware using circuit technology such as an application specific integrated circuit (ASIC), a complementary metal oxide semiconductor (CMOS), or a transistor-transistor logic (TTL) technology; or in any combination thereof.
[0050] For example, computer program code for performing the operations shown in process 600 may be written in any combination of one or more programming languages, including object-oriented programming languages such as JAVA, SMALLTALK, C++, etc., and traditional procedural programming languages such as the "C" programming language or similar programming languages. In addition, the logic instructions may include assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, state setting data, configuration data of integrated circuits, state information of electronic circuits local to personalized hardware and / or other structural components (e.g., host processors, central processing units / CPUs, microcontrollers, etc.).
[0051] Go to Fig. 6A, for process 600, illustrated processing block 610 provides for identifying patterns of memory access and computation instructions based on an input set of machine instructions via a neural network. The neural network may be a recurrent neural network (RNN). The patterns of memory access and computation instructions may include one or more of transition probabilities or sequence dependencies between instructions of the input set of instructions. Illustrated processing block 615 provides for determining, via the neural network, a sequence of instructions to be offloaded for execution by an auxiliary computing device based on the identified patterns of memory access and computation instructions. The determined sequence of instructions to be offloaded may include one or more of the following: a recurring sequence, a sequence with offload execution time benefits, a sequence that will result in high latency due to repeated memory operations, and / or a sequence that will result in cache misses. Illustrated processing block 620 provides for converting the sequence of instructions to be offloaded from instructions executable by a central processing unit (CPU) into instructions executable by an auxiliary computing device.
[0052] Illustrated processing block 630 provides for training a recurrent neural network (RNN) via one or more of: hardware emulation with a benchmark known to increase cache misses and cause memory latency issues, direct training to detect sequences that cause long latency for memory accesses, or embedded in a reinforcement learning agent. Illustrated processing block 640 provides for performing fine-tuning to continue modifying the RNN weights to optimize for a specific application.
[0053] Now go to Figure 6B , for process 650, at block 660, a check is made to determine whether the CPU is idle. If the CPU is not idle, the process terminates. If the CPU is idle, the process continues at illustrated processing block 665, which provides for assigning the CPU a task of a first process for executing a sequence of instructions to be offloaded, and for assigning a task of a second process concurrent with the first process for executing the converted offloaded instructions to the auxiliary computing device. At block 670, a determination is made as to whether the second process is completed before the first process. If so (i.e., the second process is completed before the first process), illustrated processing block 675 provides for accepting the execution result of the second process and terminating the first process. Otherwise, if the second process is not completed before the first process, illustrated processing block 680 provides for accepting the execution result of the first process and terminating the second process.
[0054] Figure 7A block diagram of an example computing system 10 for offloading memory access and computing operations to be performed by an auxiliary computing device according to one or more embodiments is shown with reference to the components and features described herein (including but not limited to the drawings and associated descriptions). The system 10 can generally be part of an electronic device / platform having the following functions: computing and / or communication functions (e.g., servers, cloud infrastructure controllers, database controllers, notebook computers, desktop computers, personal digital assistants / PDAs, tablet computers, convertible tablet devices, smart phones, etc.), imaging functions (e.g., cameras, camcorders), media playback functions (e.g., smart TVs / TVs), wearable functions (e.g., watches, glasses, headwear, footwear, jewelry), vehicle functions (e.g., cars, trucks, motorcycles), robotic functions (e.g., autonomous robots), Internet of Things (IoT) functions, etc., or any combination of these. In the illustrated example, the system 10 may include a host processor 12 (e.g., a central processing unit / CPU) having an integrated memory controller (MC) 14 that can be coupled to a system memory 20. Host processor 12 may include any type of processing device, such as a microcontroller, a microprocessor, a RISC processor, an ASIC, etc., together with associated processing modules or circuits. System memory 20 may include any non-transitory machine or computer readable storage medium, such as RAM, ROM, PROM, EEPROM, firmware, flash memory, etc., configurable logic, such as PLA, FPGA, CPLD, fixed function logic hardware using circuit technology such as ASIC, CMOS or TTL technology, or any combination thereof suitable for storing instructions 28.
[0055] The system 10 may also include an input / output (I / O) subsystem 16. The I / O subsystem 16 may communicate with, for example, one or more input / output (I / O) devices 17, a network controller (e.g., a wired and / or wireless NIC), and a storage device 22. The storage device 22 may include any appropriate non-transitory machine or computer-readable memory type (e.g., flash memory, DRAM, SRAM (static random access memory), a solid state drive (SSD), a hard disk drive (HDD), an optical disk, etc.). The storage device 22 may include a mass storage device. In some embodiments, the host processor 12 and / or the I / O subsystem 16 may communicate with the storage device 22 (all or part thereof) via a network controller 24. In some embodiments, the system 10 may also include a graphics processor 26 (e.g., a graphics processing unit / GPU) and an AI accelerator 27. In some embodiments, the system 10 may also include an auxiliary computing device 18, such as an in-memory computing (IMC) device or a near memory computing (CNM) device. In an embodiment, the system 100 may also include a visual processing unit (VPU) not shown.
[0056] The host processor 12 and the I / O subsystem 16 may be implemented together on a semiconductor die as a system on chip (SoC) 11, as shown by the solid line. The SoC 11 can therefore operate as a computing device that automatically routes memory access and computing operations to be performed by the auxiliary computing device. In some embodiments, the SoC 11 may also include one or more of a system memory 20, a network controller 24, a graphics processor 26, and / or an AI accelerator 27 (shown by the dotted line). In some embodiments, the SoC 11 may also include other components of the system 10.
[0057] The host processor 12, I / O subsystem 16, graphics processor 26, AI accelerator 27 and / or VPU may execute program instructions 28 retrieved from system memory 20 and / or storage device 22 to perform the functions described herein. Figure 6A-6B One or more aspects of the processes 600 and 650 described herein. Thus, for example, execution of the instructions 28 may cause the SoC 11 to recognize patterns of memory access and computation instructions based on an input set of machine instructions via a neural network, determine a sequence of instructions to be offloaded to be executed by an auxiliary computing device based on the recognized patterns of memory access and computation instructions via the neural network, and convert the sequence of instructions to be offloaded from instructions executable by a central processing unit (CPU) into instructions executable by an auxiliary computing device. The system 10 may implement the system as described herein with reference to Figure 1-Figure 2One or more aspects of the described computing system 100, sequence detector, decision engine and / or instruction converter. The system 10 is therefore considered to be performance-enhanced in at least the following aspects: the system intelligently routes computation and memory operations between the CPU and auxiliary computing devices to improve computational performance and reduce execution time.
[0058] The computer program code for executing the processes described above may be written in any combination of one or more programming languages and implemented as program instructions 28, including object-oriented programming languages such as JAVA, JAVASCRIPT, PYTHON, SMALLTALK, C++, etc., and / or traditional procedural programming languages such as the "C" programming language or similar programming languages. In addition, the program instructions 28 may include assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcodes, state setting data, configuration data for integrated circuits, status information of electronic circuits local to personalized hardware and / or other structural components (such as host processors, central processing units / CPUs, microcontrollers, microprocessors, etc.).
[0059] I / O devices 17 may include one or more input devices, such as a touch screen, keyboard, mouse, cursor control device, touch screen, microphone, digital camera, video recorder, video camera, biometric scanner and / or sensor; the input device can be used to input information and interact with system 10 and / or other devices. I / O devices 17 may also include one or more output devices, such as a display (e.g., a touch screen, a liquid crystal display / LCD, a light emitting diode / LED display, a plasma panel, etc.), a speaker and / or other visual or audio output device. The input and / or output devices can be used, for example, to provide a user interface.
[0060] Figure 8 A block diagram of an example semiconductor device 30 for offloading memory access and computing operations to be performed by an auxiliary computing device is shown with reference to the components and features described herein (including but not limited to the figures and associated descriptions) according to one or more embodiments. The semiconductor device 30 may be implemented, for example, as a chip, a die, or other semiconductor package. The semiconductor device 30 may include one or more substrates 32 composed of, for example, silicon, sapphire, gallium arsenide, etc. The semiconductor device 30 may also include logic 34 composed of, for example, (one or more) transistor arrays and (one or more) other integrated circuit (IC) components coupled to the (one or more) substrates 32. The logic 34 may be implemented at least in part in configurable logic or fixed function logic hardware. The logic 34 may implement the above reference Figure 7 The system on chip (SoC) 11 described herein. The logic 34 may implement one or more aspects of the processes described above, including those described herein with reference to Figure 6A-6BThe logic 34 may be implemented as described herein with reference to Figure 1-Figure 2 The computing system 100, sequence detector, decision engine, and / or instruction converter described herein are described in one or more aspects. The device 30 is therefore considered to be performance-enhanced in at least the following aspects: the system intelligently routes computation and memory operations between the CPU and auxiliary computing devices to improve computational performance and reduce execution time.
[0061] The semiconductor device 30 may be constructed using any suitable semiconductor manufacturing process or technology. For example, logic 34 may include a transistor channel region positioned (e.g., embedded) within substrate(s) 32. Thus, the interface between logic 34 and substrate(s) 32 may not be an abrupt junction. Logic 34 may also be considered to include epitaxial layers grown on an initial wafer of substrate(s) 34.
[0062] Fig. 9 1 is a block diagram illustrating an example processor core 40 according to one or more embodiments with reference to the components and features described herein (including but not limited to the figures and associated descriptions). The processor core 40 may be a core for any type of processor, such as a microprocessor, an embedded processor, a digital signal processor (DSP), a network processor, or other device that executes code. Fig. 9 Only one processor core 40 is shown, but the processing element may alternatively include more than one Fig. 9 . Processor core 40 may be a single-threaded core, or for at least one embodiment, processor core 40 may be multi-threaded in that it may include more than one hardware thread context (or "logical processor") per core.
[0063] Fig. 9 Also illustrated is a memory 41 coupled to the processor core 40. The memory 41 may be any of a variety of memories (including various layers of a memory hierarchy) known to those skilled in the art or otherwise available. The memory 41 may include one or more code 42 instructions to be executed by the processor core 40. The code 42 may implement the instructions described herein. Figure 6A-6B Processor core 40 may implement one or more aspects of processes 600 and 650 as described herein. Figure 1-Figure 2The computing system 100, sequence detector, decision engine, and / or instruction converter described herein may be described in detail. The processor core 40 follows a program sequence of instructions indicated by code 42. Each instruction may enter a front end portion 43 and be processed by one or more decoders 44. The decoder 44 may generate micro-operations such as fixed-width micro-operations of a predetermined format as its output, or may generate other instructions, micro-instructions, or control signals that reflect the original code instructions. The illustrated front end portion 43 also includes register renaming logic 46 and scheduling logic 48, which generally allocate resources and queue operations corresponding to the converted instructions for execution.
[0064] Processor core 40 is shown as including execution logic 50 having a set of execution units 50-1 to 55-N. Some embodiments may include several execution units dedicated to a specific function or set of functions. Other embodiments may include only one execution unit or an execution unit that can perform a specific function. The illustrated execution logic 50 performs the operations specified by the code instructions.
[0065] After execution of the operations specified by the code instructions is complete, back-end logic 58 retires the instructions of code 42. In one embodiment, processor core 40 allows out-of-order execution of instructions, but requires in-order retirement of instructions. Retirement logic 59 may take various forms known to those skilled in the art (e.g., a reorder buffer or the like). In this manner, processor core 40 transforms during execution of code 42 at least the outputs generated by the decoders, the hardware registers and tables utilized by register renaming logic 46, and any registers modified by execution logic 50 (not shown).
[0066] Although in Fig. 9 Although not shown in the figure, the processing element may include other elements on the chip along with the processor core 40. For example, the processing element may include memory control logic along with the processor core 40. The processing element may include I / O control logic and / or may include I / O control logic integrated with the memory control logic. The processing element may also include one or more caches.
[0067] Fig.10 is a block diagram illustrating an example of a multiprocessor-based computing system 60 according to one or more embodiments with reference to the components and features described herein, including but not limited to the figures and associated descriptions. Multiprocessor system 60 includes a first processing element 70 and a second processing element 80. Although two processing elements 70 and 80 are shown, it is to be understood that embodiments of system 60 may also include only one such processing element.
[0068] System 60 is shown as a point-to-point interconnect system, in which first processing element 70 and second processing element 80 are coupled via point-to-point interconnect 71. It should be understood that Fig.10 Any or all of the interconnects shown in may be implemented as multi-drop buses rather than point-to-point interconnects.
[0069] like Fig.10 As shown in FIG. 1 , each of processing elements 70 and 80 may be a multi-core processor including first and second processor cores (i.e., processor cores 74a and 74b and processor cores 84a and 84b). Such cores 74a, 74b, 84a, 84b may be configured to be as described above. Fig. 9 The instruction code is executed in a manner similar to that described above.
[0070] Each processing element 70, 80 may include at least one shared cache 99a, 99b. The shared caches 99a, 99b may store data (e.g., instructions) that are utilized by one or more components of the processor, such as cores 74a, 74b and 84a, 84b, respectively. For example, the shared caches 99a, 99b may cache data stored in local memory 62, 63 for faster access by the components of the processor. In one or more embodiments, the shared caches 99a, 99b may include one or more intermediate level caches, such as level 2 (L2), level 3 (L3), level 4 (L4), or other level caches, a last level cache (LLC), and / or combinations thereof.
[0071] Although shown as having only two processing elements 70, 80, it is to be understood that the scope of the embodiment is not limited thereto. In other embodiments, one or more additional processing elements may be present in a given processor. Alternatively, one or more of the processing elements 70, 80 may be an element other than a processor, such as an accelerator or a field programmable gate array. For example, (one or more) additional processing elements may include (one or more) additional processors identical to the first processor 70, (one or more) additional processors heterogeneous or asymmetric with the first processor 70, accelerators (e.g., graphics accelerators or digital signal processing (DSP) units), field programmable gate arrays, or any other processing elements. Between the processing elements 70, 80, there may be various differences in terms of the range of value metrics including architectural characteristics, microarchitecture characteristics, thermal characteristics, power consumption characteristics, etc. These differences may actually present themselves as asymmetry and heterogeneity between the processing elements 70, 80. For at least one embodiment, various processing elements 70, 80 may be present in the same die package.
[0072] The first processing element 70 may also include a memory controller logic (MC) 72 and point-to-point (PP) interfaces 76 and 78. Similarly, the second processing element 80 may include a MC 82 and PP interfaces 86 and 88. Fig.10As shown in , MCs 72 and 82 couple the processors to respective memories, namely memory 62 and memory 63, which may be part of the main memory locally attached to the respective processors. Although MCs 72 and 82 are shown as being integrated into processing elements 70, 80, for alternative embodiments, the MC logic may be discrete logic external to processing elements 70, 80, rather than integrated therein.
[0073] First processing element 70 and second processing element 80 may be coupled to I / O subsystem 90 via PP interconnects 76 and 86, respectively. Fig.10 As shown in , I / O subsystem 90 includes PP interfaces 94 and 98. In addition, I / O subsystem 90 includes interface 92 to couple I / O subsystem 90 with high performance graphics engine 64. In one embodiment, bus 73 may be used to couple graphics engine 64 to I / O subsystem 90. Alternatively, a point-to-point interconnect may couple these components.
[0074] In turn, the I / O subsystem 90 may be coupled to the first bus 65 via an interface 96. In one embodiment, the first bus 65 may be a Peripheral Component Interconnect (PCI) bus, or a bus such as a PCI Express bus or another third generation I / O interconnect bus, but the scope of the embodiment is not limited in this regard.
[0075] like Fig.10 As shown in FIG. 1 , various I / O devices 65a (e.g., biometric scanners, speakers, cameras, sensors) may be coupled to a first bus 66, and a bus bridge 66 may couple the first bus 65 to a second bus 67. In one embodiment, the second bus 67 may be a low pin count (LPC) bus. In one embodiment, various devices may be coupled to the second bus 67, including, for example, a keyboard / mouse 67a, (one or more) communication devices 67b, and a data storage unit 68, such as a disk drive or other mass storage device, which may include code 69. The illustrated code 69 may implement one or more aspects of the processes described above, including those described herein with reference to FIG. Figure 6A-6B The process 600 and 650 described. The code 69 shown can be combined with the code 42 ( Fig. 9 ) is similar. In addition, audio I / O 67c can be coupled to second bus 67 and battery 61 can supply power to computing system 60. System 60 can be implemented as described herein with reference to Figure 1-Figure 2 One or more aspects of the computing system 100, sequence detector, decision engine, and / or instruction converter are described.
[0076] Note that other embodiments are contemplated. For example, instead of Fig.10 Instead of a point-to-point architecture, the system can implement a multi-point branch bus or other such communication topology. Fig.10 As shown in the more or less integrated chips to divide Fig.10 components.
[0077] Embodiments of each of the above systems, apparatuses, components, and / or methods, including system 10, semiconductor device 30, processor core 40, system 60, computing system 100, sequence detector, decision engine, instruction converter, processes 600 and 650, and / or any other system component, may be implemented in hardware, software, or any suitable combination of these. For example, a hardware implementation may include configurable logic, such as a programmable logic array (PLA), a field programmable gate array (FPGA), a complex programmable logic device (CPLD), or fixed function logic hardware using circuit technology such as application specific integrated circuit (ASIC), complementary metal oxide semiconductor (CMOS), or transistor-transistor logic (TTL) technology, or any combination thereof.
[0078] Alternatively or additionally, all or some parts of the aforementioned systems and / or components and / or methods may be implemented in one or more modules as a collection of logic instructions stored in a machine or computer readable storage medium such as a random access memory (RAM), a read only memory (ROM), a programmable ROM (PROM), firmware, flash memory, etc., to be executed by a processor or computing device. For example, computer program code for executing the operation of the component may be written in any combination of programming languages applicable / appropriate to one or more operating systems (OS), including object-oriented programming languages such as PYTHON, PERL, JAVA, SMALLTALK, C++, C#, etc., as well as traditional procedural programming languages such as the "C" programming language or similar programming languages.
[0079] Additional notes and examples:
[0080] Example 1 includes a computing system comprising an auxiliary computing device, and an apparatus coupled to the auxiliary computing device, the apparatus comprising one or more substrates and logic coupled to the one or more substrates, wherein the logic is at least partially implemented in one or more of configurable logic or fixed-function hardware logic, the logic coupled to the one or more substrates being used to: recognize patterns of memory access and computation instructions based on an input set of machine instructions via a neural network, determine a sequence of instructions to be offloaded to be executed by the auxiliary computing device based on the recognized patterns of memory access and computation instructions via the neural network, and convert the sequence of instructions to be offloaded from instructions executable by a central processing unit (CPU) to instructions executable by the auxiliary computing device.
[0081] Example 2 includes the system of Example 1, wherein the neural network includes a recurrent neural network (RNN), wherein the pattern of memory access and computation instructions includes one or more of transition probabilities or sequential dependencies between instructions of the input set of instructions, wherein the sequence of instructions to be offloaded includes one or more of: a recurring sequence, a sequence with offload execution time benefits, a sequence that will result in high latency due to repeated memory operations, or a sequence that will result in cache misses, and wherein the logic coupled to the one or more substrates further marks the instructions to be offloaded in a program counter on the CPU.
[0082] Example 3 includes the system of Example 2, wherein the RNN is trained via one or more of: hardware emulation with a benchmark known to increase cache misses and cause memory latency issues, direct training to detect sequences that cause long latencies for memory accesses, or embedding the RNN in a reinforcement learning agent.
[0083] Example 4 includes a system as described in Example 1, wherein the logic coupled to the one or more substrates further dispatches, when the CPU is idle, a task of a first process for executing a sequence of offloaded instructions to the CPU, and dispatches a task of a second process concurrent with the first process for executing the converted offloaded instructions to the auxiliary computing device, and if the second process is completed before the first process, accepts an execution result of the second process and terminates the first process, otherwise, if the second process is not completed before the first process, accepts an execution result of the first process and terminates the second process.
[0084] Example 5 includes the system of Example 1, wherein the input set of machine instructions may be provided via a hardware prefetcher and an instruction decoder, the hardware prefetcher having a large look-ahead window to capture binary source code, and the instruction decoder decodes the captured binary source code into machine instructions.
[0085] Example 6 includes the system of any one of Examples 1 to 5, wherein the logic for translating the sequence of instructions to be offloaded comprises a hardware table comprising a direct mapping between instructions executable by the CPU and instructions executable by the secondary computing device.
[0086] Example 7 includes a semiconductor device comprising one or more substrates, and logic coupled to the one or more substrates, wherein the logic is at least partially implemented in one or more of configurable logic or fixed-function hardware logic, the logic coupled to the one or more substrates being used to: recognize patterns of memory access and computation instructions based on an input set of machine instructions via a neural network, determine a sequence of instructions to be offloaded to be executed by an auxiliary computing device based on the recognized patterns of memory access and computation instructions via the neural network, and convert the sequence of instructions to be offloaded from instructions executable by a central processing unit (CPU) to instructions executable by the auxiliary computing device.
[0087] Example 8 includes the semiconductor device of Example 7, wherein the neural network includes a recurrent neural network (RNN), wherein the pattern of memory access and computation instructions includes one or more of transition probabilities or sequential dependencies between instructions of the input set of instructions, wherein the sequence of instructions to be offloaded includes one or more of: a recurring sequence, a sequence with offload execution time benefits, a sequence that will result in high latency due to repeated memory operations, or a sequence that will result in cache misses, and wherein the logic coupled to the one or more substrates further marks the instructions to be offloaded in a program counter on the CPU.
[0088] Example 9 includes the semiconductor device of Example 8, wherein the RNN is trained via one or more of: hardware emulation with a benchmark known to increase cache misses and cause memory latency issues, direct training to detect sequences that cause long latencies for memory accesses, or embedding the RNN in a reinforcement learning agent.
[0089] Example 10 includes a semiconductor device as described in Example 7, wherein the logic coupled to the one or more substrates further assigns the CPU a task of a first process for executing a sequence of offloaded instructions when the CPU is idle, and assigns the auxiliary computing device a task of a second process concurrent with the first process for executing the converted offloaded instructions, and if the second process is completed before the first process, accepts the execution result of the second process and terminates the first process, otherwise, if the second process is not completed before the first process, accepts the execution result of the first process and terminates the second process.
[0090] Example 11 includes the semiconductor device of Example 7, wherein the input set of machine instructions can be provided via a hardware prefetcher and an instruction decoder, the hardware prefetcher having a large look-ahead window to capture binary source code, and the instruction decoder decodes the captured binary source code into machine instructions.
[0091] Example 12 includes the semiconductor device of any one of Examples 7 to 11, wherein the logic for converting the sequence of instructions to be offloaded includes a hardware table that includes a direct mapping between instructions executable by the CPU and instructions executable by the auxiliary computing device.
[0092] Example 13 includes the semiconductor device of Example 7, wherein the logic coupled to the one or more substrates includes a transistor channel region located within the one or more substrates.
[0093] Example 14 includes at least one non-transitory computer-readable medium including a set of first instructions that, when executed by a computing system, cause the computing system to: recognize patterns of memory access and computation instructions based on an input set of machine instructions via a neural network, determine a sequence of instructions to be offloaded to be executed by an auxiliary computing device based on the recognized patterns of memory access and computation instructions via the neural network, and convert the sequence of instructions to be offloaded from instructions executable by a central processing unit (CPU) to instructions executable by the auxiliary computing device.
[0094] Example 15 includes at least one non-transitory computer-readable storage medium as described in Example 14, wherein the neural network includes a recurrent neural network (RNN), wherein the pattern of memory access and computation instructions includes one or more of transition probabilities or sequential dependencies between instructions of the input set of instructions, wherein the sequence of instructions to be offloaded includes one or more of: a recurring sequence, a sequence with offload execution time benefits, a sequence that will result in high latency due to repeated memory operations, or a sequence that will result in cache misses, and wherein the first instruction, when executed, further causes the computing system to mark the instructions to be offloaded in a program counter on the CPU.
[0095] Example 16 includes at least one non-transitory computer-readable storage medium as described in Example 15, wherein the RNN is trained via one or more of: hardware emulation with a benchmark known to increase cache misses and cause memory latency issues, direct training to detect sequences that cause long latencies for memory accesses, or embedding the RNN in a reinforcement learning agent.
[0096] Example 17 includes at least one non-transitory computer-readable storage medium as described in Example 14, wherein the first instruction, when executed, also causes the computing system to assign a task of a first process for executing a sequence of unloaded instructions to the CPU when the CPU is idle, and to assign a task of a second process concurrent with the first process for executing the converted unloaded instructions to the auxiliary computing device, and if the second process is completed before the first process, accept the execution result of the second process and terminate the first process, otherwise, if the second process is not completed before the first process, accept the execution result of the first process and terminate the second process.
[0097] Example 18 includes at least one non-transitory computer-readable storage medium as described in Example 14, wherein the input set of machine instructions can be provided via a hardware prefetcher and an instruction decoder, the hardware prefetcher having a large look-ahead window to capture binary source code, and the instruction decoder decodes the captured binary source code into machine instructions.
[0098] Example 19 includes at least one non-transitory computer-readable storage medium as described in any one of Examples 14 to 18, wherein the sequence of instructions for converting to be offloaded includes reading a hardware table that includes a direct mapping between instructions executable by the CPU and instructions executable by the auxiliary computing device.
[0099] Example 20 includes a method of offloading instructions for execution, comprising: identifying patterns of memory access and computation instructions via a neural network based on an input set of machine instructions, determining via the neural network a sequence of instructions to be offloaded for execution by an auxiliary computing device based on the identified patterns of memory access and computation instructions, and converting the sequence of instructions to be offloaded from instructions executable by a central processing unit (CPU) to instructions executable by the auxiliary computing device.
[0100] Example 21 includes the method of Example 20, further comprising marking instructions to be offloaded in a program counter on the CPU, wherein the neural network comprises a recurrent neural network (RNN), wherein the pattern of memory access and computation instructions comprises one or more of transition probabilities or sequential dependencies between instructions of the input set of instructions, and wherein the sequence of instructions to be offloaded comprises one or more of: a recurring sequence, a sequence having an offload execution time benefit, a sequence that will result in high latency due to repeated memory operations, or a sequence that will result in a cache miss.
[0101] Example 22 includes the method of Example 21, wherein the RNN is trained via one or more of: hardware emulation with a benchmark known to increase cache misses and cause memory latency issues, direct training to detect sequences that cause long latencies for memory accesses, or embedding the RNN in a reinforcement learning agent.
[0102] Example 23 includes the method described in Example 20, further comprising, when the CPU is idle, assigning the CPU a task of a first process for executing a sequence of unloaded instructions, and assigning the auxiliary computing device a task of a second process concurrent with the first process for executing the converted unloaded instructions, if the second process is completed before the first process, accepting the execution result of the second process and terminating the first process, otherwise, if the second process is not completed before the first process, accepting the execution result of the first process and terminating the second process.
[0103] Example 24 includes the method of Example 20, wherein the input set of machine instructions may be provided via a hardware prefetcher and an instruction decoder, the hardware prefetcher having a large look-ahead window to capture binary source code, and the instruction decoder decodes the captured binary source code into machine instructions.
[0104] Example 25 includes the method of any of Examples 20 to 24, wherein translating the sequence of instructions to be offloaded comprises reading a hardware table comprising a direct mapping between instructions executable by the CPU and instructions executable by the secondary computing device.
[0105] Example 26 includes an apparatus comprising means for performing the method of any of Examples 20 to 24.
[0106] Thus, the adaptive techniques described herein provide for accelerating program execution and alleviating memory bottlenecks, particularly in multithreaded applications where there are many ongoing competing demands for CPU resources. The techniques improve the efficiency and adaptability of auxiliary computing hardware by automatically and intelligently determining when it will be faster to execute a given fragment of code on the CPU or on an IMC / CNM hardware device. In addition, the techniques provide for efficient handling of parallel processing tasks by avoiding frequent data exchanges between memory and processor cores, coupled with a significant reduction in data movement, thereby achieving high performance for auxiliary computing hardware.
[0107] Embodiments are applicable to use with all types of semiconductor integrated circuit ("IC") chips. Examples of these IC chips include, but are not limited to, processors, controllers, chipset components, programmable logic arrays (PLA), memory chips, network chips, systems on chip (SoC), SSD / NAND controller ASICs, and the like. In addition, in some of the drawings, signal conductors are represented by lines. Some may be different to indicate more constituent signal paths, with digital labels to indicate the number of constituent signal paths, and / or with arrows at one or more ends to indicate the main information flow direction. However, this should not be interpreted in a limiting manner. On the contrary, such added details may be used in conjunction with one or more exemplary embodiments to facilitate easier understanding of the circuit. Any signal line represented, whether or not it has additional information, may actually include one or more signals, which may propagate in multiple directions and may be implemented using any appropriate type of signal scheme, such as digital or analog lines implemented using differential pairs, optical fiber lines, and / or single-ended lines.
[0108] Example sizes / models / values / ranges may be given, although the embodiments are not limited thereto. As manufacturing techniques (e.g., photolithography) mature over time, it is expected that devices with smaller sizes can be manufactured. In addition, for simplicity of illustration and discussion, and in order not to obscure certain aspects of the embodiments, known power / ground connections to IC chips and other components may or may not be shown in the drawings. In addition, the arrangement may be shown in block diagram form to avoid obscuring the embodiments, and also taking into account the fact that the specific details of the implementation of such block diagram arrangements are highly dependent on the computing system in which the embodiments are implemented, that is, such specific details should be completely within the field of vision of those skilled in the art. In the case of elaborating specific details (e.g., circuits) in order to describe example embodiments, it should be clear to those skilled in the art that the embodiments can also be implemented without these specific details, or with variations of these specific details. Thus the specification should be considered illustrative, not restrictive.
[0109] The term "coupled" may be used herein to refer to any type of relationship between the components involved, whether direct or indirect, and may be applied to electrical, mechanical, fluid, optical, electromagnetic, electromechanical or other connections. In addition, unless otherwise indicated, the terms "first", "second", etc. may be used herein only to facilitate discussion without a specific time or sequential meaning.
[0110] As used in this application and in the claims, a list of items linked by the term "one or more of" may mean any combination of the listed terms. For example, the phrase "one or more of A, B, or C" may mean A; B; C; A and B; A and C; B and C; or A, B, and C.
[0111] Those skilled in the art will appreciate from the foregoing description that the broad techniques of the embodiments can be implemented in a variety of forms. Therefore, although the embodiments have been described in connection with their specific examples, the true scope of the embodiments should not be limited thereto, as other modifications will become apparent to those skilled in the art after studying the drawings, the specification, and the appended claims.
Claims
1. A system comprising: Memory; a central processing unit (CPU), configured to perform a first operation; executing circuitry within a memory in said memory; as well as and detector software for causing the second operation to be offloaded to the in-memory execution circuit after detecting a benefit of offloading the second operation to the in-memory execution circuit based on machine learning, the in-memory execution circuit being used to perform the second operation in the first execution path, and the CPU being used to perform the first operation in the second execution path.
2. The system of claim 1, wherein the second operation is based on the translated instruction.
3. The system of claim 1, wherein the memory is a dynamic random access memory.
4. The system of claim 1, wherein the detector software is to cause the offloading to be performed after detecting an execution time benefit of the offloading.
5. The system of claim 1 , wherein the detector software is to cause the offloading to be performed after a neural network recognizes a pattern of memory accesses and computational instructions in the second operation.
6. The system of claim 1, wherein the in-memory execution circuit is a single instruction multiple data (SIMD) processing unit.
7. The system of claim 1 , further comprising logic for: receiving a result of the second operation from execution circuitry in the memory; and When execution of the second operation is completed before execution of the first operation, the first operation at the CPU is terminated.
8. The system of claim 1, wherein the in-memory execution circuit performs the second operation in parallel with the CPU performing the first operation.
9. At least one computer-readable random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), flash memory, dynamic random access memory (DRAM), static random access memory (SRAM), solid-state drive (SSD), hard disk drive (HDD), or optical disk, comprising instructions for causing at least one processor to at least: performing a first operation; and After detecting a benefit of offloading a second operation to an in-memory execution circuit based on machine learning, offloading the second operation to the in-memory execution circuit, the in-memory execution circuit is used to perform the second operation in the first execution path, and the at least one processor is used to perform the first operation in the second execution path.
10. At least one computer-readable RAM, ROM, PROM, EEPROM, flash memory, DRAM, SRAM, SSD, HDD or optical disk according to claim 9, wherein the second operation is based on the translated instruction.
11. The at least one computer readable RAM, ROM, PROM, EEPROM, flash memory, DRAM, SRAM, SSD, HDD or optical disk of claim 9, wherein the in-memory execution circuit is in the DRAM.
12. At least one computer-readable RAM, ROM, PROM, EEPROM, flash memory, DRAM, SRAM, SSD, HDD or optical disk according to claim 9, wherein the instructions are used to cause the at least one processor to perform the offloading after detecting the execution time benefit of the offloading.
13. At least one computer readable RAM, ROM, PROM, EEPROM, flash memory, DRAM, SRAM, SSD, HDD or optical disk according to claim 9, wherein the instructions are used to cause the at least one processor to perform the offloading after the neural network recognizes a pattern of memory access and computation instructions in the second operation.
14. The at least one computer readable RAM, ROM, PROM, EEPROM, flash memory, DRAM, SRAM, SSD, HDD or optical disk of claim 9, wherein the in-memory execution circuit is a single instruction multiple data (SIMD) processing unit.
15. The at least one computer readable RAM, ROM, PROM, EEPROM, flash memory, DRAM, SRAM, SSD, HDD or optical disk according to claim 9, wherein the instructions are used to cause the at least one processor to: receiving a result of the second operation from execution circuitry in the memory; and When execution of the second operation is completed before execution of the first operation, the first operation at the at least one processor is terminated.
16. The at least one computer-readable RAM, ROM, PROM, EEPROM, flash memory, DRAM, SRAM, SSD, HDD or optical disk of claim 9, wherein the execution of the second operation by the in-memory execution circuit is in parallel with the execution of the first operation by the at least one processor.
17. A method comprising: Executing a first instruction via a central processing unit (CPU); and After detecting the benefit of offloading the second instruction to the in-memory execution circuit based on machine learning, the second instruction is offloaded to the in-memory execution circuit, the in-memory execution circuit is used to execute the second instruction in the first execution path, and the CPU is used to execute the first instruction in the second execution path. The method of claim 17 , wherein the second instruction is based on a translated instruction.
19. The method of claim 17, wherein the execution-in-memory circuit is in a dynamic random access memory.
20. The method of claim 17, wherein causing the offloading to be performed is after detecting an execution time benefit of the offloading.
21. The method of claim 17, wherein causing the offloading to be performed is after a neural network recognizes a pattern of memory accesses and computation instructions in the second instruction.
22. The method of claim 17, wherein the in-memory execution circuit is a single instruction multiple data (SIMD) processing unit.
23. The method of claim 17, further comprising: receiving a result of the second instruction from an execution circuit in the memory; and When execution of the second instruction is completed before execution of the first instruction, execution of the first instruction by the CPU is terminated.
24. The method of claim 17, wherein the execution of the second instruction by the in-memory execution circuit is in parallel with the execution of the first instruction by the CPU.