Atomic-level computing accelerator design methodology

By employing time-division multiplexing and near-memory computing layout in atomic-level computing accelerator design, the problems of excessive chip area and power consumption in traditional designs are solved, achieving high efficiency and low latency computing performance improvement.

CN122470122APending Publication Date: 2026-07-28GUANGDONG XINPEISEN TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGDONG XINPEISEN TECHNOLOGY CO LTD
Filing Date
2026-05-14
Publication Date
2026-07-28

AI Technical Summary

Technical Problem

Traditional atomic-level computing chip designs suffer from excessive chip area and power consumption, and parameters and intermediate results need to be frequently moved between computing units and external storage, resulting in high latency and high power consumption, causing the 'memory wall' problem.

Method used

By adopting a time-division multiplexing strategy and a near-memory computing layout, and configuring data input and output through a pipelined architecture, combined with the number of parallel computing units and the time-division multiplexing cycle, data movement is reduced, thus overcoming the storage wall bottleneck.

Benefits of technology

Significantly reduce hardware resource consumption, improve computing chip performance, reduce latency and power consumption, and achieve efficient computing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122470122A_ABST
    Figure CN122470122A_ABST
Patent Text Reader

Abstract

This application discloses a design method for an atomic-level computing accelerator, belonging to the field of atomic-level computing. The method includes: obtaining the computational characteristics and computational scale range of the target atomic-level computing algorithm to be designed; determining the number of parallel computing units and the number of time-division multiplexing cycles based on the computational characteristics and computational scale range, combined with hardware resource budget; configuring the data input and data output for each time-division multiplexing cycle based on a pipelined design architecture; and, by deploying storage units with a near-memory computing layout, completing the sub-module design of the atomic-level computing accelerator, thereby completing the atomic-level computing accelerator design, in conjunction with the number of parallel computing units and the time-division multiplexing cycle for completing the pipelined design. This application balances computational efficiency and hardware overhead by employing a time-division multiplexing strategy in the atomic-level computing accelerator design, significantly reducing hardware resource overhead, and reducing data movement through a near-memory computing layout, overcoming the memory wall bottleneck and effectively improving the performance of the computing chip.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of atomic-level computing technology, and in particular to a design method for atomic-level computing accelerators. Background Technology

[0002] Currently, atomic-level computation (such as molecular dynamics simulations and density functional theory calculations) is a key research tool in fields such as materials science, biomedicine, and new energy.

[0003] In related technologies, accelerators, as the core of dedicated circuit computing chips, typically require the design and development of atomic-level computing chips. However, in practical applications, it has been found that traditional design methods, which usually use fully parallel architectures and in-memory computing architectures, suffer from excessive chip area and power consumption. Parameters and intermediate results need to be frequently moved between computing units and external storage, leading to high latency and high power consumption, and causing problems such as the "memory wall".

[0004] In summary, the technical problems existing in the relevant technologies need to be improved. Summary of the Invention

[0005] This application provides a method for designing an atomic-level computing accelerator, which can significantly reduce hardware resource consumption, reduce data transfer, overcome the memory wall bottleneck, and effectively improve the performance of computing chips.

[0006] This application provides a method for designing an atomic-level computing accelerator, the method comprising the following steps: Obtain the computational characteristics and computational scale range of the target atomic-level computational algorithm to be designed; Based on the computational characteristics and computational scale range of the target atomic-level computing algorithm, and in conjunction with the preset hardware resource budget, the number of parallel computing units and the number of time-division multiplexing cycles of the atomic-level computing accelerator are determined. Based on a pipelined design architecture, configure the data input and data output for each time-division multiplexing cycle; By combining the number of parallel computing units in an atomic-level computing accelerator with the time-division multiplexing cycle for completing the pipeline design, and by deploying storage units with a near-memory computing layout, the design of each sub-module of the atomic-level computing accelerator is completed, thereby completing the design of the atomic-level computing accelerator corresponding to the target atomic-level computing algorithm to be designed.

[0007] Optionally, obtaining the computational characteristics and computational scale range of the target atomic-level computational algorithm to be designed includes: Determine the target atomic-level computation algorithm to be designed; The independent computational tasks in the target atomic-level computational algorithm are analyzed and identified as computational pattern characteristics. The data access pattern and computational intensity of the target atomic-level computational algorithm are also determined, thereby obtaining the computational characteristics of the target atomic-level computational algorithm. Based on the computational characteristics of the target atomic-level computational algorithm, the key computational processes of the target atomic-level computational algorithm are determined; The number of data channels, the scale of computation, and the accuracy requirements of the key computation processes of the target atomic-level computation algorithm are obtained, which are used as the computation scale range of the target atomic-level computation algorithm.

[0008] Optionally, determining the number of parallel computing units and the number of time-division multiplexing cycles of the atomic-level computing accelerator based on the computational characteristics and scale range of the target atomic-level computing algorithm, combined with a preset hardware resource budget, includes: Based on the computational characteristics and computational scale range of the target atomic-level computing algorithm, and in conjunction with the preset hardware resource budget, the number of parallel computing units of the atomic-level computing accelerator is determined. Based on the computation mode characteristics of the target atomic-level computing algorithm, the number of time-division multiplexing cycles of the atomic-level computing accelerator is determined.

[0009] Optionally, the pipelined design architecture, configuring the data input and data output for each time-division multiplexing cycle, includes: Based on a preset pipeline depth, the computing unit within the time-division multiplexing framework is designed so that in each clock cycle, the computing unit can input first data and output the calculation result of second data, wherein the second data is the data input in the previous clock cycle; Configure the number of clock cycles contained in each time-division multiplexing cycle to execute multiple computational tasks in the same data stream serially, and to process independent tasks in different data streams in parallel.

[0010] Optionally, the method further includes: Obtain the model parameters of the computation model corresponding to the target atomic-level computation algorithm, and store the model parameters in the on-chip registers and / or on-chip memory units of the atomic-level computation accelerator.

[0011] Optionally, the method further includes: The atomic-level computing accelerator is configured with a high-efficiency streaming protocol data interface so that the atomic-level computing accelerator can interact with external data modules through the high-efficiency streaming protocol data interface.

[0012] This application embodiment balances computational efficiency and hardware overhead by employing a time-division multiplexing strategy in the atomic-level computing accelerator design, significantly reducing hardware resource consumption, and reduces data movement through a near-memory computing layout, breaking through the memory wall bottleneck and effectively improving the performance of the computing chip. Attached Figure Description

[0013] Figure 1 This is a flowchart illustrating an atomic-level computing accelerator design method provided in an embodiment of this application; Figure 2 This is a schematic diagram of a time-division multiplexing strategy provided in an embodiment of this application; Figure 3 This is a schematic diagram of a pipeline design architecture provided in an embodiment of this application; Figure 4 This is a schematic diagram of a near-memory computing layout provided in an embodiment of this application; Figure 5 This is a schematic diagram of the structure of an atomic-level computing accelerator provided in an embodiment of this application. Detailed Implementation

[0014] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit it. In the following description, when referring to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with those of this application; they are merely examples of apparatuses and methods consistent with some aspects of the embodiments of this application as detailed in the appended claims.

[0015] It is understood that the terms “first,” “second,” etc., used in this application may be used herein to describe various concepts, but unless otherwise stated, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another. For example, without departing from the scope of the embodiments of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the words “if,” “when,” or “in response to a determination” as used herein may be interpreted as “when…” or “when…” or “in response to a determination.”

[0016] As used in this application, the terms "at least one", "multiple", "each", "any", etc., "at least one" includes one, two or more, "multiple" includes two or more, "each" refers to each of the corresponding multiples, and "any" refers to any one of the multiples.

[0017] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0018] Currently, atomic-level computation (such as molecular dynamics simulations and density functional theory calculations) is a key research tool in fields such as materials science, biomedicine, and new energy.

[0019] In related technologies, accelerators, as the core of dedicated circuit computing chips, typically require the design and development of atomic-level computing chips. However, in practical applications, it has been found that traditional design methods, which usually use fully parallel architectures and in-memory computing architectures, suffer from excessive chip area and power consumption. Parameters and intermediate results need to be frequently moved between computing units and external storage, leading to high latency and high power consumption, and causing problems such as the "memory wall".

[0020] In view of this, this application provides an atomic-level computing accelerator design method. By adopting a time-division multiplexing strategy in the atomic-level computing accelerator design to balance computing efficiency and hardware overhead, the hardware resource overhead is significantly reduced. Furthermore, by reducing data movement through a near-memory computing layout, the memory wall bottleneck is overcome, effectively improving the performance of the computing chip.

[0021] It should be noted that in all specific embodiments of this application, when processing data related to user identity or characteristics, such as user information, user behavior data, user historical data, and user location information, user permission or consent is obtained first. Furthermore, the collection, use, and processing of this data comply with relevant laws, regulations, and standards. In addition, when embodiments of this application require access to sensitive personal information of users, separate permission or consent from the user is obtained through pop-ups or redirection to confirmation pages. Only after obtaining the user's separate permission or consent is the necessary user-related data required for the proper functioning of these embodiments acquired.

[0022] The specific implementation methods of the embodiments of this application will be described in detail below with reference to the accompanying drawings.

[0023] Please refer to Figure 1 , Figure 1 This is a flowchart illustrating an atomic-level computing accelerator design method provided in an embodiment of this application, specifically including but not limited to steps 100 to 400.

[0024] Step 100: Obtain the computational characteristics and computational scale range of the target atomic-level computational algorithm to be designed.

[0025] In the embodiments of this application, the target atomic-level computing algorithm to be designed is first determined, and then the computing characteristics and computing scale range of the target atomic-level computing algorithm to be designed are analyzed and determined.

[0026] Furthermore, atomic-level computing algorithms are characterized by high computational intensity, large data throughput, and complex algorithm structure, such as molecular dynamics simulations, density functional theory calculations, and machine learning molecular dynamics algorithms.

[0027] In practical applications, once the target atomic-level computing algorithm to be designed is determined, the algorithm principle, computational characteristics and other feature data of the target atomic-level computing algorithm can be analyzed to determine the computational characteristics and computational scale range of the target atomic-level computing algorithm, which can then serve as the basis for subsequent time-division multiplexing design and reconfigurable design.

[0028] Specifically, as an optional implementation, obtaining the computational characteristics and computational scale range of the target atomic-level computational algorithm to be designed includes: Determine the target atomic-level computation algorithm to be designed; The independent computational tasks in the target atomic-level computational algorithm are analyzed and identified as computational pattern characteristics. The data access pattern and computational intensity of the target atomic-level computational algorithm are also determined, thereby obtaining the computational characteristics of the target atomic-level computational algorithm. Based on the computational characteristics of the target atomic-level computational algorithm, the key computational processes of the target atomic-level computational algorithm are determined; The number of data channels, the scale of computation, and the accuracy requirements of the key computation processes of the target atomic-level computation algorithm are obtained, which are used as the computation scale range of the target atomic-level computation algorithm.

[0029] In this application embodiment, by analyzing the computational characteristics of the target atomic-level computing algorithm to be designed, the computational scale range of key modules is determined, providing a basis for time-division multiplexing and reconfigurable design.

[0030] Specifically, the first step is to analyze and determine whether there are a large number of independent computational tasks that can be computed in parallel in the target atomic-level computational algorithm to be designed (for example, each central atom needs to calculate the interaction forces between itself and all neighboring atoms in a sphere centered on itself and with the truncation radius as its radius), which can be used as the computational mode characteristics of the target atomic-level computational algorithm.

[0031] Furthermore, by analyzing and determining the data access patterns of the target atomic-level computation algorithm, including data reuse rate (such as repeated reading of weight parameters) and data locality (such as neighboring atoms being concentrated only in adjacent grids), the data flow characteristics of the target atomic-level computation algorithm are determined.

[0032] Furthermore, by analyzing and identifying the core modules with the highest computational load in the target atomic-level computation algorithm and their operation types (such as matrix multiplication, accumulation, etc.), the computational density is determined. Finally, the computational pattern characteristics, data flow characteristics, and computational density are combined to form the computational characteristics of the target atomic-level computation algorithm.

[0033] Furthermore, based on the computational characteristics of the target atomic-level computational algorithm, the key computational processes within the algorithm are identified—those with a large computational load, high execution frequency, and significant impact on overall performance. For example, taking molecular dynamics simulations, key computational processes typically include: neighbor list construction (requiring calculation and filtering of the distance between each atom and its surrounding atoms); and potential function inference (calculating atomic energy and forces based on atomic coordinates and model parameters (such as neural network weights), involving numerous multiplication and addition operations).

[0034] Furthermore, the computational scale range of the target atomic-level computing algorithm is determined by the number of data channels, computational scale, and computational accuracy requirements of the key computational processes in the target atomic-level computing algorithm.

[0035] Among them, the number of data channels is the range of the number of nearest neighbor atoms that each atom needs to process; the computational scale is used to characterize the number of computational units that are executed sequentially or processed in parallel in the algorithm, including but not limited to: the number of hidden layers and the number of nodes in each layer of the potential function of the neural network, the number of self-consistent iterations in the density functional theory calculation, etc.; the computational accuracy requirement is the bit width required for numerical representation.

[0036] Therefore, by analyzing the target atomic-level computing algorithm to be designed and determining the computing characteristics and scale range, it is possible to identify computing modules suitable for time-division multiplexing, and then determine how to allocate hardware resources.

[0037] Step 200: Based on the computational characteristics and computational scale range of the target atomic-level computing algorithm, and in conjunction with the preset hardware resource budget, determine the number of parallel computing units and the number of time-division multiplexing cycles of the atomic-level computing accelerator.

[0038] In this embodiment, based on a preset hardware resource budget and the computational characteristics and scale range of the target atomic-level computation algorithm determined by the above steps, the number of parallel computing units is determined. Simultaneously, based on the computational mode characteristics within the computational characteristics, the number of time-division multiplexing cycles is determined, allowing different computational modes to be implemented by adjusting the number of time-division multiplexing cycles, thereby maintaining a high reuse rate of the hardware units.

[0039] Specifically, as an optional implementation, determining the number of parallel computing units and the number of time-division multiplexing cycles of the atomic-level computing accelerator based on the computational characteristics and scale range of the target atomic-level computing algorithm, combined with a preset hardware resource budget, includes: Based on the computational characteristics and computational scale range of the target atomic-level computing algorithm, and in conjunction with the preset hardware resource budget, the number of parallel computing units of the atomic-level computing accelerator is determined. Based on the computation mode characteristics of the target atomic-level computing algorithm, the number of time-division multiplexing cycles of the atomic-level computing accelerator is determined.

[0040] In this embodiment of the application, based on the computational characteristics and computational scale range of the target atomic-level computing algorithm, and in combination with the preset hardware resource budget, the number of parallel computing units of the atomic-level computing accelerator is determined. Furthermore, based on the computational mode characteristics of the target atomic-level computing algorithm, the number of time-division multiplexing cycles of the atomic-level computing accelerator is determined.

[0041] Time Division Multiplexing (TDM) refers to dividing time into several time-division multiplexing frames of equal length. Multiple independent tasks are processed in time-division within each frame, allowing multiple computing tasks to share the same set of hardware resources. This reduces the number of parallel computing units to a lower level, significantly reducing chip area and power consumption.

[0042] For example, please refer to Figure 2 , Figure 2 This is a schematic diagram of a time-division multiplexing strategy provided in an embodiment of this application. In the time-division multiplexing strategy, the same set of computing hardware repeatedly serves multiple computing objects in different time slices. Each central atom corresponds to a computing task. Multiple central atoms take turns using the same hardware circuit in time order and process different computing tasks in each time-division multiplexing frame. Thus, in atomic-level computing, for each central atom, time-division multiplexing allows the same computing core to process the tasks of multiple central atoms in sequence, enabling multiple computing tasks to share the same set of hardware resources.

[0043] In practical applications, while keeping the number of parallel computing units constant, the time-division period parameters can be configured according to the computing mode characteristics of the target atomic-level computing algorithm to achieve switching between different modes. Most functional modules are implemented through the same set of circuits, so that in a stable output state, a computing task can be completed after each time-division multiplexing frame, and the data throughput is inversely proportional to the clock cycle.

[0044] Of course, it is understandable that multiple-choice paths can be introduced in a few modules that require different logical operations.

[0045] Step 300: Based on the pipelined design architecture, configure the data input and data output for each time-division multiplexing cycle.

[0046] In this embodiment, by introducing a pipeline within the time-division multiplexing cycle framework and configuring the data input and output of each time-division multiplexing cycle, it is ensured that there is data inflow and result outflow in each clock cycle. This works in conjunction with the time-division multiplexing strategy to form an efficient computing flow, enabling the computing process to proceed continuously in time.

[0047] Therefore, this application coordinates time-division multiplexing strategy with pipeline to ensure that computing units are in working state in every cycle. This not only guarantees high reuse rate of hardware resources, but also further improves the continuity and high throughput of data processing, thereby enabling atomic-level computing accelerators to achieve the highest possible computing performance with limited hardware resources.

[0048] Specifically, as an optional implementation, the pipelined design architecture, configuring the data input and data output for each time-division multiplexing cycle, includes: Based on a preset pipeline depth, the parallel computing unit within the time-division multiplexing framework is designed so that in each clock cycle, the parallel computing unit can input first data and output the calculation result of second data, wherein the second data is the data input in the previous clock cycle. Configure the number of clock cycles contained in each time-division multiplexing cycle to execute multiple computational tasks in the same data stream serially, and to process independent tasks in different data streams in parallel.

[0049] In the embodiments of this application, please refer to Figure 3 , Figure 3 This is a schematic diagram of a pipelined design architecture provided in an embodiment of this application. The pipelined design architecture refers to breaking down a computational task into multiple smaller tasks and processing them in parallel, thereby dividing the computational process into multiple stages, each of which processes different data simultaneously.

[0050] Furthermore, a pre-defined single-stage pipeline architecture (pipeline depth of 1) can be adopted, where new data is input in each cycle and the result is directly output in the next cycle.

[0051] For example, such as Figure 3 As shown, the computation of a neuron is implemented in machine learning molecular dynamics calculations. y = ( wx + b Taking this as an example, different computing tasks are introduced in each time-division multiplexing cycle, and data input and data output are configured for each time-division multiplexing cycle, thereby realizing the computing process of serially executing the same data stream and parallel executing different data streams using the same set of circuit computing units.

[0052] Therefore, this application adopts pipelined computing and time-division multiplexing strategies to balance computing efficiency and hardware overhead, in order to meet the large computational demands of atomic-level computing. This reduces hardware resource overhead while improving data throughput, achieving a balance between high efficiency and low power consumption.

[0053] Step 400: Combining the number of parallel computing units in the atomic-level computing accelerator with the time-division multiplexing cycle for completing the pipeline design, the design of each sub-module of the atomic-level computing accelerator is completed by deploying storage units with a near-memory computing layout, thereby completing the design of the atomic-level computing accelerator corresponding to the target atomic-level computing algorithm to be designed.

[0054] In this embodiment, after determining the number of parallel computing units and the number of time-division multiplexing cycles of the atomic-level computing accelerator, and introducing a pipeline design for each time-division multiplexing cycle, the design of each sub-module of the atomic-level computing accelerator is completed by introducing storage units with a near-memory computing layout for key computing modules, thereby completing the design of the atomic-level computing accelerator corresponding to the target atomic-level computing algorithm to be designed.

[0055] For example, please refer to Figure 4 , Figure 4 This is a schematic diagram of a near-memory computing layout provided in an embodiment of this application. When designing each submodule, the storage units (such as SRAM) required for the submodule's functionality can be instantiated within the corresponding submodule. Compared to existing technologies that deploy storage units at the top level, this application, by deploying storage units using a near-memory computing layout, allows each submodule to independently read the corresponding storage unit timing library file and complete the synthesis calculation during the logic synthesis stage. Furthermore, based on the data flow direction within different storage units of the same module, the distribution area and arrangement order of the storage units can be pre-determined, guiding EDA tools to place related logic circuits near the corresponding storage units. For example, in a multi-layer network structure for machine learning molecular dynamics calculations, the storage unit groups of each layer are arranged according to the data flow order, and multiple small-capacity storage units within each storage unit group are locally dispersed, leaving layout space for logic circuits. This achieves physical proximity between storage units and computing units, reducing data transmission latency.

[0056] In practical applications, please refer to Figure 5 , Figure 5 This is a schematic diagram of the structure of an atomic-level computing accelerator provided in an embodiment of this application. The atomic-level computing accelerator is provided with multiple sub-modules for completing the target atomic-level computing algorithm and is equipped with a storage unit that adopts a near-memory computing layout.

[0057] Furthermore, the atomic-level computing accelerator is also equipped with a time-division multiplexing control module, which is used to determine the number of time-division multiplexing cycles of the atomic-level computing accelerator according to the computing mode characteristics of the target atomic-level computing algorithm. By adjusting the time-division cycle parameters, the switching between different modes can be realized, and the number of parallel computing units can be kept constant.

[0058] Therefore, this application reduces data movement and breaks through the "storage wall" bottleneck by using a module-level near-memory computing storage unit.

[0059] Optionally, the method further includes: Obtain the model parameters of the computation model corresponding to the target atomic-level computation algorithm, and store the model parameters in the on-chip registers and / or on-chip memory units of the atomic-level computation accelerator.

[0060] In this embodiment of the application, based on module-level near-memory computation, the model parameters of the computation model corresponding to the target atomic-level computation algorithm can be further stored in the on-chip registers and / or on-chip memory units of the atomic-level computation accelerator, thereby realizing system-level near-memory computation.

[0061] Specifically, by storing all the parameters that need to be loaded and the cached intermediate results in on-chip registers or memory units, the frequent communication between the computing unit and off-chip memory is reduced. The model parameters are only loaded from the outside once when the chip starts up, and are read directly from the on-chip in subsequent calculations.

[0062] Furthermore, by determining a matching computational model based on the computational characteristics and scale range of the target atomic-level computational algorithm to be designed, and by calling and loading the relevant model parameters, the frequent data transfers caused by the separation of storage and computation in the traditional von Neumann architecture can be overcome, effectively improving the performance of the computing chip.

[0063] For example, please refer to Figure 5 Atom-level computing accelerators can also be equipped with reconfigurable configuration modules. Through these modules, all parameters updated with the model (such as weights and biases) can be stored in configurable registers and memory units. By configuring the registers, the computing mode (time-division multiplexing cycle number and multiple-choice path) can be selected, improving the applicability of the computing chip and enabling it to adapt to different computing tasks.

[0064] Therefore, this application can switch between different computing modes by adjusting the number of time-division multiplexing cycles, keep the number of parallel computing units constant, use most hardware circuits universally, and only a few modules use multiple-choice paths.

[0065] Specifically, as an optional implementation, the method further includes: The atomic-level computing accelerator is configured with a high-efficiency streaming protocol data interface so that the atomic-level computing accelerator can interact with external data modules through the high-efficiency streaming protocol data interface.

[0066] In the embodiments of this application, please refer to Figure 5 Furthermore, it can be configured with a high-efficiency streaming protocol data interface for atomic-level computing accelerators, enabling atomic-level computing accelerators to interact with external data modules through the high-efficiency streaming protocol data interface.

[0067] Among them, the high-efficiency streaming protocol data interface supports burst transmission, flow control, and multi-stream multiplexing. It can continuously transmit multiple data packets after a single handshake, improving throughput. It also achieves flow control through feedback signals to ensure data transmission reliability. Combined with identifiers to mark different data packet types, it allows multiple data streams to share the same physical channel, reducing hardware resource overhead and enabling efficient data interaction between the kernel and the outside world.

[0068] Therefore, compared with existing atomic-level computing accelerator design methods, this application significantly reduces hardware resource requirements through time-division multiplexing, significantly reduces data transfer overhead through near-memory computing, and can adjust the number of time-division multiplexing cycles to adapt to various application scenarios, enabling the computing chip to efficiently process different computing tasks, while keeping the area and power consumption controllable, thus providing a feasible solution for large-scale deployment.

[0069] This application provides an atomic-level computing accelerator design method that balances computing efficiency and hardware overhead by employing a time-division multiplexing strategy in the atomic-level computing accelerator design, significantly reducing hardware resource overhead, and reducing data movement through a near-memory computing layout, thereby overcoming the memory wall bottleneck and effectively improving the performance of the computing chip.

[0070] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0071] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.

[0072] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0073] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0074] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0075] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0076] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0077] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0078] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.

Claims

1. A method for designing an atomic-level computing accelerator, characterized in that, The method includes the following steps: Obtain the computational characteristics and computational scale range of the target atomic-level computational algorithm to be designed; Based on the computational characteristics and computational scale range of the target atomic-level computing algorithm, and in conjunction with the preset hardware resource budget, the number of parallel computing units and the number of time-division multiplexing cycles of the atomic-level computing accelerator are determined. Based on a pipelined design architecture, configure the data input and data output for each time-division multiplexing cycle; By combining the number of parallel computing units in an atomic-level computing accelerator with the time-division multiplexing cycle for completing the pipeline design, and by deploying storage units with a near-memory computing layout, the design of each sub-module of the atomic-level computing accelerator is completed, thereby completing the design of the atomic-level computing accelerator corresponding to the target atomic-level computing algorithm to be designed.

2. The method according to claim 1, characterized in that, The process of obtaining the computational characteristics and computational scale range of the target atomic-level computational algorithm to be designed includes: Determine the target atomic-level computation algorithm to be designed; The independent computational tasks in the target atomic-level computational algorithm are analyzed and identified as computational pattern characteristics. The data access pattern and computational intensity of the target atomic-level computational algorithm are also determined, thereby obtaining the computational characteristics of the target atomic-level computational algorithm. Based on the computational characteristics of the target atomic-level computational algorithm, the key computational processes of the target atomic-level computational algorithm are determined; The number of data channels, the scale of computation, and the accuracy requirements of the key computation processes of the target atomic-level computation algorithm are obtained, which are used as the computation scale range of the target atomic-level computation algorithm.

3. The method according to claim 1, characterized in that, Based on the computational characteristics and scale range of the target atomic-level computing algorithm, and in conjunction with a preset hardware resource budget, the number of parallel computing units and the number of time-division multiplexing cycles of the atomic-level computing accelerator are determined, including: Based on the computational characteristics and computational scale range of the target atomic-level computing algorithm, and in conjunction with the preset hardware resource budget, the number of parallel computing units of the atomic-level computing accelerator is determined. Based on the computation mode characteristics of the target atomic-level computing algorithm, the number of time-division multiplexing cycles of the atomic-level computing accelerator is determined.

4. The method according to claim 1, characterized in that, The pipeline-based design architecture configures the data input and output for each time-division multiplexing cycle, including: Based on a preset pipeline depth, the parallel computing unit within the time-division multiplexing framework is designed so that in each clock cycle, the parallel computing unit can input first data and output the calculation result of second data, wherein the second data is the data input in the previous clock cycle. Configure the number of clock cycles contained in each time-division multiplexing cycle to execute multiple computational tasks in the same data stream serially, and to process independent tasks in different data streams in parallel.

5. The method according to claim 1, characterized in that, The method further includes: Obtain the model parameters of the computation model corresponding to the target atomic-level computation algorithm, and store the model parameters in the on-chip registers and / or on-chip memory units of the atomic-level computation accelerator.

6. The method according to claim 1, characterized in that, The method further includes: The atomic-level computing accelerator is configured with a high-efficiency streaming protocol data interface so that the atomic-level computing accelerator can interact with external data modules through the high-efficiency streaming protocol data interface.