Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

69 results about "Jazelle" patented technology

Jazelle DBX (Direct Bytecode eXecution) is an extension that allows some ARM processors to execute Java bytecode in hardware as a third execution state alongside the existing ARM and Thumb modes. Jazelle functionality was specified in the ARMv5TEJ architecture and the first processor with Jazelle technology was the ARM926EJ-S. Jazelle is denoted by a "J" appended to the CPU name, except for post-v5 cores where it is required (albeit only in trivial form) for architecture conformance.

End side model reasoning method and device based on RWKV architecture, electronic equipment and storage medium

The invention provides an end side model reasoning method and device based on an RWKV architecture, electronic equipment and a storage medium, and the method comprises the steps: obtaining a target input request of a target object, and converting the target input request into target model input data; loading a historical reasoning state corresponding to the target input request in a preset state storage space; determining a corresponding RWKV core operator according to the hardware platform type of the terminal equipment; based on the RWKV core operator, performing reasoning calculation on the target model input data and the historical reasoning state to obtain an output token sequence; wherein in the reasoning calculation process, the real-time reasoning state of the large language model is stored in a preset state accelerator memory for multiplexing; converting the output token sequence into a text format and outputting the output token sequence; and updating the historical reasoning state according to the real-time reasoning state after reasoning calculation. According to the method, calculation optimization and hardware acceleration can be carried out on the large language model of the RWKV architecture, so that the reasoning performance of the RWKV architecture model is improved on the end side.
Owner:SHENZHEN YUANSHI INTELLIGENT CO LTD

Self-optimizing and self-programming computing systems: a combined compiler, complex networks, and machine learning approach

A self-optimizing and self-programming computing system (SOSPCS) design framework that achieves both programmability and flexibility and exploits computing heterogeneity [e.g., CPUs, GPUs, and hardware accelerators (HWAs)] is provided. First, at compile time, a task pool consisting of hybrid tasks with different processing element (PE) affinities according to target applications is formed. Tasks preferred to be executed on GPUs or accelerators are detected from target applications by neural networks. Tasks suitable to run on CPUs are formed by community detection to minimize data movement overhead. Next, a distributed reinforcement learning-based approach is used at runtime to allow agents to map the tasks onto the network-on-chip-based heterogeneous PEs by learning an optimal policy based on Q values in the environment.
Owner:UNIV OF SOUTHERN CALIFORNIA

Acceleration method for accelerating multiplication and addition operation in large language model and hardware accelerator

The invention relates to the technical field of acceleration operation, in particular to an acceleration method for accelerating multiplication and addition operation in a large language model and a hardware accelerator, and the method comprises the following steps: dividing data in the form of m floating-point numbers into a floating-point number block, m being a preset positive integer; performing block floating point coding on the floating-point number block to obtain a shared index and a mantissa of the floating-point number block, and expanding bit widths of the shared index and the mantissa; performing multiplication and addition calculation based on the expanded sharing index and mantissa; the multiply-add calculation result is decoded and restored into a floating-point number form, and a decoding result is obtained; according to the acceleration method for accelerating the multiplication and addition operation in the large language model and the hardware accelerator, relatively high calculation precision can be kept.
Owner:GUANGDONG UNIV OF TECH

Instruction-level simulation and performance modeling system for parallel computing architecture

The invention provides an instruction-level simulation and performance modeling system for a parallel computing architecture, and belongs to the technical field of computer architecture and simulation verification, and the system comprises an instruction modeling layer which is used for analyzing and executing an intermediate instruction set defined by the architecture; the scheduling execution layer is used for simulating a multi-thread and multi-core parallel execution process; the storage access layer is used for constructing a hierarchical storage access and bandwidth and delay model; and the performance analysis layer is used for collecting and counting key indexes such as an execution period, an instruction utilization rate and memory access delay, and realizing accurate performance modeling of the parallel architecture. According to the method, the performance bottleneck of the design scheme can be rapidly evaluated in the early stage of architecture design, the simulation speed is high, the module configurability is high, the modeling precision is adjustable, and the method is suitable for functional verification, micro-architecture exploration and compiler performance analysis of parallel computing architectures, accelerator chips, heterogeneous multi-core processors and the like.
Owner:YUANQIXIN (SHANDONG) SEMICONDUCTOR TECHNOLOGY CO LTD

Compilation optimization device and method supporting Page Attention and medium

The invention discloses a compiling optimization device and method supporting Page Attention and a medium, and belongs to the field of compiling optimization devices.The device comprises a multi-core AI accelerator, and each calculation core comprises a tensor component, a vector component, a DMA component and an IMM component; the compiling optimization device is optimized through the following modules: (a) a KV Cache management module, which is used for storing page table items corresponding to different KV Cache blocks according to a batch sequence and generating standardized page table items; (b) a DMA control module which supports a linear reading and remapping mode, carries discrete KV Cache blocks from an external memory to an on-chip weight buffer area in a segmented manner according to an index table, and hides data loading delay by adopting a double-buffer technology; and (c) a multi-batch reasoning scheduling module which dynamically allocates computing resources according to the data scale in a code stage. According to the invention, through deep cooperation of a hardware architecture and functional logic, the problems of low KV Cache management efficiency, complex multi-batch dynamic scheduling and insufficient computing resource utilization rate on an ASIC platform are solved.
Owner:BEIJING YIXIN YIYU MICROELECTRONICS TECH CO LTD

Softmax function hardware accelerator based on dynamic pruning and compression lookup table

The invention provides a softmax function hardware accelerator based on dynamic pruning and compression of a lookup table, which retains an input value of an interval near a maximum value through a dynamic pruning method, reduces a to-be-processed data volume and compresses the size of the lookup table, then calculates an e index by using the compressed lookup table, and calculates a reciprocal of summation by using a dynamic Newton iteration method, so as to obtain the softmax function hardware accelerator. And the division method is replaced by multiplication and shift calculation. According to the method, high precision of the softmax function is guaranteed, meanwhile, the calculation amount is remarkably reduced, the hardware overhead is reduced, and the calculation speed of the softmax function is increased.
Owner:SOUTHEAST UNIV

Elliptic curve digital signature verification accelerator based on ARM and FPGA integrated heterogeneous platform

The invention relates to the technical field of integrated circuits, in particular to an elliptic curve digital signature verification accelerator based on an ARM and FPGA integrated heterogeneous platform, which comprises an ARM processor used for executing a software program and connected with the elliptic curve digital signature verification accelerator on the FPGA through an AXI 4 bus interface; wherein the signature verification accelerator comprises a Hash calculation module, a large number reduction module, a public key decoding module and a double scalar multiplication module; the hash calculation module and the large number reduction module calculate and simplify a source operand; the public key decoding module converts operands into coordinates on an elliptic curve; and the double-scalar multiplication module verifies the reliability of the signature and returns a calculation result to the ARM processor, so that a verification task of one-time digital signature is realized. According to the system, the calculation speed of digital signature verification can be accelerated through hardware under limited hardware resource consumption, and the calculation burden of the ARM processor is relieved.
Owner:FUDAN UNIVERSITY

Heterogeneous accelerator uniform interface firmware system and working method thereof

The invention provides a heterogeneous accelerator unified interface firmware system and a working method thereof, and relates to the technical field of computers, the system comprises four modules: a unified interface module which provides a standardized programming interface, shields underlying hardware differences such as a GPU and an NPU, and supports programming models such as OpenCL; the fine-grained task scheduling module is used for decomposing tasks and reasonably distributing the tasks according to task requirements, hardware loads and complexity so as to realize load balancing; the resource management module virtualizes resources of the computing core and the like, and dynamically optimizes multi-task resource allocation; and the expansion support module is used for assisting the integration and drive maintenance of the novel accelerator through a plug-in interface. According to the method, the problems of complex programming, unreasonable task scheduling, low resource utilization efficiency and the like of the heterogeneous accelerator are solved, and the overall performance and development convenience of a heterogeneous computing system are improved.
Owner:HEFEI SUMICROELECTRONICS TECH CO LTD

Large Language Model-Based Inference Acceleration Method, Medium, And Device

A large language model (LLM)-based inference acceleration method, a medium, and a device are disclosed. The method includes: determining LLMs deployed on a plurality of first hardware accelerators, multiple submodels obtained through segmenting the LLMs and respectively deployed on a plurality of second hardware accelerators, and input data of a to-be-executed task; performing first processing on the input data through the plurality of LLMs to obtain a first processing result; performing second processing on the first processing result through the multiple submodels to obtain a second processing result; and determining an execution result of the to-be-executed task in response to the second processing result meeting a stop inference condition.
Owner:XG TECHNOLOGIES PTE LTD

Multimodal large language model-oriented compiling method and system for generating hardware accelerator executable codes, terminal and storage medium

The invention relates to the technical field of data processing, and discloses a compiling method and system for generating hardware accelerator executable codes facing a multi-mode large language model, a terminal and a storage medium. The method comprises the following steps: constructing a computational graph according to the structure of a to-be-deployed large language model and operators supported by a hardware platform; traversing the computational graph to deduce output tensor type information of each operator, and verifying the type consistency between adjacent operators; based on the verified type information, executing static memory allocation to determine a static storage address of the weight data, and executing dynamic memory allocation to determine a dynamic storage address of the state data; then combining the verified type information and the static and dynamic storage addresses to generate an intermediate representation instruction sequence for configuring a hardware acceleration platform register; and finally, converting the intermediate representation instruction sequence into a target code file which can be directly loaded and executed by a hardware acceleration platform. The compiling efficiency, the deployment reliability and the resource utilization rate are improved.
Owner:SHENZHEN MAITEXIN TECH CO LTD

Heterogenous accelerators for efficient generative LLM inference using phase splitting

A system and method for splitting a prompt and token generation phase in a generative large language model (LLM) inference onto separate virtual machines (VMs) is provided. Two separate pools of VMs for prompt and token processing are maintained. The VMs in each of the pools are pre-loaded with a model of choice. A scheduler allocates an inference to a prompt VM from a pool of prompt VMs and a token VM from a pool of token VMs. Context generated from layers of the generative LLM during the prompt computation is saved in a key-value (KV) cache that is transferred from the prompt VM to token VM as it is used for all the future token generation iterations.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Method and system for executing virtual instruction set

The present invention discloses a method for executing a virtual instruction set to build an operator library to realize a deep learning accelerator and solve the problem of cross-generation compatibility. The method provides a virtual instruction set execution unit, that is, a virtual instruction set, and then the compiler compiles the virtual instruction set into vector instructions and / or special instructions, and finally the virtual instruction set execution unit executes the vector instructions and / or special instructions. The virtual instruction set execution unit is used to execute the virtual instruction set, which includes a general vector calculation unit and a special instruction set calculation unit. The general vector calculation unit is used to execute the vector instruction set, and the special instruction set calculation unit is used to execute the special instruction set. The virtual instruction set includes a first C++ template and a second C++ template. The first C++ template defines a first C++ function. The first C++ function can be compiled by the compiler into a vector instruction in the vector instruction set. The second C++ template defines a second C++ function. The second C++ function can be compiled by the compiler into a special instruction in the special instruction set.
Owner:奕行智能科技(广州)有限公司

Application programming interface to indicate accelerator error handlers

Apparatuses, systems, and techniques to execute one or more application programming interfaces (APIs) to perform one or more operations for one or more accelerators within a heterogeneous processor. In at least one embodiment, one or more processors are to perform one or more instructions in response to one or more APIs to indicate one or more functions to be performed in response to one or more errors from one or more accelerators within a heterogeneous processor.
Owner:NVIDIA CORP

A Reinforcement Learning Parallel Processing Accelerator, an Acceleration Method and an Electronic Device

The present invention discloses a reinforcement learning parallel processing accelerator, an acceleration method and an electronic device, which relate to the technical field of accelerators. The controller writes the current view features, at least two groups of batch historical view features and instruction sequences into corresponding memories respectively. The instruction loading and distribution component reads and decodes the instruction sequences and distributes parameters and start calculation instructions to the calculation components. The data loading control component selects the required feature data from the memory according to the parameters of each calculation layer and loads it into the corresponding feature cache. After receiving the start calculation instruction, the calculation components read multiple view feature data simultaneously and perform parallel processing in combination with the parameters, which can greatly improve the data processing efficiency. Through the mutual cooperation of the above components, the characteristics of the reinforcement learning model can be deeply analyzed, and according to different reinforcement learning tasks and data characteristics, the parameters and processing flow can be flexibly adjusted, so as to improve the processing efficiency of the reinforcement learning model, reduce the resource utilization rate of the accelerator at the same time, and have a fast response speed.
Owner:LANGCHAO ELECTRONIC INFORMATION IND CO LTD

VTA acceleration core software and hardware cooperative tuning method and system

The invention discloses a VTA acceleration core software and hardware collaborative tuning method and system, and the method comprises the steps: initializing a hardware design parameter of a multifunctional tensor accelerator to be a default minimum value, and reading a resource utilization rate of an EDA comprehensive report; hardware parameters are dynamically adjusted under the condition that hardware design constraint conditions are met, and the FPGA resource utilization rate is maximized; a tensor shape of a neural network operator is converted according to hardware parameters, and a VTA calculation unit is adapted; analyzing software parameters, and calculating buffer area requirements; verifying the legality of the software parameter according to the buffer area requirement, and judging the software parameter as an illegal software parameter when the buffer area requirement does not meet a set condition; illegal software parameters are triggered to be regenerated, and optimization interruption caused by resource out-of-limit is avoided; a multifunctional tensor accelerator drive program is recompiled for legal hardware parameters; and iteratively executing until the collaborative optimization of the software and hardware parameters reaches the standard. Through a double closed-loop mechanism of hardware parameter dynamic adjustment and software parameter legality verification, the method is suitable for deployment optimization of the compute-intensive model in the power edge equipment.
Owner:CHINA ELECTRIC POWER RESEARCH INSTITUTE CO LTD +2

A large language model quantization method based on orthogonal characteristics and accelerator architecture

The application belongs to the technical field of large language model quantization, and particularly relates to a large language model quantization method based on orthogonal characteristics and an accelerator architecture. The quantization method divides the activation tensor of the large language model into multiple column blocks, and allocates an FP4 quantization format to the entire activation tensor with the column block as the granularity. The concept of the column block is defined as follows: the matrix of the activation tensor is divided into multiple segments with the same number of elements, wherein each element in the segment is arranged continuously in the same row in the first dimension of the matrix, and arranged in multiple continuous columns in the second dimension; the column block includes multiple columns in the second dimension, and the number of columns in each column block is consistent with the number of elements in the segment. The application overcomes the defects existing in the existing large language model grouping quantization technology, and solves the contradiction between the precision of the large language model and the hardware efficiency.
Owner:NORTHWESTERN POLYTECHNICAL UNIV

Fluid pulsation accelerator architecture for reinforcement learning of large language model

The invention discloses a fluid pulsation accelerator architecture for reinforcement learning of a large language model, and belongs to the technical field of artificial intelligence hardware acceleration. The system aims at solving the problems that when a current universal processor executes an RLHF working load, an instruction driving normal form is not matched with data flow calculation, fixed parallel granularity cannot adapt to a dynamic load, and the efficiency is low due to the fact that a framework does not sense data statistical characteristics. The core of the system is a liquid systolic array calculation fabric capable of being dynamically reconstructed, execution is triggered through data flow, and kernel boundaries and instruction overhead are eliminated. According to the system, a global lookup table subsystem is fused, and nonlinear calculation is optimized through merging query and parallel lookup by using data distribution prior; and a dynamic scheduling unit is configured, and elastic parallelism is realized by adopting a resource allocation algorithm supporting work stealing. According to the method, the throughput and the energy efficiency of RLHF training and a large language model reasoning stage can be remarkably improved, and a new design normal form is provided for a next-generation AI special computing architecture.
Owner:BEIJING UNIV OF CHEM TECH

Accelerating a fully homomorphic encryption (FHE) operation with an on-chip systolic array

Provided are techniques for accelerating a Fully Homomorphic Encryption (FHE) operation with an on-chip systolic array. A computer processing chip comprises an Artificial Intelligence (AI) accelerator comprising a direct memory access and a systolic array, a Level 3 (L3) cache connected to the AI accelerator, and a core connected to the AI accelerator and the L3 cache. The AI accelerator receives AI accelerator code from the core, where the AI accelerator code comprises new instructions, where the systolic array executes the new instructions using first data by executing a BMUL instruction to perform multiplication and generate first results, a BSUB instruction to perform subtraction using the first results to generate second results, and a BADDSUB instruction to perform modular correction on the second results to generate final results, and where the direct memory access prefetches second data for the systolic array.
Owner:INTERNATIONAL BUSINESS MACHINE CORPORATION

RISC-V virtual prototype platform based on Crypto-VP architecture

The invention relates to an RISC-V virtual prototype platform based on Crypto-VP architecture, which is characterized in that a CPU (Central Processing Unit) kernel in the RISC-V virtual prototype platform is respectively connected with a debugging and monitoring interface controller and a coprocessor model, and is connected with a transaction-level modeling 2.0 standard bus through a transaction-level modeling 2.0 standard bus interface; one end, far away from the CPU kernel, of the debugging and monitoring interface controller is connected with the main memory; the local interrupt controller and the platform-level interrupt controller are connected to a transaction-level modeling 2.0 standard bus through a transaction-level modeling 2.0 standard interface; the accelerator model is connected with a transaction-level modeling 2.0 standard bus through a transaction-level modeling 2.0 standard interface and is allocated to a specified address range of the CPU core; one end of the executable and linkable format loader obtains executable and linkable format files of the RISC-V architecture generated by the external GNU compiler set, and the other end of the executable and linkable format loader is connected with the main memory; and each module is integrated with a power consumption statistics and visualization module. According to the RISC-V virtual prototype platform, the simulation and exploration speed is greatly improved.
Owner:GUANGDONG UNIV OF TECH +1

A self-isomorphic mapping hardware accelerator and acceleration method based on full homomorphic encryption

The application discloses a hardware accelerator and an acceleration method based on a full homomorphic encryption isomorphism mapping, and the accelerator comprises a polynomial storage area, a Tag and Sel reloading unit, a read address calculation unit, a permutation network and a cache unit. The method comprises the following steps: according to the architecture parallelism configured during compilation and the maximum polynomial degree to be supported, after compilation, the isomorphism mapping of the polynomial can be completed according to the actual polynomial degree and the isomorphism parameter of the input. The embodiment of the application can reduce the domain conversion overhead and the overhead of the bit inversion arrangement coefficient, and reduce the calculation burden of the homomorphic rotation. The application can be widely applied to the field of homomorphic encryption technology.
Owner:SUN YAT SEN UNIV

Machine learning model updates to ML accelerators

Examples herein describe a peripheral I / O device with a hybrid gateway that permits the device to have both I / O and coherent domains. As a result, the compute resources in the coherent domain of the peripheral I / O device can communicate with the host in a similar manner as CPU-to-CPU communication in the host. The dual domains in the peripheral I / O device can be leveraged for machine learning (ML) applications. While an I / O device can be used as an ML accelerator, these accelerators previously only used an I / O domain. In the embodiments herein, compute resources can be split between the I / O domain and the coherent domain where a ML engine is in the I / O domain and a ML model is in the coherent domain. An advantage of doing so is that the ML model can be coherently updated using a reference ML model stored in the host.
Owner:XILINX INC

Hardware-optimized matrix multiplication operations for large language models

A processing system configured to implement a large language model (LLM) includes an accelerator unit (AU) having hardware configured to perform matrix multiplication operations for the LLM using sets of predetermined matrix dimensions. Further, to help optimize the LLM for the processing system, the processing system includes a processor that modifies one or more matrix multiplication operations of the LLM based the sets of predetermined matrix dimensions supported by the hardware of the AU. The processor then recompiles the LLM using the modified multiplication operations and implements the recompiled LLM.
Owner:XILINX INC

Heterogeneous functional processing architectures and methods to design and fabricate the same

A heterogeneous architecture system and method to design and fabricate the same. The heterogeneous architecture system includes at least one processing unit, a plurality of accelerator units, a memory subsystem connected to the at least one processing unit and the plurality of accelerator units, and at least one function arbiter connected to the at least one processing unit and the plurality of accelerator units.In certain embodiment, the heterogeneous architecture system integrates various types of processing components, each optimized for specific tasks, allowing for a more efficient execution of diverse workloads.
Owner:HUI RONALD CHI CHUN

Apparatus and method for an attack-resistant hardware accelerator

PendingUS20260189400A1MultiplexingJazelle
Apparatus and method for an attack resistant hardware accelerator. For example, one embodiment comprises an attack-resistant masked HMAC-SHA-2 accelerator with: (a) multiplexer-based implementations of masked choose (Ch) logic and majority (Maj) logic with implicit masked-domain non-linear computations; (b) Secure Binary to arithmetic (BtoA) conversion logic integrated within a masked carry-save-adder gate; and (c) an area-efficient arithmetic-to-Binary (AtoB) mask conversion circuit using a secure masked sparse-tree adder.
Owner:INTEL CORP

An IBConvert program module for an accelerator excitation curve conversion system

The present invention relates to the technical field of accelerators, and particularly refers to an IBConvert program module for an accelerator excitation curve conversion system; the entire system includes two parts: an IBConvert program module and an EpicsDBGenerator program module. Among them, the IBConvert program is a standard EPICS software IOC that needs to run continuously to achieve real-time conversion of I and B; the EpicsDBGenerator program module is used to generate the DB files required by the IBConvert program. When there are changes in the number of power supplies, names, and excitation curve fitting coefficients, this program needs to be run to generate new DB files; the described IBConvert program module is designed according to EPICS specifications and mainly includes the design of EPICS records and the design of functions for processing waveform data; the system of the present invention can be applied to large scientific installations, has low requirements for equipment, strong compatibility, is simple to install and use, and can simultaneously achieve real-time conversion of DC power supplies, pulsed power supplies with preset values, pulsed power supplies for bump tracks without preset values, and pulsed power supplies with floating power supplies.
Owner:INST OF HIGH ENERGY PHYSICS CHINESE ACAD OF SCI

Pipelined horizontal parallelism for large language models

A disclosed computer-implemented method may include generating, via a hardware accelerator included in a plurality of hardware accelerators that includes the hardware accelerator and at least one additional hardware accelerator, a first result tensor segment by executing a tensor operation on a first activation tensor segment included in an activation tensor. The method may also include executing, via the hardware accelerator, a collective communication operation with the at least one additional hardware accelerator as to the first result tensor segment and, during execution of the collective communication operation with the at least one additional hardware accelerator, generating, via the hardware accelerator, a next result tensor segment by executing the tensor operation on a next activation tensor segment included in the activation tensor. Various other methods, systems, and computer-readable media are also disclosed.
Owner:ADVANCED MICRO DEVICES INC

Fully homomorphic encryption (FHE) accelerator permutation application integrated circuit

PendingUS20260189362A1JazelleData stream
A method and system of the device may include a plurality of permute units, where each of the plurality of permute units is configured to perform local permutations on polynomial coefficients of an FHE program executed by the FHE accelerator; where the plurality of permute units are controlled by a set of permute operations, and where the permute operations are derived from instructions complied to minimize dataflows of coefficients in the FHE accelerator while utilizing variable resources of the permutation module and an internal fabric for different permutations.
Owner:CHAIN REACTION LTD

Compiler for machine learning programs

In variants, a program optimization method can include: receiving a program; determining a set of proxy inputs for the program; generating a set of intermediate traces for the program; and generating a set of execution traces for the program, wherein the set of execution traces can include executor fusions associated with device kernels for hardware accelerators. During runtime, program results can be computed by executing an execution trace instead of executing the program.
Owner:GRID AI INC

Scheduling tile program operations

PendingKR1020260140162AJazelleParallel computing
Processors, systems, and technologies for generating schedules of operations of tile-based programs according to constants for the use of function units of an accelerator. In at least one embodiment, a compiler generates a dependency graph representing the operations of a tile program and determines the function units of an accelerator to be used to perform the operations in a first order represented by the dependency graph. The compiler further determines a second order for performing the operations of the tile program based, at least partially, on constraints for the use of the determined function units, and generates code to perform the tile program based on this second order, at least partially.
Owner:NVIDIA CORP

FPGA-ASIC (Field Programmable Gate Array-Application Specific Integrated Circuit) hybrid heterogeneous architecture energy consumption optimization method and device

The invention relates to the technical field of computers, and provides an FPGA-ASIC hybrid heterogeneous architecture energy consumption optimization method and device, and the method comprises the steps: constructing a plurality of initial virtual models of an FPGA-ASIC hybrid heterogeneous architecture; wherein the FPGA-ASIC hybrid heterogeneous architecture comprises a plurality of hardware accelerators, and the distribution states of all the hardware accelerators in all the initial virtual models are different; each initial virtual model is controlled to execute a target task, and execution efficiency information of each initial virtual model is recorded; determining a target virtual model in the initial virtual models based on the execution efficiency information of the initial virtual models; and adjusting the distribution state of a virtual hardware accelerator in the target virtual model so as to optimize the energy consumption of the FPGA-ASIC hybrid heterogeneous architecture. According to the method, the energy consumption of the FPGA-ASIC hybrid heterogeneous architecture is optimized.
Owner:SHENZHEN CITY MAIDIJIE ELECTRONICS TECH