Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

78 results about "Overhead (computing)" patented technology

In computer science, overhead is any combination of excess or indirect computation time, memory, bandwidth, or other resources that are required to perform a specific task. It is a special case of engineering overhead. Overhead can be a deciding factor in software design, with regard to structure, error correction, and feature inclusion. Examples of computing overhead may be found in functional programming, data transfer, and data structures.

Hybrid expert model reasoning method based on cooperation of CPU and GPU

The invention discloses a hybrid expert model reasoning method based on cooperation of a CPU (Central Processing Unit) and a GPU (Graphics Processing Unit), and belongs to the field of deep learning. According to the method, a CPU-GPU computing framework of a hybrid expert model is constructed, heterogeneous computing resource loads are effectively balanced, and the hardware utilization rate is remarkably increased; an intelligent cache management mechanism based on dynamic priority scores is provided, high-demand experts are reserved preferentially, and the transmission overhead caused by cache missing is reduced; through pipeline parallel design for separating calculation and transmission tasks, CPU calculation and PCIe transmission are overlapped in the GPU execution period, and delay is effectively hidden. In addition, in combination with a multi-layer expert activation prediction prospective prefetching mechanism, the expert cache hit rate is improved. The method is compatible with hybrid expert models of different scales and structures, and stable and efficient reasoning acceleration is realized on a resource-limited heterogeneous platform.
Owner:PEKING UNIV

Neural network large model efficient reasoning method based on multiple GPGPUs

The invention belongs to the technical field of artificial intelligence and high-performance computing, and particularly relates to a neural network large model efficient reasoning method based on multiple GPGPUs. The method aims to solve the problems of high communication overhead, non-uniform load, low resource utilization rate, high data transmission delay and the like among multiple processors. Dividing a calculation task into a plurality of sub-graphs through static analysis and mixed granularity partitioning of a model calculation graph; distributing the sub-graphs to the optimal GPGPU based on a weighted cost function in combination with heterogeneous resource perception and a dynamic mapping strategy; a global pipeline scheduling plan is constructed by using communication topology perception, and calculation and communication overlap are maximized; data are loaded in advance through a host side hierarchical caching and asynchronous prefetching mechanism, and transmission delay is hidden; multi-stream concurrent execution and event-based lightweight synchronization are adopted on each GPGPU, so that waiting overhead is reduced. According to the method, the reasoning delay can be remarkably reduced, the throughput and the hardware utilization rate are improved, and the method has good adaptivity and expandability.
Owner:BEIJING TOPMOO TECH

Reverse confusion resisting method and system for deep learning model of end-side equipment

The invention relates to an anti-reverse confusion method and system for an end-side equipment deep learning model, and the core process comprises the steps: firstly inputting an original model obtained through the training of a model training frame into a deep learning compiling frame, and extracting three types of key information, namely, a model operator, a topological structure, parameters and dimensions, through the characteristics of a compiler; then constructing a feature analysis module to evaluate model features, dynamically matching a confusion scheme from a strategy library, and balancing safety and performance; the confusion module is embedded into a plurality of different levels such as a computational graph level, an operator template level and tensor intermediate expression through a hierarchical compiling mechanism, and a complex scheme can be jointly implemented across multiple levels; and finally, the compiler synchronously completes confusion reinforcement when generating the target code. According to the method, hardware adaptation is not needed, low-overhead confusion is achieved through a native pass mechanism of a compiler, fine-grained customized protection is supported, a model structure, parameters and computational logic can be effectively hidden, reverse engineering attacks can be resisted, and the method is particularly suitable for end-side equipment scenes with limited computing power.
Owner:WUHAN UNIV

Approximate floating point multiplier, chip and computing device

PendingCN120104094ADigital data processing detailsBinary multiplierAnd logic unit
The invention discloses an approximate floating point multiplier, a chip and computing equipment, an approximate mantissa multiplier of the approximate floating point multiplier comprises an AND logic unit and a compressor unit, and the AND logic unit is used for performing AND operation on two input operands bit by bit to generate a partial product array with the size of 11 rows and 21 columns; the compressor unit is used for compressing the 11th column to the 21st column by column to obtain a final approximate mantissa, the compressor unit comprises two novel approximate 4-2 compressors ignoring carry design, and the error rate of the approximate 4-2 compressors is within an acceptable range by utilizing mutual compensation inside the compressors. The invention aims to excavate and use the characteristics of floating point multiplication to further improve the energy efficiency of floating point multiplication, realize the optimization of the approximate floating point multiplier in the overhead aspects of precision, power consumption, area and the like, and solve the problems of relatively complex circuit and low compression efficiency of the traditional approximate 4-2 compressor.
Owner:NAT UNIV OF DEFENSE TECH

Power grid field-oriented lightweight large language model fine tuning method

The invention relates to the field of large model fine tuning, in particular to a power grid field-oriented lightweight large language model fine tuning method. The method specifically comprises the following steps: acquiring local power grid data at an edge node, and reasoning through a lightweight large language model to obtain operation state data; and uploading the operation state data to a central node, selecting a matched fine tuning model version from a lightweight model version library, generating a fine tuning strategy, and ensuring efficient screening of training samples and reasonable allocation of computing power resources. Through the incremental training and LoRA technology, the fine tuning process is optimized, the training efficiency is greatly improved, and the calculation overhead is reduced. According to the method, the real-time performance of a power grid task can be guaranteed, meanwhile, limited computing resources of edge nodes are fully utilized, efficient fine adjustment of a lightweight large language model in the power grid field is achieved, and the performance of the model in tasks such as monitoring, fault diagnosis and operation optimization is improved.
Owner:JIANGSU ELECTRIC POWER INFORMATION TECH

Fusion operator execution method, electronic device, storage medium and program product

The invention relates to the technical field of artificial intelligence, and provides a fusion operator execution method, electronic equipment, a storage medium and a program product, and the method comprises the steps: determining a plurality of to-be-fused target operators; according to the tensors of the multiple target operators, a pseudo address mapping table is generated, and the pseudo address mapping table is used for recording the mapping relation between the actual memory addresses of the tensors and pseudo addresses; operator fusion is carried out on the multiple target operators, a fusion operator is obtained, and the memory address of the tensor of the fusion operator is a continuous pseudo address; and executing the fusion operator, converting the pseudo address into an actual memory address according to the pseudo address mapping table in the execution process, and accessing data of the tensor according to the actual memory address. According to the method, a large amount of extra memory copy overhead caused by physical movement and rearrangement of the original tensor in a traditional operator fusion scheme is avoided, and the computing resource utilization rate and the execution efficiency are remarkably improved.
Owner:SHANGHAI BIREN TECH CO LTD

Task scheduling method and system based on distributed scheduling framework

The invention particularly relates to a task scheduling method and system based on a distributed scheduling framework, and relates to the technical field of distributed computing and task scheduling. The resource coordinator is connected with a conflict detection module; a conflict resolution arbitration module; and a final consistency state synchronization module. According to the invention, the lock-free design of the optimistic scheduling decision module and the local cache mechanism break through the dependence of the traditional scheduling on the global real-time consistent state, and realize the top-speed decision in the high-concurrency scene; the local cache with timeliness deviation adopts a hierarchical index structure and a dual-mode updating mechanism, so that a scheduler is supported to complete calculation only by depending on local data while the cache freshness and the updating overhead are balanced, cross-node real-time communication is not needed, and the decision delay is lower.
Owner:成都菁蓉联创科技有限公司

Communication link establishment method and device, equipment, medium and program product

The embodiment of the invention relates to the technical field of artificial intelligence chips, provides a communication link establishment method and device, equipment, a medium and a program product, and is used for reducing cluster establishment cost and avoiding delay overhead caused by inter-node communication while expanding the single-node set communication scale. The method comprises the steps that attribute information of all computing devices deployed in a node is acquired, a plurality of computing devices are deployed in the node, and each computing device comprises a plurality of bare chips; for each computing device, respectively executing the following operations: based on the attribute information of the computing device, configuring the attribute information of each bare chip in the computing device, and obtaining the attribute information of a plurality of bare chips; and constructing a point-to-point communication link among the plurality of bare chips according to the attribute information of the plurality of bare chips.
Owner:SHANGHAI BIREN TECH CO LTD

Lightweight large language model fine-tuning method for power grid field

The present application relates to the field of large model fine-tuning, and specifically relates to a lightweight large language model fine-tuning method for the power grid field. The specific method of the present application comprises: obtaining local power grid data at the edge node and performing inference through a lightweight large language model to obtain operating state data; uploading the operating state data to the center node and selecting a matched fine-tuning model version from the lightweight model version library to generate a fine-tuning strategy, ensuring efficient screening of training samples and reasonable allocation of computing resources. Through incremental training and LoRA technology, the fine-tuning process is optimized, significantly improving training efficiency and reducing computational overhead. This method can ensure the real-time performance of the power grid task while fully utilizing the limited computing resources of the edge node, achieving efficient fine-tuning of lightweight large language models in the power grid field and improving the performance of the model in tasks such as monitoring, fault diagnosis and operation optimization.
Owner:JIANGSU ELECTRIC POWER INFORMATION TECH

Prediction model dynamic scheduling system and method based on heterogeneous computing platform

The invention relates to the technical field of heterogeneous computing, in particular to a prediction model dynamic scheduling system and method based on a heterogeneous computing platform. Running the prediction model in the edge host, and abstracting a calculation-intensive operator in the model into a candidate unloading task; and constructing a time cost model containing host execution time, FPGA execution time and transmission overhead, and dynamically updating in combination with execution statistics. And the scheduling module compares the predicted completion time of the local and unloading paths according to a time cost model, and selects to execute a prediction task on the host or the FPGA. A page locking DMA buffer area is arranged in a system memory, and the FPGA directly reads and writes historical data; and the FPGA side schedules tasks according to the priority and the deadline, and returns to a traditional closed-loop correction time cost model. Compared with the prior art, the method can dynamically adapt to illumination conditions and system load changes, reduces the average reasoning time delay and high quantile tail time delay of tasks, and improves the performance and stability of edge side photovoltaic power prediction.
Owner:GANSU ELECTRIC POWER TIANSHUI POWER SUPPLY

Reconstruction method and corresponding reconstruction system, storage medium and program product

The invention provides a computer-implemented reconstruction method and a corresponding system, a medium and a product. The method is used for reconstructing an original program based on an original variable to obtain a reconstructed program which retains functions of the original program and is based on a target variable different from the original variable. The method includes: receiving an original program, a first mapping specification for converting an original variable or a representation based on the original variable into a corresponding representation using a target variable, and a second mapping specification for converting a target variable or a representation based on the target variable into a corresponding representation using the original variable; the original program is traversed to identify statements in the original program involving the original variable, and a reconstruction operation is performed on the identified statements using the first and second mapping specifications to reconstruct them into corresponding reconstructed statements involving only the target variable. The scheme provided by the invention has the beneficial effects of reducing the processing and computing overhead required by reconstruction, improving the reconstruction efficiency, reducing the requirements on computer hardware and the like.
Owner:THE HONG KONG UNIV OF SCI & TECH

Compression method of sparse data

The invention discloses a sparse data compression method, and belongs to the field of big data and real-time calculation. The method comprises the following steps of: 1, optimizing the layout of an UnsafeRow data memory; 2, encoding the UnsafeRow data; step 3, carrying out UnsafeRow data decoding, and carrying out UnsafeRow data decoding; step 4, carrying out serialization and deserialization on the compressed UnsafeRow; step 5, ORC write-in optimization is carried out; wherein the ORC is an efficient and high-performance column type storage format; and step 6, ORC reading optimization is carried out. By optimizing the memory layout of UnsafeRow and compressing the memory use of sparse data, the memory use efficiency of the data is improved, the overhead of a network and a disk is reduced, and the data processing efficiency of Spark SQL (Structured Query Language) is improved; meanwhile, the storage space of data storage is reduced through encoding optimization of ORC.
Owner:XIAN FIBERHOME SOFTWARE TECH CO LTD

Task scheduling method and system based on heterogeneous pulse neural network on GPU (Graphics Processing Unit)

The invention discloses a task scheduling method and system based on a heterogeneous pulse neural network on a GPU (Graphics Processing Unit), which are used for modeling a scheduling problem into a reward maximization task with constraints so as to optimize resource allocation and parameter adjustment. According to the priority-based multi-preemptive scheduling framework provided by the invention, the heterogeneous SNN tasks on the GPU are dynamically managed by adopting dynamic priority calculation, a self-adaptive time segmentation mechanism and an overhead perception optimization strategy. The effectiveness of the method is verified through a large number of simulation experiments, and the performance under different complexity proportions, arrival modes, load levels and time slice configurations is tested. Results show that the priority-based multi-preemptive scheduling framework provided by the invention is superior to the existing scheduling algorithm in the aspects of energy efficiency, throughput and resource utilization rate. The work promotes the practical application of the SNN workload in a service computing environment, and promotes the development of a more efficient and extensible neuromorphic computing solution.
Owner:ZHEJIANG UNIV

A method and system for distributed parallel processing of mass network data

The application provides a kind of mass network data distributed parallel processing method and system, it is related to data processing field, solve the problem of low efficiency of distributed parallel processing caused by the mismatch of data characteristics and computing power characteristics and the large overhead of global rearrangement in prior art.The method comprises: obtaining network data to be processed;Based on the first feature information of network data, the time consumption of prediction calculation is determined based on the time consumption of calculation, and a plurality of data fragments are obtained based on the slice boundary;Based on the mapping relationship between the first feature information of data fragment and the second feature information of computing node, the data fragment is distributed to the corresponding computing node;In the parallel processing process of data fragment, in response to the deviation between the actual progress and the estimated progress of computing node exceeds the preset threshold, the task of computing node that has not been completed is divided into subtask slices, and is migrated to other computing nodes for local redistribution.The application is used for mass network data processing.
Owner:ANHUI TELECOMM ENG

Intelligent task configuration method based on big data SQL code computing power analysis

An intelligent task configuration method based on big data SQL code computing power analysis performs computing power evaluation on sql codes to be operated, and the core lies in that a query processing process is decomposed into three key dimensions: an operation type, a data scale and processing complexity. Each dimension is quantized through well-designed parameters, and finally a unified computing power demand calculation formula is integrated. The formula not only considers a single-thread execution scene, but also particularly adds extra overhead brought by parallel processing, so that an evaluation result is closer to an actual operation environment. The SQL code is further stored in a notebook and kept as a file, the hash value of the file is calculated to serve as an identification code, a plan of an sql code running task is configured according to the authority level and the real-time computing power condition of a user, and computing power resource allocation is intelligently optimized.
Owner:NANTONG JIUWEI SOFTWARE TECH CO LTD

Method and apparatus for heterogeneous parallel computing

The present invention provides a heterogeneous parallel computing method and apparatus. Among them, the method includes distributing each operator to a corresponding heterogeneous engine group according to different computing power requirements of the operators; numbering the operators distributed to each heterogeneous engine group; parsing the data dependency relationships between the operators in each heterogeneous engine group, and inserting a counter comparison command word into the operator queue, wherein the data dependency relationships are recorded by the numbers of the operators in each heterogeneous engine; performing in-group synchronization on each sub-engine in the same heterogeneous engine group; and performing inter-group synchronization on each heterogeneous engine group. The technical solution provided by the present invention can reduce the overhead of synchronization events while making full use of the computing power of GPUs and AI accelerators, thereby significantly improving the utilization rate of hardware resources and enhancing the overall performance of the system.
Owner:VASTAI TECH (SHANGHAI) INC

Hybrid hardware-software consistency framework

The accelerator device (140) shares the same consistency domain as the hardware elements in the host computing device (105). The mix of hardware and software consistency reduces the overhead of managing data when large blocks of data are moved from the host to the accelerator device. An accelerator application (125) executing on the host identifies a data set (130) that it wishes to transfer to the accelerator device for processing. The accelerator application transfers ownership from the home agent (135) in the host to the accelerator device. The slave agent (155) can then take ownership of the data. As a result, any memory operation request received from the request agent (145) in the accelerator device can obtain access to the data set in local memory (160) via the slave agent without the slave agent having to obtain permission from the home agent in the host.
Owner:XILINX INC

Computing power leasing scheme generation method and device and storage medium

The invention discloses a computing power leasing scheme generation method and device and a storage medium, and relates to the technical field of data processing, and the disclosed computing power leasing scheme generation method comprises the steps: determining computing power levels corresponding to all tasks of a target business scene, and determining cluster types and cluster configuration information corresponding to all the computing power levels; constructing a computing power cluster corresponding to each computing power level according to the cluster type and the cluster configuration information corresponding to each computing power level; scheduling each task to the corresponding computing power cluster for execution to obtain total execution overhead; and if the total execution overhead meets the expected execution overhead, according to the cluster type and the cluster configuration information corresponding to each computing power level, generating a computing power leasing scheme adaptive to the target business scene, thereby solving the problem of high computing power leasing cost of a current computing power leasing scheme generated based on a specific business scene, and reducing the computing power leasing cost.
Owner:SHENZHEN JIEYI TECH CO LTD

Computing Offloading and Resource Allocation Method Applicable to CPU-GPU Heterogeneous Clusters

The present invention relates to a computing offloading and resource allocation method applicable to a CPU-GPU heterogeneous cluster, which considers joint computing offloading and resource allocation in a CPU-GPU heterogeneous network to achieve lower system overhead and higher GPU utilization. Each task is decomposed into a serial segment and a parallel segment, which can be offloaded to the CPU and GPU respectively. Based on the resource sharing technology of the GPU, the computing power of the GPU is discretized, and the computing resource allocation is formulated as an integer programming. Then the task scheduling is modeled as a mixed integer non-linear programming problem to minimize the total overhead composed of latency and energy consumption. We decompose the mixed integer non-linear programming problem so that the computing offloading and resource allocation can be alternately optimized, which leads to an algorithm combining simulated annealing and convex optimization. Numerical simulations are carried out to evaluate the performance of the proposed solution, which is optimal in terms of system overhead, the number of benefited UEs and speedup ratio compared with traditional methods.
Owner:SHANGHAI TECH UNIV

Federal learning acceleration method based on parallel sampling and training of in-batch real-time data of assembly line

PendingCN121684100AMachine learningKnowledge based modelsEvent synchronizationAlgorithm
The invention discloses a federated learning acceleration method and system based on pipeline in-batch real-time data sampling and training. According to the method, in-batch data sampling is provided on the algorithm level aiming at the problems that computing resources of edge equipment are limited and importance is outdated and gradient deviation is caused by existing static sampling: the importance of samples in a fixed mini-batch is evaluated in real time by utilizing a latest model in each round of iteration, and a dynamic micro-batch is constructed; and through a gradient correction coefficient based on a sampling probability reciprocal, distribution deviation is eliminated, and unbiased training is realized. On the system level, an assembly line parallel mechanism based on a CPU-GPU heterogeneous architecture is designed, overlapping execution of sampling and training is achieved through a double-thread-double-flow concurrent model, data competition is solved through an annular buffer area and an event synchronization mechanism, and overhead is reduced in cooperation with mixing precision reasoning. According to the method, the hardware utilization rate can be remarkably improved, the model convergence precision is improved while the training time is greatly shortened, and the method is suitable for various edge computing scenes.
Owner:NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Fair resource management method and system for multi-user shared large language model reasoning

The invention discloses a fair resource management method and system for multi-user shared large language model reasoning, and the method is characterized in that the method employs a cold data recognition elimination mechanism and a fair cache distribution mechanism, carries out the unified quantification of the resource consumption of a request in GPU calculation, decoding and cache transmission, and carries out the fair scheduling based on the accumulated CPI value of each user; the system comprises a user request management module, a CPI fair scheduling module, a model reasoning execution module, a resource monitoring module and a cache management module. Compared with the prior art, the method has the advantages that users can share computing and caching resources fairly, efficient operation performance of the system is guaranteed, users with high cache hit rate are prevented from excessively occupying GPU execution opportunities due to low apparent computing overhead, the overall reasoning efficiency and the cache hit rate of the system are remarkably improved, and the user experience is improved. Unified and fair distribution of computing resources and cache resources is achieved, the effect is prominent especially in a multi-tenant high-load scene, and the method has good application prospects and commercial development value.
Owner:EAST CHINA NORMAL UNIV

System and method for dynamic redundancy-aware blockchain-based partial computation offloading for metaverse within computing environment in network

The present disclosure relates to a system and method for dynamic redundancy-aware blockchain-based partial computation offloading for a metaverse within a computing environment in a network, and according to the present disclosure, in order to improve the QoS of a metaverse service, it is possible to provide an environment that may perform existing traditional task offloading, perform optimal offloading through an in-network computing agent, and provide an expandable network and an ultra-low latency service. In addition, it is possible to maximize incentives while minimizing the overhead of computation execution costs incurred when a user performs a task, and satisfy the constraints on latency and blockchain offloading costs.
Owner:IND FOUND OF CHONNAM NAT UNIV +1

AI data handling optimization method based on multi-core system

The invention discloses an AI data handling optimization method based on a multi-core system. The data handling optimization method comprises the steps that S1, initialization is carried out; s2, asynchronous request; s3, asynchronous processing; and S4, parallel execution. The processor A for executing the AI operation is always in the user state and does not need to fall into the kernel state, so that the system calling overhead is avoided; in the process that the processor A executes the AI operation, the processor B schedules DMA hardware to complete transmission processing of next data required by the AI reasoning operation, the DMA hardware and the AI reasoning operation are completely parallel, and when DMA transmission time is shorter than AI reasoning operation time, time overhead of DMA data transmission is completely masked; based on the software and hardware collaboration framework, the utilization rate of the computing part of the processor A can be close to 100%.
Owner:JINDIE SPACETIME (BEIJING) TECHNOLOGY CO LTD

Polynomial chain verification method and device for outsourcing neural network inference

The present application relates to the field of specific computing model and information security technology, aiming at the problems of high computing cost and high cost in existing outsourcing reasoning technology, a polynomial chain verification method and device for outsourcing neural network reasoning are proposed.The method comprises: data input;in the preprocessing stage, the client generates the mask matrix of the input data, the model end generates the mask and the verification auxiliary matrix of the model parameter, the input data, the mask matrix and the Beaver triple are divided into multiple shares and distributed to different computing servers;by constructing a verifiable polynomial chain, the verification amount and the mask are generated under additive secret sharing and returned to the client, and by designing a polynomial form of verification equation, the verification of the reasoning process is completed.The present application completes the reasoning verification by designing a polynomial equation, ensures the integrity of the whole process and the privacy, and realizes the low-overhead, low-cost step-level verification with constant-level online verification calculation time of parallel computing.
Owner:NAT UNIV OF DEFENSE TECH

Parallel computing communication method, device and equipment based on distributed many-core processor

The invention relates to the technical field of communication, and discloses a parallel computing communication method, device and equipment based on a distributed many-core processor, high-precision time sequence synchronization can be realized through a central polling controller, a partition shared memory, dual-port extension and combination of a standardized process, the reliability of a computing result is improved, the software complexity is greatly reduced, and the communication efficiency is improved. The communication flexibility is ensured, the communication energy efficiency ratio is remarkably improved, data conflicts are also avoided, and power consumption and time sequence overhead caused by bus state switching are reduced. Data processing is continuous due to assembly line work, the overall throughput rate of the system is increased, the deployment process is greatly simplified, and the applicability and usability of the platform are improved.
Owner:SHANDONG INSPUR SCI RES INST CO LTD

A distributed computing method and system based on MDS coding and flexible scheduling strategy

The present invention discloses a distributed computing method and system based on MDS encoding and a flexible scheduling strategy. This method is oriented towards a master-slave distributed computing framework. Based on different task arrival rates, and considering the non-negligible overhead of revoking redundant tasks, the method, by designing an appropriate model encoding scheme and task scheduling strategy, alleviates the stragglers problem in the distributed system while balancing the number of redundant tasks with the system load, thereby reducing the average execution time of the overall task. Furthermore, the method considers the situation in which new and old tasks exist after adjusting the model encoding scheme and task scheduling strategy. By designing a compatible solution that distinguishes task types, it is possible to avoid invalidating the calculation results of tasks before the adjustment.
Owner:NANJING UNIV

An Execution Management Method and Device for Spark SQL Query Plan Trees Based on DPU

The present invention provides an execution management method and device for a Spark SQL query plan tree based on DPU, which allows hybrid computing of DPU and CPU for query tasks, improves computing performance and resource utilization rate. When the adaptive query execution mode is not running, the query plan tree is deployed as a whole. When all operators conform to the types supported by DPU, the operators are preferentially offloaded to DPU for operation, otherwise the whole is handed over to CPU for processing. When the adaptive query execution mode is running, first judge whether the currently intercepted sub-plan tree has an adaptive query execution operator. If so, first judge whether the entire query plan tree can be offloaded and marked in the configuration sheet. If not, query the mark in the configuration sheet and hand over the current sub-plan tree to DPU or CPU for processing according to the mark. This execution management method deploys the whole for a single query task, avoiding the introduction of additional row-column conversion operators and the data replication overhead brought by the row-column conversion operators, saving resources and improving computing performance.
Owner:YUSUR TECH CO LTD

Heterogeneous computing scheduling method, system, device and storage medium

The disclosure provides a heterogeneous computing scheduling method, system, device and storage medium, and relates to the technical field of artificial intelligence. In some embodiments of the disclosure, according to the task descriptor, the page table mapping relationship of the input-output memory management unit is configured without the participation of the system core, and address translation preprocessing is completed; the virtual address interval accessed subsequently by the neural network processing unit is predicted; before the neural network processing unit initiates a memory access request, according to the virtual address interval, the page table hardware traversal of the input-output memory management unit is triggered, the page table entry is obtained, and the page table entry is loaded into the input-output conversion buffer; a start signal is sent to the neural network processing unit, so that the neural network processing unit accesses the shared main memory through the input-output memory management unit based on the page table mapping relationship. Through the dual-core division of the system core and the control core, the system overhead caused by the participation of the operating system in the underlying scheduling is reduced, and the additional power consumption overhead is reduced.
Owner:BEIJING VCORE TECH CO LTD

Tile block instruction set architecture and processing method

The invention relates to the technical field of processors, in particular to a tile block instruction set architecture and a processing method, and the tile block instruction set architecture comprises a tile block parallel processing system control instruction, a tile block descriptor register configuration instruction, a tile block data handling instruction, a tile block data copying instruction and a tile block arithmetic instruction. According to the tile block instruction set architecture processing method, one or more tile block operation units, one or more tile block storage management units and one or more tile block calculation task scheduling and synchronization units are combined into a tile block parallel processing system, so that tile block instructions are allowed to be processed in parallel, and tile block-based data operation is completed. The architecture aims at solving the problems that when an existing processing system processes two-dimensional data blocks, instruction overhead is too large, data abstraction is insufficient, and the coupling degree of software and hardware is too high.
Owner:NANJING UNIV

Optimization processing method, system, device and medium of ai accelerator

The application provides an AI accelerator processing optimization method, system, device and medium, and the method is applied to an AI accelerator in communication connection with a main memory. Through the collaborative architecture of the AI accelerator and the 3D DRAM, the application realizes reduction of transmission overhead and power consumption in the data carrying process, improvement of the computing throughput and the energy efficiency ratio, and simultaneously supports large-scale parallel computing tasks relying on the high bandwidth characteristics of the 3D DRAM.
Owner:SHANGHAI GUANGYU XINCHEN TECHNOLOGY CO LTD