Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

103 results about "Computational RAM" patented technology

Computational RAM or C-RAM is random-access memory with processing elements integrated on the same chip. This enables C-RAM to be used as a SIMD computer. It also can be used to more efficiently use memory bandwidth within a memory chip.

Key-value cache management system for inference processes

A key-value cache management system processes inference requests of a generative model. An inference request requests a starting virtual memory address and a number of layers in the generative model. In response to receiving an inference request, contiguous virtual memory space is reserved for the key-value cache in accordance with the inference request by assigning the starting virtual memory address in the key-value cache and calculating memory pointers for each layer in the model and each block so that the one or more blocks are written sequentially. The generated memory pointers are outputted to the generative model so that each self-attention layer writes computed key-value pairs at physical addresses in the key-value cache specified by the memory pointers.
Owner:BYTEDANCE TECHNOLOGY LTD +1

Compute in memory three-dimensional non-volatile NAND memory for neural networks with weight and input level expansions

A non-volatile memory device for performing compute in memory operations for a neural network uses a three dimensional NAND architecture. Multi-bit weight values are stored encoded as sets of threshold voltages for sets of memory cells. A weight value is stored in multiple memory cells on the same word line and connected between a bit line and a source line, each of the memory cells programmed to one of multiple threshold voltages. When multiplying an input value with the weight value, the word line is biased so that, for at least one of the threshold voltages, the memory cell will be in the linear operation region. Input values are encoded as a set of one or more voltage levels applied to a corresponding set of bit lines, each bit line connected memory cells also storing the weight value, connected to the word line, and connected to the source line.
Owner:SANDISK TECHNOLOGIES LLC

Protocol-Aware Provisioning of Resource over CXL Fabrics

Dynamic provisioning of resources in a datacenter improves the utilization efficiency of compute, memory, storage, and network resources, while maintaining flexibility to meet changing demands. Embodiments herein disclose protocol-aware provisioning of resources over Compute Express Link (CXL) fabrics, enabling multi-protocol pooling and disaggregation of resources, including memory and workload-specific accelerators such as GPUs and DSAs. In some embodiments, a Resource Provisioning Unit (RPU) facilitates intent-based protocol translations and mappings between address spaces, potentially enabling the creation of large-scale compute-memory fabrics that may utilize both coherent and non-coherent Non-Transparent Bridging (NTB) between multiple protocols in a single system, serving as the underlying infrastructure for executing workloads such as Large-Language Models (LLMs) and Deep Learning Recommendation Models (DLRMs). Some embodiments also optimize low-latency communication between processes running on different nodes by enabling host-to-host memory provisioning for libraries such as OpenMP, Pthreads, or CUDA, which can utilize shared memory.
Owner:HYATT GAYA OPAL MS +1

Data processing method and system of low-energy-consumption large language model based on momentum mechanism and multiple types of experts

The invention discloses a data processing method and system for a low-energy-consumption large language model based on a momentum mechanism and multiple types of experts, and belongs to the technical field of large language models, and the method comprises the steps: obtaining a hybrid expert model which comprises a gating network, a combination network and a plurality of expert networks; obtaining target data based on a lazy loading mechanism; obtaining fitness scores of the target data corresponding to the expert networks through the gating network, and obtaining the expert network meeting the sorting requirement as a target network; and obtaining output data of the target data in the target network, and carrying out weighted summation on the multiple output data through the combined network. According to the method, the multiple expert networks are arranged in the hybrid expert model, the fitness ranking is performed on the fitness scores of the target data, and the output of the expert network meeting the ranking requirement is selected to perform weighted summation so as to obtain the final output, so that the floating-point operation and calculation memory overhead are reduced, the calculation efficiency is improved, and the memory resource waste is reduced.
Owner:NANJING UNIV OF INFORMATION SCI & TECH

Heterogeneous resource dynamic sensing and self-adaptive scheduling method

The invention provides a heterogeneous resource dynamic sensing and self-adaptive scheduling method, which comprises the following steps of: dynamically acquiring basic attributes such as resource types, core configuration, memory capacity and bandwidth and real-time information such as load state and temperature through a heterogeneous resource feature extraction module by utilizing a hardware interface and software probe technology, and storing data in a feature table after processing, accurate scheduling decision basis is ensured; a self-adaptive scheduling algorithm is designed, matching with resource characteristics is carried out based on task calculation, memory and parallelism requirements, proper resources are preferentially allocated, competition is avoided by checking a load state, and the algorithm supports dynamic adjustment so as to optimize allocation efficiency; and finally, through a real-time monitoring and dynamic adjustment mechanism, regularly collecting resource and task state information, detecting a competition or low-efficiency condition, executing operations such as task migration, priority adjustment or resource reservation and the like, and ensuring efficient operation of the system. The heterogeneous resource scheduling efficiency and stability are improved, and the method is suitable for a modern computing environment.
Owner:JIANGSU WANWEI AISI NETWORK INTELLIGENT IND INNOVATION CENT

Static random-access memory (SRAM) device and related SRAM-based compute-in-memory devices

An SRAM cell includes a first inverter cross-coupled to a second inverter. The first inverter includes a first pull-up transistor and a first pull-down transistor, having coupled drains that define a first storage node. The SRAM cell further includes a first N-type pass-gate transistor having a first drain coupled to a write bit line, a first source coupled to the first storage node, and a first gate coupled to a first write word line. The SRAM cell further includes a first P-type pass-gate transistor having a second drain coupled to the write bit line and a second source coupled to the first storage node. The SRAM cell further includes a P-type transistor having a third drain, coupled to a second gate of the first P-type pass-gate transistor, a third source coupled to a second write word line, and a third gate coupled to an enable signal.
Owner:TAIWAN SEMICONDUCTOR MANUFACTURING CO LTD

Multi-channel data acquisition system and method based on FPGA (Field Programmable Gate Array)

The invention relates to the technical field of signal processing, and discloses a multichannel data acquisition system and method based on an FPGA (Field Programmable Gate Array), an acquisition module drives a global counter by using a global synchronous clock, writes acquired data into an annular buffer area with a physical address and a counter value in a linear mapping relationship, and establishes implicit time index storage. When a trigger event occurs, the system broadcasts the locked global trigger timestamp, and the acquisition module backtracks and reads historical data according to the global trigger timestamp, and packages the historical data into a sparse matrix type data packet in combination with the channel validity mask. And after receiving the data packet, the data processing module directly calculates a memory mapping address by using global time information in the data packet, and writes a data load into a corresponding position of the waveform reconstruction buffer area. According to the method, through strict binding of the physical address and the absolute time, independent time label redundancy is eliminated, high-precision synchronization of distributed multiple channels is ensured, and out-of-order automatic in-situ recombination and efficient waveform reconstruction of data of a receiving end are achieved.
Owner:CHANGCHUN TESTING MASCH RES INST

Adaptive gateway flow scheduling method based on multi-dimensional information

The invention discloses a multi-dimensional information-based adaptive gateway traffic scheduling method, which comprises the following steps of: acquiring server performance indexes, network delay, bandwidth and geographic position data in real time, and establishing a historical database; performance load weights are generated through weighted calculation of CPU, memory and disk indexes, and routing priorities are dynamically adjusted by adopting exponential function mapping; performing normalization processing on network bandwidth, delay and geographic distance, and calculating a service access weight to optimize a transmission path in combination with a step function; analyzing historical traffic data based on an LSTM model to predict a future trend, and generating a traffic prediction weight through a sigmoid function to realize pre-equalization; and combining the three types of weights, weighting to generate dynamic routing configuration, eliminating abnormal nodes in real time in combination with a health check mechanism, triggering high-load early warning and linking a predictive balancing strategy. According to the method, multi-dimensional index collaborative decision-making is realized, routing distribution is dynamically optimized, network delay is effectively reduced, and system throughput and stability in a high-concurrency scene are improved.
Owner:河北省体育局射击射箭运动中心(河北省军事体育运动学校) +1

Network security functions for dynamic construction and programmatic placement

A system and method are provided for placing network functions among respective locations in a network. The locations at which the network functions are placed can be nodes and network devices within the network. These nodes can be selected, e.g., based on which network devices have available capacity and or specialized hardware (e.g., accelerator sin a data processing units (DPUs)) that is optimized for particular network functions. The network functions can include an inline network function that is provisioned directly in a data plane of one of the network devices (e.g., in-lined directly in a hardware offload device without a virtual machine and without a container). The decision of where to place the network functions can be based on a performance metric (e.g., representing available computational / memory resources at the network nodes) and / or a network-function metric (e.g., representing consumed computational / memory resources by the network functions) to improve system performance.
Owner:CISCO TECHNOLOGY INC

Tree-Based Network Architecture for Accelerating Machine Learning Collective Operations

Aspects of the disclosure are directed to a tree-based network architecture for serving and / or training machine learning models. The architecture includes one or more multi-chip packages having a plurality of compute-memory stacks connected via an input / output (I / O) die. The I / O die includes an aggregator to aggregate computations from the compute-memory stacks. The architecture can further include a plurality of the multi-chip packages connected on a server via a server level aggregator and a plurality of the servers connected on a rack via a rack level aggregator for further aggregation of the computations from the compute-memory stacks. The tree-based network architecture allows for fewer hops, resulting in lower latency and savings in bandwidth when serving and / or training machine learning models.
Owner:GOOGLE LLC

Edge server load evaluation method based on analytic hierarchy process and performance loss

The invention relates to the technical field of edge computing resource management, in particular to an edge server load evaluation method based on analytic hierarchy process and performance loss. According to the method, firstly, edge servers are divided into a common type, a computing type, a memory type and an I / O type according to heterogeneous types; aiming at each type of edge server, constructing a judgment matrix by adopting an analytic hierarchy process and solving a feature vector so as to calculate load weight coefficients of CPU (Central Processing Unit), memory and disk I / O (Input / Output) resources; calculating a real-time load based on the weight coefficient, wherein the real-time load is the weighted sum of the resource utilization rates; the performance loss is further calculated according to the nonlinear relation between the real-time load and the total load ratio; and finally, restraining the real-time load through the load threshold value, and evaluating the performance state of the server in combination with the performance loss. According to the method, the accuracy and reliability of load evaluation of the heterogeneous edge server are effectively improved, the resource utilization rate is optimized, overload of the server is avoided, and the system stability is enhanced.
Owner:GUIZHOU INST OF TECH

High-concurrency virtual service transfer scheduling method and system based on cloud edge collaborative architecture

The invention discloses a high-concurrency virtual service transfer scheduling method and system based on a cloud edge collaborative architecture, and relates to the technical field of cloud computing and edge computing collaboration, and the method comprises the steps: constructing a service gene analysis model at an edge node, responding to a virtual service access request, extracting the resource consumption characteristics of a service flow in real time, and calculating a service separation potential energy index. The virtual business is dynamically decoupled into a connection anchoring phase and a computing power free phase based on the index; establishing a load inertia monitoring mechanism, modeling resource consumption of edge nodes into a kinematics system, collecting speed and acceleration of load change in real time, calculating system pressure momentum, and generating a graded unloading instruction when the momentum exceeds a preset critical collapse threshold value; starting a convergent differential injection mechanism, pre-constructing a logic shadow at a cloud end, and calculating a convergence ratio of a memory dirty page generation rate to a network transmission rate in real time; the method has the advantages of being high in practicability, having the pre-judgment capacity and being capable of achieving lossless switching.
Owner:JIANGSU AIYUSEN INFORMATION TECHNOLOGY CO LTD

Dual-mode computing memory controller and memory system for PIM-DRAM

The invention belongs to the technical field of semiconductors, and particularly relates to a dual-mode computing memory controller and a memory system for a PIM-DRAM. The dual-mode computing memory controller used for the PIM-DRAM comprises an address conversion module used for processing address mapping between a host and a PIM core; the command translation module is used for converting the host command into a PIM operation instruction; the command generation module is used for generating an address mapping command; the data translation module is used for managing data streams among the DDR bus, the on-chip bus and the DRAM; the multi-mode state machine is used for dynamically switching to a memory mode or a PIM mode according to different memory commands sent by the host; and the dual-mode computing access module is used for selecting an access space of the DRAM according to the address mapping command or the data route. According to the method, efficient and low-delay PIM operation is achieved by dynamically switching the memory and the calculation mode, meanwhile, the method is compatible with a traditional DDR protocol and can be directly integrated to an existing system, and a host interface does not need to be modified.
Owner:CORE ARK (SHANGHAI) INTEGRATED CIRCUIT CO LTD

Heterogeneous BIM model lightweight server intelligent scheduling method and system

The invention discloses a heterogeneous BIM model lightweight server intelligent scheduling method and system. The method comprises the steps that a data storage structure is constructed, and an engine information table, a server information table and a task information table are designed; according to the engine identification code, obtaining unit task memory consumption and server memory configuration information from the engine information table and the server information table, counting the total number of server tasks and the number of lightweight task success from the task information table, calculating a lightweight success rate, and counting the number of tasks in processing from the cache structure; based on the unit task memory consumption, the server memory configuration, the lightweight success rate and the number of tasks in processing, calculating a memory usage amount, a memory occupancy rate and a residual computing power, and determining a weight factor according to the lightweight success rate to obtain a weighted residual computing power; and realizing server scheduling based on the calculated memory occupancy rate and weighted residual computing power. By adopting the method, model resource consumption and server residual computing power can be dynamically matched, and the average utilization rate of a cluster memory is improved.
Owner:DONGHUI (ZHEJIANG) TECHNOLOGY CO LTD

Protocol-aware provisioning of resource over CXL fabrics

Dynamic provisioning of resources in a datacenter improves the utilization efficiency of compute, memory, storage, and network resources, while maintaining flexibility to meet changing demands. Embodiments herein disclose protocol-aware provisioning of resources over Compute Express Link (CXL) fabrics, enabling multi-protocol pooling and disaggregation of resources, including memory and workload-specific accelerators such as GPUs and DSAs. In some embodiments, a Resource Provisioning Unit (RPU) facilitates intent-based protocol translations and mappings between address spaces, potentially enabling the creation of large-scale compute-memory fabrics that may utilize both coherent and non-coherent Non-Transparent Bridging (NTB) between multiple protocols in a single system, serving as the underlying infrastructure for executing workloads such as Large-Language Models (LLMs) and Deep Learning Recommendation Models (DLRMs). Some embodiments also optimize low-latency communication between processes running on different nodes by enabling host-to-host memory provisioning for libraries such as OpenMP, Pthreads, or CUDA, which can utilize shared memory.
Owner:HYATT GAYA OPAL MS +1

Cache capacity control method, device, equipment and system

The invention discloses a cache capacity control method, device, equipment and system, relates to the technical field of computers, and can improve the utilization rate of memory resources. The method can be applied to a cluster, the cluster comprises a scheduler management node and a plurality of computing nodes, the method is executed by the scheduler management node, and the method comprises the following steps: the scheduler management node receives a cache resource request sent by a first computing node, and allocates a cache resource to the first computing node according to the cache resource request and a resource allocation strategy; determining a target resource quantity allocated to a cache system of the first computing node from a computing memory of the first computing node to obtain a resource allocation result; and sending the resource allocation result to the first computing node. Wherein the resource allocation result is used for indicating that a target resource quantity in idle resources of the computing memory is divided into the cache system, and the resource allocation strategy comprises at least one of the following items: an occupation parameter of the computing memory, an occupation parameter of the cache system and a reserved resource quantity of the computing memory.
Owner:HUAWEI TECH CO LTD

Fully analog compute-in-memory architecture for neural networks

A Fully Analog STate-Space Compute-In-Memory (FAST-CIM) architecture implements cascaded CIM arrays that maintain continuous analog signal flow throughout neural network computations. Sequential data patches are processed through cascaded CIM arrays, where a feedforward CIM array performs initial vector-matrix multiplication using analog computations within memory cells, and the analog output flows directly through a gain circuit to a recurrent CIM array without digital conversion. The gain circuit converts analog current signals to voltage signals, enabling direct analog connection between the cascaded arrays. The recurrent CIM array executes State-Space Model (SSM) computations, processing the cascaded analog signals along with previous state information maintained by a State Write and Propagate (SWAP) circuit that stores states using capacitive elements. The cascaded array configuration eliminates analog-to-digital and digital-to-analog converters between processing stages, achieving reduced power consumption and latency compared to conventional implementations that require digital interfaces between CIM arrays.
Owner:GEORGIA TECH RES CORP

Data loading method, device, computer equipment and storage medium

The present application is applicable to the field of data loading technology and provides a data loading method, apparatus, computer equipment, and storage medium, wherein the method comprises: obtaining a first target storage capacity required for loading target data from a disk in a terminal into a memory; obtaining the remaining computational storage capacity of the memory, and when it is determined that the remaining computational storage capacity is less than the first target storage capacity, calculating the ratio of the sum of the used computational storage capacity of the memory and the first target storage capacity to the actual total available amount of the memory to obtain a calculation result; based on the calculation result, when it is determined that the terminal is in a target operating state, loading the target data into the memory. This solution can reduce the frequency of heating of electronic devices, increase the service life of electronic devices, and enhance the user experience.
Owner:CHINA PING AN PROPERTY INSURANCE CO LTD

Apparatus, engine, system and method for predictive analytics in a manufacturing system

A predictive analytics apparatus, engine, system and method capable of providing real time analytics in a manufacturing system that may include a data input capable of receiving raw data output from at least one machine operable to effect the manufacturing system embodiments, and a processor to execute code from a computing memory. The code may comprise an adaptor to push the received raw data to a database to processed data; an extractor to extract the processed data from the database; predictive analytics to receive the extracted processed data and apply thereto a predictive model comprised of target data for the at least one machine, and to provide feedback to the at least one machine to modify performance of the at least one machine based on the application of the predictive model; and a visualizer capable to provide at least a visualization of the feedback and the performance.
Owner:JABIL INC

Floating-point computing-in-memory device, exponent computing memory module and mantissa computing memory module

A floating-point computing-in-memory device, an exponent computing memory module and a mantissa computing memory module are provided. The floating-point computing-in-memory device includes the exponent computing memory module and the mantissa computing memory module. The exponent computing memory module includes a plurality of weighting exponent memory circuits, a plurality of exponent computing circuits and a comparison circuit. The exponent computing circuits are used to obtain a plurality of exponent products. The mantissa computing memory module includes a bit shifting circuit, a plurality of weighting mantissa memory circuits, a plurality of mantissa computing circuits, a shift-and-addition circuit, a plurality of weighting sign memory circuits, a plurality of sign computing circuits and an addition circuit. The mantissa computing circuits and the shift-and-addition circuit are used to obtain a plurality of mantissa products. The sign computing circuits are used to obtain a plurality of sign products.
Owner:IND TECH RES INST

A large file export optimization method and system based on real-time memory monitoring

The application relates to the technical field of data processing, and discloses a large file export optimization method and system based on real-time memory monitoring. The method comprises the following steps: monitoring a memory state in real time and calculating a memory growth rate; taking the growth rate as an independent risk dimension, and giving a warning when the rate exceeds a threshold value and the memory usage rate does not reach a threshold value; performing adjustment in response to the warning, including dynamically updating the throughput control of the data shard size, and triggering the degradation mode to reduce the export field to the data precision control of the core field; and attaching an instruction page to the export file during degradation. Through trend prediction and dynamic control, the application reduces the risk of memory overflow and guarantees core data export.
Owner:山东齐鲁壹点传媒有限公司 +1

Method for Obtaining Periodic Structure Wideband RCS Based on CBFM and CAT

The present invention proposes a method for obtaining the broadband RCS of a periodic structure based on CBFM and CAT, and the implementation steps are as follows: 1) Divide the periodic structure target; 2) Obtain the incident wave electric field of each sub-region; 3) Calculate the Chebyshev frequency sampling points within the frequency band f; 4) Use the characteristic basis function method CBFM to solve the induced current at each Chebyshev frequency sampling point; 5) Use the Chebyshev approximation method CAT to calculate the surface current at the wave number corresponding to each frequency within the frequency band f; 6) Obtain the radar cross section RCS(ks) of the periodic structure target at the wave number corresponding to each frequency. The present invention divides the periodic structure target into N sub-regions, uses the Chebyshev approximation method to calculate the induced current at each frequency sampling point, represents each sub-region with a characteristic basis function, compresses the impedance matrix into an N-dimensional matrix, solves the problem of large computing memory when calculating the periodic structure, and improves the computing efficiency.
Owner:XIDIAN UNIV

A kinetic solution method and system for multi-scale particle transport simulation

This invention belongs to the technical field of multi-scale particle transport simulation, and discloses a kinetic solution method and system for multi-scale particle transport simulation. The method includes: a particle kinetic model based on discrete velocity form; selecting discrete points in velocity space; calculating the distribution function along the direction of the discrete points in velocity space at the center of the grid according to a compact scheme; accumulating the contribution values ​​of the distribution function to macroscopic physical quantities; scanning the downstream grid sequentially based on the direction of the discrete points in velocity space; changing the discrete points in velocity space and repeating the operation until the contribution values ​​of the grid at all discrete points in velocity space are accumulated, and obtaining the final macroscopic physical quantities; outputting the calculation results of macroscopic physical quantities when the convergence condition is met. This invention discloses a kinetic solution method and system for multi-scale particle transport simulation that balances computational memory, accuracy, and efficiency, and is particularly suitable for multi-scale particle transport simulation with a large number of discrete points in velocity space and strong heterogeneity.
Owner:HUAZHONG UNIV OF SCI & TECH

Data processing method, distributed system, electronic device, and storage medium

Embodiments of the present disclosure provide a data processing method, a distributed system, an electronic device and a storage medium. The method is applied to a distributed system comprising a master node and a plurality of slave nodes, and comprises: the master node determining task information allocated to each slave node according to a hierarchical attribute of a target model and a computing power state of the plurality of slave nodes, wherein the task information is used to specify a model layer processed by the corresponding slave node; each slave node performing weight quantization processing of the corresponding model layer based on the allocated task information to obtain complete quantized weights of the target model. The method realizes full-link collaborative optimization from computing, memory to storage, significantly improves the execution efficiency of weight quantization tasks and the utilization rate of hardware resources, and provides technical support for efficient deployment of large-scale models.
Owner:SHANGHAI BIREN TECH CO LTD

Cross-data center large model training system architecture and resource allocation method and system

PendingCN121957895AImplement collaborative trainingEfficient collaborative utilizationResource allocationBiological modelsWide areaData center
The invention provides a system architecture for cross-data center large model training and a resource allocation method and system, and belongs to the technical field of cross-wide area distributed large model training. According to the method, large-scale model cooperative training across multiple data centers can be realized, and the bottleneck that the computing power of a single data center is limited is broken through. Through unified modeling and scheduling of calculation, memory and network resources, task loads can be intelligently allocated according to hardware performance and network bandwidth of different data centers, and efficient collaborative utilization of computing power resources is realized. The training task of the super-large-scale model can be rapidly completed in the heterogeneous computing power environment, and the training time is remarkably shortened. The provided flexible parallelism degree allocation method can be adaptive to different task and resource conditions, the proportion of data parallelism, model parallelism and pipeline parallelism is automatically adjusted, the parallelism efficiency is improved, and the communication overhead is reduced. A training time estimation function is integrated, the overall time delay and resource requirements can be predicted before task execution, and a basis is provided for scheduling decision making.
Owner:BEIJING JIAOTONG UNIV

A category-based 6g network multi-dimensional resource ai model dynamic deployment optimization method

The application discloses a kind of 6G network multidimensional resource AI model dynamic deployment optimization methods based on category theory, belongs to intelligent collaborative optimization technical field;Method is: the cross-layer consistency dependency of end-to-end AI reasoning service is formalized by functor form;Establish the joint optimization model with long-term average end-to-end delay minimization as target, while being constrained by multidimensional resource and service quality;Convert long-term random optimization problem into time-slot online decision problem;Get AI model dynamic deployment and task scheduling result.The application realizes cross-layer consistency description to task scheduling and model deployment through category theory unified modeling and functor composite mechanism, reduces the inconsistency and redundant constraint caused by hierarchical modeling, improves the structured degree and explainability of joint decision;Under the constraint of multidimensional resources such as calculation, memory, storage and bandwidth, dynamic adaptive optimization is realized, node resource over-limit and load imbalance are effectively avoided, and congestion and queuing delay are reduced.
Owner:NANJING UNIV OF POSTS & TELECOMM

Wireless baseband modem datapath using compute in memory

System and method embodiments are disclosed for merging computing hardware datapath to reduce energy consumption and improve throughput of modem signal processing functions on a Modem. Compute-in-memory (CIM) is used for data access and processing in parallel using processing elements integrated within memory chip. The processing throughput may significantly improve since data is accessed directly from the computational memory storage in parallel for processing and the clock speed for data processing may run at a higher spec than conventional approaches. Energy consumption for Modem processing is also reduced.
Owner:EDGEQ INC

An NTT hardware implementation system based on an FPGA platform

The present application belongs to the technical field of post-quantum cryptography engineering, and specifically relates to an NTT hardware implementation system based on an FPGA platform. The system is used for accelerating the polynomial multiplication operation on a ring with the highest complexity and the longest time consumption in a lattice-based cryptography scheme; the hardware circuit supporting multiple parameters is designed for standard NTT, pruned NTT and hybrid NTT, and specifically includes: a compact butterfly operation unit, a reconfigurable modulo reduction unit and a cross-storage type memory access mode, the polynomial multiplication operation is completed through a multi-parallel acceleration structure and a multiplexed butterfly calculation unit; forward NTT, corresponding point coefficient multiplication and inverse NTT operations can be realized. The butterfly operation unit is internally in a pipeline mode, thereby reducing the time delay of the critical path; the modulo reduction hardware unit uses addition and shift to replace the multiplication operation therein, thereby reducing the consumption of DSP resources; the cross-storage type memory access mode can simultaneously support three NTT calculations, and the memory utilization rate reaches 100%.
Owner:FUDAN UNIVERSITY

Task scheduling method and device for large language model, equipment and storage medium

The invention discloses a task scheduling method and device for a large language model, equipment and a storage medium, and the method comprises the steps: carrying out the tensor segmentation of a to-be-segmented tensor in the large language model, so as to obtain a CC tensor, a CG tensor and a GG tensor; and scheduling hardware resources corresponding to the CC tensor, the CG tensor and the GG tensor to cooperatively work so as to control the large language model to execute tasks. According to the method, the tensor is divided into the CC tensor, the CG tensor and the GG tensor which use different hardware resources, and then all the hardware resources work cooperatively to make full use of available computing, memory and communication resources, so that the reasoning efficiency of a small computing system can be improved under the condition that the reasoning accuracy of a large language model is ensured.
Owner:HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)

Rapid line resistance network analysis method for nonvolatile memory array

The invention belongs to the technical field of line resistance network analysis, and discloses a fast line resistance network analysis method for a nonvolatile memory array, which comprises the following steps of: establishing electrode KCL equations of all rows and columns in a line resistance network of the nonvolatile memory array; a voltage line and a coefficient matrix in the KCL equation are transformed and decomposed, a fixed point iterative algorithm is introduced, and voltage distribution of nodes at the top and the bottom of the array is iteratively calculated. An original complex large-scale linear equation set problem is converted into two small-scale sub-problems. And solving is carried out through mutual iteration of the two sub-problems, so that the calculation complexity and the calculation memory are reduced. According to the method, the influence of IR-drop on neural network calculation acceleration can be predicted and evaluated in the design stage. The method has the advantages that the accuracy is high, and the influence of IR-drop can be simulated with the accuracy exceeding 95%; the calculation complexity is low, and the calculation complexity is reduced from O ((2mn) 2) to O (mn); the memory load is small, and the use of the memory is minimized while high accuracy is kept.
Owner:HUAZHONG UNIV OF SCI & TECH