Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

313 results about "Memory bandwidth" patented technology

Memory bandwidth is the rate at which data can be read from or stored into a semiconductor memory by a processor. Memory bandwidth is usually expressed in units of bytes/second, though this can vary for systems with natural data sizes that are not a multiple of the commonly used 8-bit bytes.

Instruction execution method and device for artificial intelligence chip

The invention discloses an instruction execution method and device for an artificial intelligence chip, and relates to the technical field of artificial intelligence chips, and the method comprises the steps: carrying out the deep learning driven feature recognition and resource demand prediction of an input task, and generating a demand prediction report of the task for computing resources through the analysis of a computational graph and a data dependency relationship of the task; a computing unit and memory resources are intelligently scheduled, an optimal instruction execution path is dynamically selected, and meanwhile a caching strategy is optimized; automatically generating a micro instruction set corresponding to the task according to the computing resource demand, the computing characteristic and the intelligent scheduling result of the task; when multiple tasks are executed in parallel, the execution sequence of the multiple tasks is dynamically adjusted according to the calculation load and the resource sharing condition of the tasks, and resource allocation is optimized. According to the method, the computing resources and the memory bandwidth required by each task can be accurately predicted through the deep learning driving analysis of the task computing graph, so that the allocation of the computing resources is optimized.
Owner:BEIJING LEKAIWENYU TECHNOLOGY CO LTD

Security isolation method for edge computing nodes of Internet of Things

The invention relates to the technical field of computer security, in particular to a security isolation method for an edge computing node of the Internet of Things, which comprises the following steps of: identifying a trust domain in the edge computing node of the Internet of Things, defining a sensitive data source, and generating a sensitive data identifier set, and based on the sensitive data identifier set, planning minimum central processing unit time and maximum memory bandwidth for each isolation domain to form an isolation domain resource limit, and establishing initial security isolation configuration. According to the method, the identification mechanism based on the sensitive data identification set is introduced into the edge computing node of the Internet of Things, the definition of the trust domain can be realized, and the resource limit is constructed in combination with the minimum central processing unit time and the maximum memory bandwidth of each isolation domain, so that the basic distribution mode of computing resources is restrained, and the computing efficiency is improved. And resources are prevented from being occupied by high-frequency low-priority tasks. After a data entry is positioned on a system level and stain marks are applied according to sensitive attributes, sensitive data images have identifiable features on a memory level.
Owner:JIANGSU SHENGZHITUO INFORMATION ENGINEERING CO LTD

AI model training acceleration method and system based on computing power service scheduling

The invention provides an AI model training acceleration method and system based on computing power service scheduling, and relates to the technical field of computing power service.Firstly, historical training process data of an AI model and structural features of a current training task model are collected, a computing power demand prediction model is constructed, the structural features of the current training task model are input to generate a computing power resource demand distribution sequence, and a computing power demand prediction model is constructed; the method comprises the following steps of: training a demand change curve of each stage of calculation core quantity, memory bandwidth and data transmission rate, screening a matched computing power resource combination scheme from a computing power service cluster according to a sequence, scheduling computing nodes to execute a training task, and collecting actual computing power resource consumption data in real time; and finally, comparing actual data with the prediction sequence, generating a deviation value, and dynamically adjusting the number of enabled computational nodes, thereby realizing accurate prediction and dynamic optimization scheduling of the computing power demand, improving the utilization rate of computing power resources, and accelerating AI model training.
Owner:SICHUAN BOCHUANGHUI FRONTIER TECH CO LTD +1

Convolution operation method and device, electronic equipment and storage medium

The invention relates to the technical field of artificial intelligence chips, and provides a convolution operation method and device, electronic equipment and a storage medium, and the method comprises the steps: traversing convolution kernel elements, and determining a current to-be-loaded data block based on the shape of a target data block and the coordinates of the traversed current convolution kernel element block; based on the current to-be-loaded data block, determining a to-be-covered data block and a newly-added data block, and loading the newly-added data block into the shared memory; based on the initial position, reading the input data block from the shared memory, applying the input data block and the current convolution kernel element block, performing matrix multiply-accumulate operation, and accumulating an operation result to an output result; and determining a convolution operation result based on an output result obtained by accumulation after traversal is completed. According to the method and the device, only the newly added data blocks are loaded into the shared memory, so that repeated data loading can be avoided, the data volume loaded each time is reduced, and the memory bandwidth is saved.
Owner:SHANGHAI BIREN TECH CO LTD

Multi-thread high-throughput data flow channel separation method and system based on zero copy

The invention belongs to the technical field of data transmission and processing, and discloses a zero-copy-based multi-thread high-throughput data stream channel separation method and system, and the method comprises the steps: directly writing a mixed data stream into a front-end buffer region configured as an annular structure through a data receiving module by adopting direct memory access; then, a multi-thread processing module dynamically allocates a plurality of processing threads from a thread pool to separate channel data in parallel, each thread adopts a zero copy algorithm based on pointer offset, positions the channel data in a memory, creates pointer reference and associates the channel data to a corresponding rear-end buffer area, and logic separation is achieved without physical copy; and finally, the data storage module efficiently writes the separated data into persistent storage in an asynchronous I / O mode. According to the method, zero-copy, multi-thread parallel and two-stage dynamic buffering strategies are combined, the data separation efficiency is remarkably improved, CPU occupation and memory bandwidth are greatly reduced, and the real-time performance and stability of high-throughput data processing are guaranteed.
Owner:CHINA JILIANG UNIV

DCU-based high-performance sparse stiffness matrix vector multiplication method

The invention provides a DCU-based high-performance sparse stiffness matrix vector multiplication method, which comprises the following steps of: according to a sparse stiffness matrix, dividing a non-zero element into a plurality of calculation unit blocks by rows, pre-loading non-zero element data to an L1 shared memory or a register file through an on-chip shared memory controller of the DCU, a high-bandwidth crossbar switch of the DCU is used for realizing data copying and transmission; starting multi-row fusion execution for a short row of which the row non-zero element is lower than a DCU single-instruction multi-data width threshold value; constructing a wavefront scheduler based on a DCU asynchronous computing engine: binding an independent instruction cache region for each wavefront, and loading a multiply-add operation instruction set in advance through a prefetch instruction queue; a calculation unit state register is established, and when a wavefront scheduler ready signal is triggered, a scalar unit of the DCU is activated to execute calculation; and realizing cross-thread block reduction by adopting a DCU atomic operation accelerator. According to the method, the calculation throughput and the memory bandwidth utilization rate of large-scale structural mechanics stiffness matrix vector multiplication are effectively improved.
Owner:HENAN POLYTECHNIC

Multi-level page table traversal acceleration method, system, equipment and medium

The invention provides a multi-level page table traversal acceleration method, system and device and a medium, and belongs to the technical field of computers. The method comprises the following steps: inputting a virtual address into a path predictor and a dynamic weight hierarchical cache, and collecting an instruction access proportion, a TLB hit rate and a pre-fetching success rate in real time; performing hash calculation according to a preset bit of the virtual address through a path predictor, and querying a historical record table to obtain a prediction path and a confidence level so as to determine a query range; when the cache is hit, scores of hit entries are calculated, priority ranking and elimination decision making of the cache entries are carried out, and meanwhile, the cache entries are managed by adopting a hybrid replacement algorithm; under a preset condition, starting a prefetch operation through an adaptive prefetch engine, calculating and dynamically expanding a prefetch step length, and preventing memory bandwidth contention through two-stage monitoring; when the instruction access proportion, the TLB hit rate and the prefetching success rate meet conditions, an instruction priority mode is activated, and scores of hit entries are updated by adjusting related calculation parameters.
Owner:SHANDONG INSPUR SCI RES INST CO LTD

Memory writing method and device, storage medium and program product

The invention discloses a memory writing method and device, a storage medium and a program product, and relates to the technical field of memory access, and the method comprises the following steps: determining a memory step length of an initial image tensor corresponding to a target image, determining a target vectorization width according to the memory step length and a memory data width of a target graphics processor, and determining a target data dimension of the target graphics processor, determining a target thread grid corresponding to the initial image tensor according to the target data dimension, performing image tensor slicing by using the target thread grid to obtain a target image tensor, and writing the target image tensor into a memory. According to the method, the memory layout (step length) of the input tensor can be dynamically analyzed, so that the optimal vectorization width is adaptively determined, the memory access efficiency is maximized, slicing is performed according to the data dimension when the target graphics processor outputs the data, the method has adaptability and high efficiency, the GPU video memory bandwidth is fully utilized, and the memory bandwidth utilization rate is improved.
Owner:LANGCHAO ELECTRONIC INFORMATION IND CO LTD

BIOS parameter debugging method, program product, electronic equipment and storage medium

The invention discloses a BIOS parameter debugging method, a program product, electronic equipment and a storage medium, and relates to the technical field of BIOS tuning, the BIOS parameter debugging method comprises the following steps: respectively inputting each group of candidate BIOS parameter combinations into a hybrid model; obtaining at least one of a CPU utilization rate sub-score, a memory bandwidth sub-score, a delay sub-score and a power consumption sub-score corresponding to each group of candidate BIOS parameter combinations output when the hybrid model runs in the target business scene; and based on the CPU utilization rate sub-score, the memory bandwidth sub-score, the delay sub-score and the power consumption sub-score, determining a target BIOS parameter combination corresponding to the target business scene from each group of candidate BIOS parameter combinations. According to the BIOS parameter debugging method, the synergistic effect between hardware is comprehensively considered, and the technical effects of relatively high tuning efficiency, relatively good tuning effect and dynamic tuning capability are achieved.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Data processing method, product, electronic equipment and computer readable storage medium

The invention discloses a data processing method, a product, electronic equipment and a computer readable storage medium, relates to the technical field of computers, and aims to solve the problems of large GPU (Graphic Processing Unit) space occupation and high access delay in data processing in related technologies. Reading corresponding historical key value cache data from a key value cache of the memory extension equipment according to a historical key value cache data acquisition request sent by a host end; returning the historical key value cache data to the host end, and enabling the host end to process the current input data based on the historical key value cache data by adopting a local model to obtain key value data; and obtaining key value cache data corresponding to the key value data sent by the host end, and sending the key value cache data to a key value cache of the memory extension equipment for storage. The key value cache data is stored in the key value cache of the memory extension equipment, and the historical key value cache data is read from the key value cache, so that the occupation of the memory space of the GPU and the consumption of the memory bandwidth can be reduced, and the access delay is reduced.
Owner:LANGCHAO ELECTRONIC INFORMATION IND CO LTD

Artificial intelligence model training resource adaptive distribution system

The invention belongs to the technical field of artificial intelligence, and discloses an artificial intelligence model training resource adaptive distribution system. The method comprises the following steps: acquiring and calculating graph structure data and hardware resource state data in real time, and calculating a data reuse rate and generating a candidate operator fusion scheme by constructing an operator execution time sequence constraint matrix and identifying an operator cluster of data locality characteristics; a resource competition hotspot prediction mechanism is introduced, memory bandwidth occupation fluctuation characteristics are analyzed, a resource conflict probability is calculated for a fusion scheme, and a dynamic balance optimization model of fusion income and resource conflicts is constructed. An optimal operator fusion decision sequence and a resource allocation strategy are generated through iterative solution, and accurate dynamic adjustment of computing resources in the training process is achieved. The training efficiency and the resource utilization rate are improved, the energy consumption is reduced, and the system stability is enhanced.
Owner:YANGZHOU HUAZHISHENG INFORMATION TECHNOLOGY CO LTD

Memory access method and device, storage medium and program product

The invention discloses a memory access method and device, a storage medium and a program product, and relates to the technical field of memory access, comprising: monitoring memory access information of a target memory; the memory access information comprises a missing page address sequence during memory access, an address distribution characteristic when an address conversion failure event occurs in an address conversion lookaside buffer, a cross-node access event and a virtual memory access frequency; determining corresponding target feature information; the target feature information comprises spatial locality, thermal density, access dispersion and unbalance degree data; and determining a current memory access mode based on the target feature information, and determining a target access strategy from preset memory access strategies configured with different memory page table prefetching rules to access the memory. The memory access feature information is quantified through the memory access information, the corresponding memory access mode is determined, and the corresponding memory page table prefetching rule is executed, so that the traversal delay of the multi-level page table can be reduced, and the memory access performance is improved by fully utilizing the memory bandwidth.
Owner:LANGCHAO ELECTRONIC INFORMATION IND CO LTD

FPGA-based high-performance large language model accelerator and reasoning method

The invention discloses an FPGA (Field Programmable Gate Array)-based high-performance large language model accelerator and a reasoning method, which adopt a structure of combining a plurality of computing units (CU) and a matrix processing unit (MPE), and can efficiently allocate computing tasks and realize parallel processing by virtue of the parallel computing capability of the FPGA, thereby greatly improving the computing efficiency. Through parallel processing of a plurality of tasks, the reasoning speed is obviously improved, and the calculation bottleneck in a traditional scheme is avoided. And secondly, in the aspect of memory management, through a hybrid storage strategy of a high-bandwidth memory (HBM) and an off-chip memory (such as DDR), utilization of memory bandwidth is optimized, efficient flowing of data among a plurality of computing units is realized, delay in data transmission is reduced, and efficient data access in the computing process is ensured.
Owner:BEIJING UNIV OF POSTS & TELECOMM

Transformer models with optimized first layer

This specification discloses systems and methods for enhancing the efficiency of transformer models during inference and training by precomputing and storing in memory a significant portion of operations in the first transformer layer. The stored precomputed outputs are retrieved from memory during runtime, reducing computational complexity and memory bandwidth requirements. This approach results in decreased latency, increased throughput, and lower cost-per-token. The disclosed techniques are particularly advantageous for transformer models that incorporate positional encodings within the attention mechanism, such as Rotary Position Embedding (RoPE) and other relative position encoding schemes. The method of offline precomputing involves calculating the outputs of the eliminated operations and components for each of the original vocab_size embedding-vectors, where vocab_size is the size of the embedding vocabulary. One embodiment of the invention removes the feedforward network and the attention query, key, and value projections from the first transformer layer of the encoder and the decoder stacks.
Owner:GRAEF NILS

ARM-based embedded image decoding display system

The invention discloses an embedded image decoding display system based on an ARM, and relates to the technical field of computers, and the system comprises a dynamic sensing module which is used for monitoring the CPU occupancy rate, the cache hit rate and the memory bandwidth utilization rate of an ARM processor in real time, establishing a three-dimensional resource vector containing a dynamic weight coefficient, and calculating a threshold boundary; the adaptive decoding engine comprises an optimization unit which is used for reconstructing a bit stream processing channel and a register data exchange mechanism based on an SIMD expansion instruction of an ARMv8.2 instruction set; the prediction unit is used for predicting instruction-level hotspot distribution of the decoding task by adopting a TinyLSTM model and generating an instruction transmitting strategy; the hardware optimization module is used for executing a heterogeneous collaboration strategy of the NEON coprocessor and the MaliGPU, and the heterogeneous collaboration strategy comprises the following steps: dividing DCT / IDCT (discrete cosine transform / inverse discrete cosine transform) calculation granularity; determining a mixing precision conversion assembly line of the GPU shader according to the color gamut of the display equipment; and the display driving module is used for implementing a decoding parameter dynamic confusion algorithm in the security domain and forming a tamper-proof closed loop with the optimization unit.
Owner:TRONLONG

Data processing method and device for large model parameters, equipment and medium

The embodiment of the invention provides a data processing method and device for large model parameters, equipment and a medium, and belongs to the technical field of artificial intelligence. The method comprises the following steps: obtaining a quantization matrix and a meta-parameter of a target large model from different partitions of a first data storage area of a memory, reading memory access granularity representing the data processing capability of a processor, and determining the size of a filling area between the insertion meta-parameter and the quantization matrix based on the memory access granularity; dividing a second data storage area different from the first data storage area in the memory, and sequentially writing the meta-parameter, preset filling information matched with the filling area in size and the quantization matrix into the second data storage area to obtain an initial data atomic block; and continuously arranging all the initial data atomic blocks in the second data storage area to obtain a target data atomic block. By optimizing the data storage mode, the memory bandwidth utilization rate of the processor in the reasoning process is improved, and then the data calculation efficiency of the processor is improved.
Owner:PENG CHENG LAB

Semiconductor assembly for providing an enhanced memory bandwidth and methods for forming the same

A plurality of processor dies may be attached to an interposer structure including interposer dielectric material layers having formed therein interposer metal interconnect structures, A dielectric matrix may be formed around the plurality of processor dies over the interposer structure. A router die may be attached to the plurality of processor dies. The router die includes router dielectric material layers having formed therein router metal interconnect structures and a router substrate having formed therein router through-substrate via structures therein. A backside of the router substrate may be thinned to expose end surfaces of the router through-substrate via structures. Memory dies may be attached to the router substrate after thinning the backside of the router substrate. Bonding pads of the memory dies are electrically connected to the router through-substrate via structures.
Owner:TAIWAN SEMICONDUCTOR MANUFACTURING CO LTD

Model performance test method and device, electronic equipment and storage medium

The invention discloses a model performance test method and device, electronic equipment and a storage medium, and relates to the technical field of artificial intelligence. The theoretical maximum lexical throughput of a target large language model is calculated based on the video memory bandwidth of a graphics processor, the model parameter quantity, the byte number corresponding to the quantization precision and the video memory bandwidth utilization rate; meanwhile, the benchmark performance throughput is obtained, a theoretical corresponding first concurrency number is calculated in combination with the theoretical maximum lexical unit throughput and the concurrency competition loss coefficient, then the model test is executed based on the first concurrency number to obtain the actual maximum lexical unit throughput and a corresponding second concurrency number, and a model performance test result is generated. The problems that in the prior art, due to the fact that manual testing is conducted depending on manual intervention, a continuous approaching attempt mode is adopted, a reasonable test starting point is not deduced in combination with hardware core bottlenecks and key parameters, evaluation is time-consuming and labor-consuming, the result is prone to being affected by artificial factors, and accuracy and consistency are poor can be solved.
Owner:JINAN INSPUR DATA TECH CO LTD

Method for accelerating secure metadata access in secure memory system, memory controller and system

The invention discloses a method for accelerating secure metadata access in a secure memory system, a memory controller and a system, and belongs to the field of secure memory systems, and the method comprises the following steps: when a page table item corresponding to a logic page where data accessed by a processor is located does not hit a TLB, obtaining the page table item from a memory page table, extracting a physical page address from the page table item, and storing the physical page address in a memory; a counter corresponding to a physical page where the data to be accessed is located and a father node of the counter in the integrity tree are prefetched through the physical page address; adding a replacement dirty block address in a miss request sent by the last level of cache, after receiving the miss request containing a field of the replacement dirty block address, executing conventional memory reading and decryption, positioning a counter corresponding to the replacement dirty block address, and performing prefetching by using an idle memory bandwidth; in addition to the secure metadata cache, the prefetching queue is maintained to temporarily store the prefetched metadata. The cache hit rate of the security metadata in the security memory system can be improved, and the performance overhead caused by the cache miss can be reduced.
Owner:HUAZHONG UNIV OF SCI & TECH

Semiconductor device for providing improved memory bandwidth and method for forming the same

A plurality of processor dies may be attached to an interposer structure including interposer dielectric material layers in which interposer metal interconnect structures are formed. A dielectric matrix may be formed around the plurality of processor dies over the interposer structure. A router die may be attached to the plurality of processor dies. The router die includes router dielectric material layers in which router metal interconnect structures are formed and a router substrate in which router substrate via structures are formed. A backside of the router substrate may be thinned to expose end surfaces of the router substrate via structures. Memory dies may be attached to the router substrate after thinning the backside of the router substrate. The bonding pads of the memory dies are electrically connected to the router substrate via structures.
Owner:TAIWAN SEMICONDUCTOR MANUFACTURING CO LTD

Layered self-adaptive full block pre-filling scheduling method and system for large language model reasoning

The invention discloses a hierarchical self-adaptive full block pre-filling scheduling method and system for large language model reasoning, and the method comprises the steps: carrying out the hierarchical portrait analysis of a to-be-served model, and dividing the to-be-served model into partitions with different calculation characteristics according to the calculation intensity and memory access characteristics of each layer; then, making a layering and partitioning strategy based on a partitioning result, allocating a larger partitioning size to a calculation-intensive partition, allocating a smaller partitioning size to a memory bandwidth-intensive partition, and generating a layering and partitioning mapping table; and finally, when the online scheduling is executed, querying the mapping table according to the request processing progress to determine the block target size, and jointly forming a batch processing unit by the decoding task and the pre-filled block with the heterogeneous size under the constraint of the iteration time budget to be executed. According to the method, accurate matching of calculation and bandwidth resources is achieved, the system throughput can be effectively improved, tail delay and fluctuation thereof can be remarkably reduced, bubbles under pipeline parallelism are reduced, and the method is suitable for various attention mechanisms and distributed reasoning architectures.
Owner:ZHEJIANG LAB

Optimization calculation method and device for attention mechanism

The invention provides an attention mechanism optimization calculation method and device, and relates to the technical field of artificial intelligence, and the method comprises the steps: constructing a packaging mask tensor of a target batch based on the length information of a plurality of input sequences in the target batch; when attention weight calculation is executed on the target batch based on the calculation unit, the packaged mask tensor and the attention score tensor are calculated, and the attention score tensor after mask processing is obtained; and determining an attention calculation result of the target batch based on the attention score tensor after mask processing. According to the method, real-time dynamic judgment on the effectiveness of sequence elements in the attention calculation process is replaced by pre-constructing the packaged mask tensor, and complex conditional judgment logic is converted into simple tensor operation. According to the invention, the branch prediction overhead and thread differentiation in the calculation process are greatly reduced, the parallel processing efficiency of the calculation unit is improved, and the occupation of memory bandwidth is reduced, so that the calculation efficiency of the attention mechanism is improved.
Owner:SHANGHAI BIREN TECH CO LTD

Power management method, device and equipment for graphics card and storage medium

The invention relates to the technical field of video card power management, and discloses a power management method, device and equipment for a video card and a storage medium. The method comprises the following steps: collecting time sequence data of SM unit frequency of a graphics card, GDDR video memory bandwidth and VRM output voltage in real time, and establishing a GPU power consumption topology model; based on the GPU power consumption topology model, extracting a time sequence feature vector of the load change of the rendering pipeline; constructing a sparse representation matrix based on the time sequence feature vector; performing power consumption gradient threshold judgment and delay type classification by using the sparse representation matrix, and identifying a delay fault type in the current graphics card power management system; and adjusting a switching parameter of the GPU, a VRM compensation current and a fan rotating speed according to the identified delay fault type, and updating a delay detection threshold value and a repair strategy strength parameter. According to the invention, three different delay fault types of frequency climbing, voltage regulation and temperature control current limiting can be accurately distinguished, and continuous optimization of power management performance is realized.
Owner:SHENZHEN XIANGSHENG INTELLIGENT MANUFACTURING CO LTD

Tensor transpose processor

The present invention relates to a processor designed to optimize memory bandwidth utilization for tensor transpositions in machine learning. An example processor includes an input tensor shift buffer, a staging buffer, and an output tensor shift buffer. The input tensor shift buffer reads an input tensor from input memory and performs multiple cycles of input tensor shifting. The shifted tensor data is then written into the staging buffer. The output tensor shift buffer reads the shifted tensor data from the staging buffer and performs multiple cycles of output tensor shifting. Finally, the result is written to the output memory. This configuration facilitates efficient handling and transformation of tensor data, optimizing the computational processes required in machine learning tasks.
Owner:MOFFETT TECH CO LTD

Industrial multi-protocol heterogeneous real-time video collaborative analysis method, equipment and medium

The invention provides an industrial multi-protocol heterogeneous real-time video collaborative analysis method. The method comprises the steps that a lightweight compiler is developed, an operator dependency graph analysis module is integrated, instruction-level parallelism of a neural network processing unit is predicted through a graph neural network, and optimization intermediate representation is dynamically generated; constructing a protocol classification graph neural network model, carrying out protocol type self-identification and streaming media metadata automatic extraction, and converting a heterogeneous video stream into a standardized frame sequence; designing a multi-agent reinforcement learning model, and generating an optimal scheduling strategy through offline simulation training by taking a neural network processing unit calculation unit utilization rate, a memory bandwidth occupancy rate and a task queue depth as state spaces; predicting a video key region by adopting a hybrid network, and performing resolution downsampling on a non-key region in combination with motion vector analysis; the hardware-level scheduling adopts time slice rotation preemptive scheduling, and time slices are dynamically distributed according to the priority of an algorithm.
Owner:BEIJING ENGINEERING DIGITAL INTELLIGENCE (BEIJING) TECHNOLOGY CO LTD

Affine motion model restrictions for memory bandwidth reduction of enhanced interpolation filter

A method for coding a video implemented in an encoder or a decoder including the enhanced interpolation filter, EIF, for motion compensation, the method comprising: i) determining control point motion vectors, CPMVs, for a block according to affine inter-prediction, the block being an affine block or a sub-block of the affine block; ii) for a predefined sub-block size determining a reference area for a sub-block with the predefined sub-block size according to values of the CPMVs; iii) comparing the determined reference area with a predefined threshold; iv) applying EIF for motion compensation, comprising deriving the pixel-based motion vector field for the block; wherein if the determined reference area is larger than the threshold, deriving the pixel-based motion vector field for the block further comprises motion vector clipping, wherein motion vector clipping range is determined based on motion model of the block and the size of the block.
Owner:HUAWEI TECH CO LTD

Vector and matrix calculation-oriented memory access system

The invention provides a memory access system oriented to vector and matrix calculation, the system comprises a vector memory access unit, a matrix memory access unit, a vector register group and a matrix register group, the vector memory access unit is connected with a memory interface and the vector register group, reads elements of a one-dimensional data structure or a two-dimensional data structure from a memory, and stores the elements of the one-dimensional data structure or the two-dimensional data structure; vector data are generated through data reorganization operation, a matrix access unit is connected with a memory interface and a matrix register set, matrix block data of a two-dimensional data structure are read from a memory together with the matrix access unit, matrix data are generated after data reorganization, and the matrix data are broadcasted to one or more computing units according to rows or columns. And performing calculation on the matrix data and the vector data. In the memory access system, the vector memory access unit and the matrix memory access unit can load data in parallel, and the utilization rate is improved through a plurality of computing units, so that the problems of low memory access efficiency and low memory bandwidth utilization rate are solved.
Owner:NANJING UNIV

Bandwidth adjustment method and device, storage medium and electronic equipment

The invention provides a bandwidth adjustment method and device, a storage medium and electronic equipment. The method comprises the following steps: identifying a user state scene; wherein the user state scene is associated with the current process number and the current process identifier; and sending a bandwidth adjustment instruction to a system on chip (SOC), so that the SOC adjusts the memory bandwidth to the target bandwidth after determining the target bandwidth matched with the user state scene. According to the method and the device, the rationality of memory bandwidth allocation can be improved, the performance and the cruising ability of the electronic equipment can be improved, and the availability is high.
Owner:BEIJING XIAOMI MOBILE SOFTWARE CO LTD

Extending dynamic resource controllers for sub-NUMA clustering patterns

The invention relates to extending a dynamic resource controller for a sub-NUMA clustering pattern. Examples include an apparatus having: a plurality of clusters of processor cores; a plurality of memory bandwidth allocators, each cluster corresponding to at least one of the plurality of memory bandwidth allocators to apply memory bandwidth settings to the processor cores of an associated cluster to dynamically adjust the priority of memory bandwidth allocated for workloads to be processed by the processor cores of the associated cluster; a plurality of memory controllers, each cluster of processor cores corresponding to at least one of the plurality of memory controllers; a plurality of performance monitors, each cluster corresponding to at least one of the plurality of performance monitors to generate performance monitoring statistics by monitoring performance of a workload processed by the processor core based at least in part on the performance monitoring configuration parameters; and a hardware dynamic resource controller configurable as a plurality of virtual dynamic resource controllers.
Owner:INTEL CORP

Post-training calibration for activation sparsity

The first token prediction of a large language model is bottlenecked by compute and second token predictions onwards are bottlenecked by memory bandwidth. Inferences can be made more efficient through activation sparsity. An activation tensor is pruned using an importance threshold value. The mode of the activation tensor is centered in a lossless manner using an estimated mode value to improve activation sparsity further. Pruning and mode-centering mechanisms can be inserted into a neural network strategically and post-training to implement sparsification. A two-stage greedy grid search algorithm is implemented to determine the calibrated importance threshold values of various pruners and the estimated mode values using a small dataset. A modified neural network with pruning and lossless mode-centering can be deployed onto hardware.
Owner:INTEL CORP +6