Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

230 results about "Memory bandwidth" patented technology

Memory bandwidth is the rate at which data can be read from or stored into a semiconductor memory by a processor. Memory bandwidth is usually expressed in units of bytes/second, though this can vary for systems with natural data sizes that are not a multiple of the commonly used 8-bit bytes.

DCU-based high-performance sparse stiffness matrix vector multiplication method

The invention provides a DCU-based high-performance sparse stiffness matrix vector multiplication method, which comprises the following steps of: according to a sparse stiffness matrix, dividing a non-zero element into a plurality of calculation unit blocks by rows, pre-loading non-zero element data to an L1 shared memory or a register file through an on-chip shared memory controller of the DCU, a high-bandwidth crossbar switch of the DCU is used for realizing data copying and transmission; starting multi-row fusion execution for a short row of which the row non-zero element is lower than a DCU single-instruction multi-data width threshold value; constructing a wavefront scheduler based on a DCU asynchronous computing engine: binding an independent instruction cache region for each wavefront, and loading a multiply-add operation instruction set in advance through a prefetch instruction queue; a calculation unit state register is established, and when a wavefront scheduler ready signal is triggered, a scalar unit of the DCU is activated to execute calculation; and realizing cross-thread block reduction by adopting a DCU atomic operation accelerator. According to the method, the calculation throughput and the memory bandwidth utilization rate of large-scale structural mechanics stiffness matrix vector multiplication are effectively improved.
Owner:HENAN POLYTECHNIC

Data processing method, product, electronic equipment and computer readable storage medium

The invention discloses a data processing method, a product, electronic equipment and a computer readable storage medium, relates to the technical field of computers, and aims to solve the problems of large GPU (Graphic Processing Unit) space occupation and high access delay in data processing in related technologies. Reading corresponding historical key value cache data from a key value cache of the memory extension equipment according to a historical key value cache data acquisition request sent by a host end; returning the historical key value cache data to the host end, and enabling the host end to process the current input data based on the historical key value cache data by adopting a local model to obtain key value data; and obtaining key value cache data corresponding to the key value data sent by the host end, and sending the key value cache data to a key value cache of the memory extension equipment for storage. The key value cache data is stored in the key value cache of the memory extension equipment, and the historical key value cache data is read from the key value cache, so that the occupation of the memory space of the GPU and the consumption of the memory bandwidth can be reduced, and the access delay is reduced.
Owner:LANGCHAO ELECTRONIC INFORMATION IND CO LTD

Artificial intelligence model training resource adaptive distribution system

The invention belongs to the technical field of artificial intelligence, and discloses an artificial intelligence model training resource adaptive distribution system. The method comprises the following steps: acquiring and calculating graph structure data and hardware resource state data in real time, and calculating a data reuse rate and generating a candidate operator fusion scheme by constructing an operator execution time sequence constraint matrix and identifying an operator cluster of data locality characteristics; a resource competition hotspot prediction mechanism is introduced, memory bandwidth occupation fluctuation characteristics are analyzed, a resource conflict probability is calculated for a fusion scheme, and a dynamic balance optimization model of fusion income and resource conflicts is constructed. An optimal operator fusion decision sequence and a resource allocation strategy are generated through iterative solution, and accurate dynamic adjustment of computing resources in the training process is achieved. The training efficiency and the resource utilization rate are improved, the energy consumption is reduced, and the system stability is enhanced.
Owner:YANGZHOU HUAZHISHENG INFORMATION TECHNOLOGY CO LTD

Memory access method and device, storage medium and program product

The invention discloses a memory access method and device, a storage medium and a program product, and relates to the technical field of memory access, comprising: monitoring memory access information of a target memory; the memory access information comprises a missing page address sequence during memory access, an address distribution characteristic when an address conversion failure event occurs in an address conversion lookaside buffer, a cross-node access event and a virtual memory access frequency; determining corresponding target feature information; the target feature information comprises spatial locality, thermal density, access dispersion and unbalance degree data; and determining a current memory access mode based on the target feature information, and determining a target access strategy from preset memory access strategies configured with different memory page table prefetching rules to access the memory. The memory access feature information is quantified through the memory access information, the corresponding memory access mode is determined, and the corresponding memory page table prefetching rule is executed, so that the traversal delay of the multi-level page table can be reduced, and the memory access performance is improved by fully utilizing the memory bandwidth.
Owner:LANGCHAO ELECTRONIC INFORMATION IND CO LTD

Data processing method and device for large model parameters, equipment and medium

The embodiment of the invention provides a data processing method and device for large model parameters, equipment and a medium, and belongs to the technical field of artificial intelligence. The method comprises the following steps: obtaining a quantization matrix and a meta-parameter of a target large model from different partitions of a first data storage area of a memory, reading memory access granularity representing the data processing capability of a processor, and determining the size of a filling area between the insertion meta-parameter and the quantization matrix based on the memory access granularity; dividing a second data storage area different from the first data storage area in the memory, and sequentially writing the meta-parameter, preset filling information matched with the filling area in size and the quantization matrix into the second data storage area to obtain an initial data atomic block; and continuously arranging all the initial data atomic blocks in the second data storage area to obtain a target data atomic block. By optimizing the data storage mode, the memory bandwidth utilization rate of the processor in the reasoning process is improved, and then the data calculation efficiency of the processor is improved.
Owner:PENG CHENG LAB

Model performance test method and device, electronic equipment and storage medium

The invention discloses a model performance test method and device, electronic equipment and a storage medium, and relates to the technical field of artificial intelligence. The theoretical maximum lexical throughput of a target large language model is calculated based on the video memory bandwidth of a graphics processor, the model parameter quantity, the byte number corresponding to the quantization precision and the video memory bandwidth utilization rate; meanwhile, the benchmark performance throughput is obtained, a theoretical corresponding first concurrency number is calculated in combination with the theoretical maximum lexical unit throughput and the concurrency competition loss coefficient, then the model test is executed based on the first concurrency number to obtain the actual maximum lexical unit throughput and a corresponding second concurrency number, and a model performance test result is generated. The problems that in the prior art, due to the fact that manual testing is conducted depending on manual intervention, a continuous approaching attempt mode is adopted, a reasonable test starting point is not deduced in combination with hardware core bottlenecks and key parameters, evaluation is time-consuming and labor-consuming, the result is prone to being affected by artificial factors, and accuracy and consistency are poor can be solved.
Owner:JINAN INSPUR DATA TECH CO LTD

Method for accelerating secure metadata access in secure memory system, memory controller and system

The invention discloses a method for accelerating secure metadata access in a secure memory system, a memory controller and a system, and belongs to the field of secure memory systems, and the method comprises the following steps: when a page table item corresponding to a logic page where data accessed by a processor is located does not hit a TLB, obtaining the page table item from a memory page table, extracting a physical page address from the page table item, and storing the physical page address in a memory; a counter corresponding to a physical page where the data to be accessed is located and a father node of the counter in the integrity tree are prefetched through the physical page address; adding a replacement dirty block address in a miss request sent by the last level of cache, after receiving the miss request containing a field of the replacement dirty block address, executing conventional memory reading and decryption, positioning a counter corresponding to the replacement dirty block address, and performing prefetching by using an idle memory bandwidth; in addition to the secure metadata cache, the prefetching queue is maintained to temporarily store the prefetched metadata. The cache hit rate of the security metadata in the security memory system can be improved, and the performance overhead caused by the cache miss can be reduced.
Owner:HUAZHONG UNIV OF SCI & TECH

Layered self-adaptive full block pre-filling scheduling method and system for large language model reasoning

The invention discloses a hierarchical self-adaptive full block pre-filling scheduling method and system for large language model reasoning, and the method comprises the steps: carrying out the hierarchical portrait analysis of a to-be-served model, and dividing the to-be-served model into partitions with different calculation characteristics according to the calculation intensity and memory access characteristics of each layer; then, making a layering and partitioning strategy based on a partitioning result, allocating a larger partitioning size to a calculation-intensive partition, allocating a smaller partitioning size to a memory bandwidth-intensive partition, and generating a layering and partitioning mapping table; and finally, when the online scheduling is executed, querying the mapping table according to the request processing progress to determine the block target size, and jointly forming a batch processing unit by the decoding task and the pre-filled block with the heterogeneous size under the constraint of the iteration time budget to be executed. According to the method, accurate matching of calculation and bandwidth resources is achieved, the system throughput can be effectively improved, tail delay and fluctuation thereof can be remarkably reduced, bubbles under pipeline parallelism are reduced, and the method is suitable for various attention mechanisms and distributed reasoning architectures.
Owner:ZHEJIANG LAB

Optimization calculation method and device for attention mechanism

The invention provides an attention mechanism optimization calculation method and device, and relates to the technical field of artificial intelligence, and the method comprises the steps: constructing a packaging mask tensor of a target batch based on the length information of a plurality of input sequences in the target batch; when attention weight calculation is executed on the target batch based on the calculation unit, the packaged mask tensor and the attention score tensor are calculated, and the attention score tensor after mask processing is obtained; and determining an attention calculation result of the target batch based on the attention score tensor after mask processing. According to the method, real-time dynamic judgment on the effectiveness of sequence elements in the attention calculation process is replaced by pre-constructing the packaged mask tensor, and complex conditional judgment logic is converted into simple tensor operation. According to the invention, the branch prediction overhead and thread differentiation in the calculation process are greatly reduced, the parallel processing efficiency of the calculation unit is improved, and the occupation of memory bandwidth is reduced, so that the calculation efficiency of the attention mechanism is improved.
Owner:SHANGHAI BIREN TECH CO LTD

Power management method, device and equipment for graphics card and storage medium

The invention relates to the technical field of video card power management, and discloses a power management method, device and equipment for a video card and a storage medium. The method comprises the following steps: collecting time sequence data of SM unit frequency of a graphics card, GDDR video memory bandwidth and VRM output voltage in real time, and establishing a GPU power consumption topology model; based on the GPU power consumption topology model, extracting a time sequence feature vector of the load change of the rendering pipeline; constructing a sparse representation matrix based on the time sequence feature vector; performing power consumption gradient threshold judgment and delay type classification by using the sparse representation matrix, and identifying a delay fault type in the current graphics card power management system; and adjusting a switching parameter of the GPU, a VRM compensation current and a fan rotating speed according to the identified delay fault type, and updating a delay detection threshold value and a repair strategy strength parameter. According to the invention, three different delay fault types of frequency climbing, voltage regulation and temperature control current limiting can be accurately distinguished, and continuous optimization of power management performance is realized.
Owner:SHENZHEN XIANGSHENG INTELLIGENT MANUFACTURING CO LTD

Tensor transpose processor

The present invention relates to a processor designed to optimize memory bandwidth utilization for tensor transpositions in machine learning. An example processor includes an input tensor shift buffer, a staging buffer, and an output tensor shift buffer. The input tensor shift buffer reads an input tensor from input memory and performs multiple cycles of input tensor shifting. The shifted tensor data is then written into the staging buffer. The output tensor shift buffer reads the shifted tensor data from the staging buffer and performs multiple cycles of output tensor shifting. Finally, the result is written to the output memory. This configuration facilitates efficient handling and transformation of tensor data, optimizing the computational processes required in machine learning tasks.
Owner:MOFFETT TECH CO LTD

Affine motion model restrictions for memory bandwidth reduction of enhanced interpolation filter

A method for coding a video implemented in an encoder or a decoder including the enhanced interpolation filter, EIF, for motion compensation, the method comprising: i) determining control point motion vectors, CPMVs, for a block according to affine inter-prediction, the block being an affine block or a sub-block of the affine block; ii) for a predefined sub-block size determining a reference area for a sub-block with the predefined sub-block size according to values of the CPMVs; iii) comparing the determined reference area with a predefined threshold; iv) applying EIF for motion compensation, comprising deriving the pixel-based motion vector field for the block; wherein if the determined reference area is larger than the threshold, deriving the pixel-based motion vector field for the block further comprises motion vector clipping, wherein motion vector clipping range is determined based on motion model of the block and the size of the block.
Owner:HUAWEI TECH CO LTD

Vector and matrix calculation-oriented memory access system

The invention provides a memory access system oriented to vector and matrix calculation, the system comprises a vector memory access unit, a matrix memory access unit, a vector register group and a matrix register group, the vector memory access unit is connected with a memory interface and the vector register group, reads elements of a one-dimensional data structure or a two-dimensional data structure from a memory, and stores the elements of the one-dimensional data structure or the two-dimensional data structure; vector data are generated through data reorganization operation, a matrix access unit is connected with a memory interface and a matrix register set, matrix block data of a two-dimensional data structure are read from a memory together with the matrix access unit, matrix data are generated after data reorganization, and the matrix data are broadcasted to one or more computing units according to rows or columns. And performing calculation on the matrix data and the vector data. In the memory access system, the vector memory access unit and the matrix memory access unit can load data in parallel, and the utilization rate is improved through a plurality of computing units, so that the problems of low memory access efficiency and low memory bandwidth utilization rate are solved.
Owner:NANJING UNIV

Post-training calibration for activation sparsity

The first token prediction of a large language model is bottlenecked by compute and second token predictions onwards are bottlenecked by memory bandwidth. Inferences can be made more efficient through activation sparsity. An activation tensor is pruned using an importance threshold value. The mode of the activation tensor is centered in a lossless manner using an estimated mode value to improve activation sparsity further. Pruning and mode-centering mechanisms can be inserted into a neural network strategically and post-training to implement sparsification. A two-stage greedy grid search algorithm is implemented to determine the calibrated importance threshold values of various pruners and the estimated mode values using a small dataset. A modified neural network with pruning and lossless mode-centering can be deployed onto hardware.
Owner:INTEL CORP +6

Hybrid bonding 3D stacked accelerator and acceleration method oriented to Transform reasoning

The invention belongs to the field of Transform model acceleration, and discloses a hybrid bonding 3D stacked accelerator and acceleration method oriented to Transform reasoning, the accelerator adopts a 3D stacked architecture, and comprises a DRAM storage bare chip and a logic bare chip, the DRAM storage bare chip comprises 16 DRAM storage blocks, the logic bare chip comprises 16 processing groups, each processing group corresponds to one DRAM storage block, and the processing groups correspond to the DRAM storage blocks. Each processing group is integrated with two attention processing units, and the DRAM storage blocks are connected with the processing groups through 32 copper-copper hybrid bonding vertical channels. According to the method, the problems of dense calculation and high memory bandwidth requirement in the reasoning process are solved, efficient model reasoning acceleration is realized, and the method is suitable for natural language processing, computer vision, video analysis and other application scenes depending on the Transform model.
Owner:NANJING UNIV OF POSTS & TELECOMM

System memory peak bandwidth measurement method and electronic equipment

The invention discloses a system memory peak bandwidth measurement method and electronic equipment, and relates to the technical field of bandwidth measurement, the memory bandwidth measurement process is optimized through cooperation of dynamic selection of a maximum width vector instruction, forced non-cache access and NUMA perception binding, and the bandwidth measurement efficiency is improved by utilizing the vector processing capacity of a CPU (Central Processing Unit). The data throughput of a single operation is improved to 512 bits or even higher, through cooperation of a non-temporary instruction and a memory barrier instruction, a cache level is bypassed, interference of cache hit or jitter on a measurement result is eliminated, it is ensured that the measurement result truly reflects the performance of a memory system, and the measurement accuracy is improved. Through an automatic NUMA binding mechanism, delay and congestion caused by cross-node access are avoided, and the accuracy, stability and repeatability of a test result are improved.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Pipelined read-modify-write operations in cache memory

Providing memory bandwidth compression using compressed memory controllers (CMCs) in a central processing unit (CPU)-based system is disclosed. In this regard, in some aspects, a CMC is configured to receive a memory read request to a physical address in a system memory, and read a compression indicator (CI) for the physical address from a master directory and / or from error correcting code (ECC) bits of the physical address. Based on the CI, the CMC determines a number of memory blocks to be read for the memory read request, and reads the determined number of memory blocks. In some aspects, a CMC is configured to receive a memory write request to a physical address in the system memory, and generate a CI for write data based on a compression pattern of the write data. The CMC updates the master directory and / or the ECC bits of the physical address with the generated CI.
Owner:TEXAS INSTRUMENTS INC

Server multi-core computing processor load optimization test method, device and equipment and medium

The invention relates to the technical field of server testing, in particular to a server multi-core computing processor load optimization testing method, device and equipment and a medium, and the method comprises the following steps: automatically identifying hardware configuration information of a server; generating a multi-dimensional pressure test plan based on the hardware configuration information; executing the test plan, and monitoring the load rate, the memory bandwidth, the cache hit rate and the I / O performance index of each CPU core in real time; dynamically adjusting a task allocation strategy based on a set load difference threshold to realize task dynamic optimization, and if the difference between the maximum core load rate and the minimum core load rate is monitored to exceed the threshold, triggering a task migration operation; system performance data are collected again, performance indexes before and after optimization are compared and analyzed, and system performance bottlenecks are recognized; and automatically generating a performance test report. The utilization rate and the overall task throughput of the multi-core processor are improved, and the problem that a static task allocation strategy is difficult to adapt to dynamic load changes is solved.
Owner:SHANDONG CHAOYUE DATA CONTROL ELECTRONICS CO LTD

Efficient dynamics simulation analysis method based on Fourier neural operator

The invention belongs to the technical field of model simulation, and particularly relates to an efficient dynamics simulation analysis method based on a Fourier neural operator. The present invention proposes FNO-Speed, and a series of comprehensive solutions for inefficient operations that the FNO solver does not fully utilize hardware. According to the method, two unique optimization methods are adopted, and comprise a multi-level parallel implicit image-to-column general matrix multiplication optimization strategy and a user-defined size high-frequency signal filtering algorithm. According to the method, efficient general matrix multiplication is achieved through an implicit image-to-column and data division strategy to replace pointwise convolution, and fragmentary calculation of frequency domain local linear transformation is eliminated through the latter. The FNO-Speed makes full use of the memory bandwidth, improves the calculation efficiency, and aims to solve the problems of low utilization rate of calculation resources and delay influence caused by large-scale data access calculation.
Owner:CHINA UNIV OF PETROLEUM (EAST CHINA)

Processing apparatus, storage management method, and related device

Disclosed in embodiments of the present application are a processing apparatus, a storage management method, and a related device. The apparatus may comprise a processor core, a last-stage cache LLC, and a memory controller. The memory controller is used for: receiving a first memory access request sent by the processor core and used for reading first data, the first memory access request carrying a first memory access address; on the basis of the first memory access address, reading memory access data from a memory coupled to the processing apparatus; and if the memory access data comprises the first data and second data in a compressed state, in response to the first memory access request, sending the first data to the processor core by means of the LLC, and caching the second data in the compressed state to the LLC. The processing apparatus of the present application can reduce the overhead of storage resources on the basis of memory bandwidth compression technology.
Owner:HUAWEI TECH CO LTD

Data processing method and device, medium and program product

The invention relates to the technical field of artificial intelligence, and provides a data processing method and device, a medium and a program product.The method comprises the steps that a target calculation task is executed on input data, and a first precision result is generated; before the first precision result is written into a global memory, calculating local absolute value maximum values corresponding to all result blocks in the first precision result in parallel, and determining a global absolute value maximum value in the first precision result based on the local absolute value maximum values of all the result blocks; and based on the global absolute value maximum value, calculating to obtain a scaling factor, and performing quantization processing on the first precision result by using the scaling factor to obtain a second precision result. Before the first precision result is written into the global memory, the on-chip memory is used for calculating the maximum value of the local absolute value of the first precision result, and redundant memory read-write operation on the first precision result is effectively avoided, so that memory bandwidth occupation and processing delay are remarkably reduced, and the overall execution efficiency is improved.
Owner:SHANGHAI BIREN TECH CO LTD

Method for reasoning optimization of pre-training model and electronic equipment

The invention discloses a reasoning optimization method of a pre-training model and electronic equipment, and relates to the technical field of reasoning of the pre-training model.The priority of each subtask is determined on the basis of load data of multiple task stages of the pre-training model.The calculation units are allocated to the subtasks according to the priorities of the subtasks, and therefore the calculation efficiency of the subtasks is improved. By dynamically decoupling each task stage, efficient allocation of computing resources is realized, and the resource utilization rate is improved. And on the other hand, the first data of the two adjacent task stages are transmitted through a double-buffering mechanism, the problem that the inter-layer data transmission efficiency is low in the related technology is solved, and the hardware bandwidth utilization rate is increased. Therefore, the technical problems of low utilization rate of idle memory bandwidth resources and low interlayer transmission efficiency are solved, and the technical effect of improving the resource utilization rate and the bandwidth utilization rate is achieved.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Multi-platform dirty disk performance consistency testing method for solid state disk

The invention relates to the technical field of data storage, and discloses a multi-platform dirty disk performance consistency testing method for a solid state disk, which comprises the following steps: acquiring configuration information of a to-be-tested solid state disk and recording a SMART data baseline; the method comprises the following steps: initializing a solid state disk, performing steady-state preprocessing in an empty disk state of the initialized solid state disk, and recording performance data at the moment as an empty disk reference; filling the dirty disk proportion of the solid state disk to a target value in a sequential or random write-in mode, applying a corresponding dynamic load model under the dirty disk proportion, and continuously performing performance sampling, recording performance data and generating a performance attenuation curve; and starting a CPU pressure test, a memory bandwidth test and a GPU load test at the background while running the solid state disk performance benchmark test. According to the SSD performance compatibility testing method and system, cross testing is carried out on two mainstream platforms of AMD and Intel, and the performance compatibility of the SSD in different system environments can be systematically evaluated.
Owner:SHENZHEN JINGCUN TECH CO LTD

Inter-board communication method, network device and electronic device

The application provides an inter-board communication method, a network device and an electronic device. The inter-board communication in the application does not depend on a socket mechanism, but places message receiving and sending, message fragmentation and fragment recombination in a user state. In the user state, zero-copy can be used for inter-board communication, without the need for copy actions between a kernel state and the user state, thereby reducing CPU resource and memory bandwidth consumption caused by copy actions between the kernel state and the user state.
Owner:NEW H3C TECH CO LTD

Big language model-based reasoning method and device, electronic equipment and storage medium

The embodiment of the invention relates to the field of artificial intelligence, and discloses a reasoning method and device based on a large language model, electronic equipment and a storage medium. The method comprises the following steps: inputting a last token of a first token set into a prediction module, and outputting a second token set; carrying out parallel calculation on the second token set in a decoder layer, and outputting the next reasoning token of each token in the second token set; and inputting all the reasoning tokens into a prediction result decision device, for each prediction branch sequence in the second token set, matching the next reasoning token of the tokens in the sequence with the next token stage by stage from the first stage, and outputting the longest token sequence obtained by matching as a third token set by the prediction result decision device. Through a prediction-parallelization-judgment process, the problem of video memory bandwidth bottleneck caused by the fact that a large language model calculates tokens one by one is solved.
Owner:SHANGHAI JIANQI TECHNOLOGY CO LTD

On-demand regulation of memory bandwidth utilization to service requirements of display

Systems, apparatuses, and methods for prefetching data by a display controller. From time to time, a performance-state change of a memory are performed. During such changes, a memory clock frequency is changed for a memory subsystem storing frame buffer(s) used to drive pixels to a display device. During the performance-state change, memory accesses may be temporarily blocked. To sustain a desired quality of service for the display, a display controller is configured to prefetch data in advance of the performance-state change. In order to ensure the display controller has sufficient memory bandwidth to accomplish the prefetch, bandwidth reduction circuitry in clients of the system are configured to temporarily reduce memory bandwidth of corresponding clients.
Owner:ATI TECHNOLOGIES ULC +1

File decompilation method and electronic device

The application discloses a file unshelling method and electronic equipment, and relates to the technical field of computer security, and comprises the following steps: inserting a probe at a target system call of a kernel, and injecting the probe into a target program; identifying a suspiciously shelled process through the target program; monitoring a memory permission change behavior performed by the suspiciously shelled process through the target program, and counting a permission change frequency of a memory region, an entropy value of the memory region and a call chain depth; when it is confirmed that the suspiciously shelled process is in an unshelling stage according to the permission change frequency, the entropy value and the call chain depth, writing data decrypted by the suspiciously shelled process into a target buffer through the target program; reading the data from the target buffer in a user mode, and recombining the data into a memory image according to a base address and a dirty page bitmap and writing the memory image back to a file. The method can avoid I / O bottlenecks caused by full memory dumping, reduce invalid data in the user mode, and reduce CPU occupation and memory bandwidth.
Owner:LANGCHAO ELECTRONIC INFORMATION IND CO LTD

Video coding and video distribution

Motion compensation requires a significant amount of memory bandwidth, especially for smaller prediction unit sizes. The worst case bandwidth requirements can occur when bi-predicted 4×8 or 8×4 PUs are used. To reduce the memory bandwidth requirements for such smaller PUs, methods are provided for restricting inter-coded PUs of small block sizes to be coded only in a uni-predictive mode, i.e., forward prediction or backward prediction. More specifically, PUs of specified restricted sizes in bi-predicted slices (B slices) are forced to be uni-predicted.
Owner:TEXAS INSTRUMENTS INC

Memory access optimization method and system for neural radiation field rendering

The invention discloses a neural radiation field real-time rendering-oriented memory access optimization method and system, and mainly solves the problems of low rendering efficiency and difficulty in real-time rendering on edge equipment in the prior art. According to the scheme, the method comprises the steps that all zero value units of a trained model are removed, and index-effective value key value pairs of compressed grids are constructed and stored on a chip; the occupation grid query of all light advancing sampling points is completed on a chip through two-dimensional mapping; dividing a high-resolution spatial hash table in the multi-resolution hash code into sub-grids, and mapping spatially adjacent 3D points to a continuous memory address; all input light rays are partitioned, the light rays in each block query a hash table in parallel in a similar voxel region, and original random and scattered memory access is converted into a predictable access mode; and a two-stage Cache structure is adopted for caching and repeatedly using multi-resolution hash addresses and values. According to the method, the memory bandwidth can be reduced, the rendering speed can be increased, and the method can be applied to augmented reality, virtual reality and other scenes with high real-time requirements.
Owner:XIDIAN UNIV