Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

515 results about "Direct memory access" patented technology

Direct memory access (DMA) is a feature of computer systems that allows certain hardware subsystems to access main system memory (random-access memory), independent of the central processing unit (CPU).

Complex scene-oriented AI large model lightweight deployment method

The invention provides a complex scene-oriented AI large model lightweight deployment method, and relates to the technical field of edge computing, and the method comprises the steps: carrying out the structured pruning of a pre-trained Transform network based on the attention head importance score, carrying out the dynamic sparsification of the activation state of a feedforward network according to the input tensor entropy value, employing the dynamic mixing precision quantization, and carrying out the reconstruction of an AI large model. Obtaining network parameters after pruning quantization; deploying the pruned and quantized network parameters to an edge computing device, distributing a feature extraction operator to a neural network processor through a heterogeneous computing scheduler, and unloading a classification operator to a multi-core central processing unit; and managing an on-chip memory in combination with a virtual memory paging mechanism, realizing zero-copy data transmission by utilizing a direct memory access controller, and outputting a reasoning result tensor. According to the method, efficient and reliable operation of the large model at the resource-constrained edge node is realized.
Owner:XIAN XINGXUN INTELLIGENT COMM TECH CO LTD

Vehicle-mounted edge computing data recording system and method based on protocol adaptive analysis

The invention relates to the technical field of network communication, and discloses a vehicle-mounted edge computing data recording system and method based on protocol adaptive analysis, and the system comprises a multi-protocol network access controller which is used for capturing an original data frame and extracting a physical feature triple; the identification analysis engine is used for executing hash operation on the triple to generate a physical feature code and retrieving a physical logic address mapping table; the direct memory access controller is used for responding to a target memory address pointer hit by retrieval and directly writing a data load into an input buffer area of the functional operation module, and by constructing a direct addressing mechanism based on Hash mapping, thorough decoupling of vehicle-mounted heterogeneous network physical topology and edge computing logic is achieved.
Owner:SHANGHAI JUPO TECH CO LTD

Two-level context caching and eviction for scatter-gather DMA

One aspect of the instant disclosure may provide a system and method for processing scatter-gather direct memory access (S-G DMA) instructions. During operation, the system may receive an S-G DMA instruction associated with a message and gather instruction context for the S-G DMA instruction. An S-G DMA processor may process the S-G DMA instruction based on the gathered instruction context and determine whether there exists a pending S-G DMA instruction associated with the message. In response to the presence of the pending S-G DMA instruction, the system stores the instruction context in a hot context cache at an address corresponding to the pending S-G DMA instruction. In response to the absence of the pending S-G DMA instruction, the system stores the instruction context in a cold context cache.
Owner:HEWLETT PACKARD ENTERPRISE DEV LP

Method for quickly forwarding message from PCIE interface to WIA interface

The invention relates to the technical field of industrial communication networks, in particular to a method for quickly forwarding messages from a PCIE (Peripheral Component Interface Express) interface to a WIA (Wireless Interface Architecture) interface, which comprises the following steps of: simultaneously receiving PCIE messages and WIA-FA protocol messages through a hardware logic circuit; analyzing the message in a data link layer to obtain an analysis result comprising data load and address information; packaging the analysis result according to a pre-configuration rule to generate a data frame conforming to a target protocol; and constructing a sending descriptor for the generated data frame, and sending the data frame from a corresponding interface through a direct memory access mechanism. According to the method, parallel analysis and protocol conversion of messages are achieved through a hardware logic circuit, fast address conversion is achieved through a pre-configured address mapping table, and zero-copy data transmission is achieved through a direct memory access mechanism. According to the invention, software protocol stack processing is replaced by a hardware processing flow, so that communication delay and CPU resource consumption are effectively reduced.
Owner:BONCHREE (SHANGHAI) COMMUNICATION CO LTD

Overhead reduction using address translation in direct memory accesses

Techniques to reduce direct memory access (DMA) overhead may include retrieving an address translation descriptor from a descriptor queue of a DMA engine, and updating an address translation table in the DMA engine with address translation information obtained from the location indicated by the address translation descriptor. A set of memory descriptors is then obtained from the descriptor queue. The set of memory descriptors can be processed by determining that the addresses in the set of memory descriptors are to be translated using the address translation table, and performing memory access operations by using the address translation table to translate the addresses in the set of memory descriptors.
Owner:AMAZON TECH INC

Efficient regional ocean forecasting method

The invention provides an efficient regional ocean forecasting method, and belongs to the technical field of ocean forecasting. A regional ocean mode four-dimensional variational assimilation system is constructed, and an adjoint mode calculation framework is established; after calculation modules are grouped according to dependency dimensions, a slave core parallel scheme is designed, and data transmission and calculation assembly line overlapping are realized by adopting a direct memory access step access and double-buffering technology; introducing a multi-scale time step adaptive integral algorithm and a gradient convergence acceleration model, predicting an optimal convergence path according to a historical iteration trajectory, dynamically adjusting a search direction and a step factor, completing four-dimensional variation assimilation, outputting an optimized ocean initial field, and performing forward integral forecasting to generate ocean state field variable forecasting data; the technical problem of insufficient timeliness of regional ocean forecasting caused by low calculation efficiency of the adjoint mode is solved.
Owner:青岛国实科技集团有限公司

Naked eye 3D binocular image acquisition and real-time processing system based on hardware synchronous triggering

The invention relates to a naked eye 3D binocular image acquisition and real-time processing system based on hardware synchronous triggering, and belongs to the technical field of image processing and three-dimensional display. According to the system, a hardware synchronous triggering mechanism is adopted, a time sequence synchronous control unit is used for sending a synchronous pulse signal to a binocular image sensor, and strict synchronization of left and right viewpoint image acquisition is ensured. The image signal processing unit processes collected original data, the stereo parallax correction module performs epipolar correction, and the sub-pixel interleaving module generates a composite view frame according to grating parameters of the display terminal. The system realizes high-speed data transmission through a double-buffer direct memory access DMA mechanism. The problems of visual tearing and weak stereoscopic impression caused by asynchronous binocular image acquisition in the prior art are solved, nanosecond-level synchronization precision is realized, optical crosstalk is reduced, and smoothness and comfort of stereoscopic display are ensured.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Data scheduling and control method based on mainboard memory direct connection

The invention discloses a data scheduling and control method based on mainboard memory straight-through, and relates to the technical field of computer hardware, the method comprises the following steps: pre-allocating a straight-through memory area for mounting equipment in a firmware initialization stage, establishing a mapping relation between equipment identification and a physical address of an exclusive memory segment, and through address translation and access control rules of a firmware layer, the device bypasses a system cache level and accesses the exclusive memory segment in a direct memory access mode. Constructing a firmware scheduling table based on the mapping relation to manage access requests of all devices, monitoring transmission delay in real time, and dynamically adjusting priority weight and bandwidth allocation in the scheduling table according to a monitoring result; and sequencing the requests according to the updated scheduling table to complete scheduling and control of the memory access behavior of the equipment. According to the method and the device, the delay bottleneck caused by multi-level copying and firmware scheduling lag in mainboard data forwarding is effectively solved.
Owner:HUAIAN COLLEGE OF INFORMATION TECH +1

High-real-time interrupt management system and method based on RISC-V architecture

The invention relates to the technical field of integrated circuit design, computer system structures and embedded systems, in particular to a high-real-time interrupt management system and method based on an RISC-V. The system comprises an improved platform-level interrupt controller and an optimized processor internal interrupt processing unit. The system and the method are in tight coupling cooperative work through a special interruption field interface and a system bus, and the system further comprises a core local interrupter, an SRAM, an APB bus matrix and other components. The improved platform-level interrupt controller is responsible for sampling, gating, arbitration, flow control and direct memory access transmission of external interrupt, and a plurality of functional modules are arranged in the improved platform-level interrupt controller; an interrupt processing unit in the processor is responsible for hardware handshake, vector jump, nested control and tail biting mechanism implementation. Through hardware field management, multi-level priority arbitration, hardware nesting and tail biting mechanisms, delay and overhead caused by software intervention in a traditional architecture are eliminated, nanosecond response is achieved, the system throughput rate is increased, the software development threshold is lowered, and the method is suitable for strong real-time scenes.
Owner:FUDAN UNIVERSITY

Virtual machine memory processing method and device, product, virtualization server and medium

The invention discloses a virtual machine memory processing method and device, a product, a virtualization server and a medium, and relates to the field of server virtualization. In the method, through a virtual input and output rear-end module arranged on a virtual machine monitor in a host and a virtual input and output front-end module on a virtual machine, it is ensured that the virtual input and output rear-end module can obtain information of a memory allocation event of direct memory access; secondly, after acquiring the information of the memory allocation event of the direct memory access, the virtual input / output rear-end module locks the address related to the memory allocation event, namely, the locked address meets the direct memory access requirement, and the memory of the virtual machine is dynamically locked, so that the memory utilization rate is improved, and the waste of resources is reduced; and only part of the memory of the virtual machine is locked, and only the locked memory part is migrated after the host is replaced, so that the amount of data to be migrated after the host is replaced is reduced.
Owner:JINAN INSPUR DATA TECH CO LTD

Data processing method and electronic device using scatter gather DMA

A data processing method using a scatter gather direct memory access (SG DMA), the method comprising: obtaining information for a rule table for specific subtasks of a SG DMA from a host, deriving the rule table for the specific subtasks based on the information, deriving source addresses, destination addresses and data sizes for the specific subtasks based on the rule table, and performing a SG DMA operation to transfer data of the data sizes located at the source addresses of a first memory to data spaces of the data sizes located at the destination addresses of a second memory.
Owner:REBELLIONS INC

Key-value cache compression method and system for accelerating large language model inference

Provided is an accelerator for accelerating batch large language model (LLM) inference via key-value cache compression. An accelerator, according to one embodiment, may comprise a plurality of compute cores. Here, each of the plurality of compute cores may comprise a plurality of processing units for processing an LLM inference operation on a per-token basis, and a direct memory access unit for managing operations of reading weights from a memory and writing key-value activation data back to the memory. In addition, the direct memory access unit may comprise a compression engine for managing online compression of the key-value activation data when the key-value activation data is written to the memory, a decompression engine for decompressing the compressed key-value activation data retrieved from the memory, and a memory management unit for managing reading and writing of the compressed key-value activation data in the memory.
Owner:HYPERACCEL CO LTD

Method and apparatus for hardware resource sharing in direct memory access controller

A direct memory access controller (DMAC) includes a virtual channel, a physical channel, and a context manager. The virtual channel is configured to generate a flow control signal for executing a direct memory access (DMA) command. The physical channel includes read and write control logic to initiate data transmission from the source device to the destination device in accordance with the DMA command in response to the flow control signal. In one embodiment, a context manager includes: allocation logic configured to arbitrate between allocation requests from virtual channels and to allocate physical channels to selected virtual channels; and routing logic configured to route a flow control signal from the selected virtual channel to the allocated physical channel, and to route a status signal between the allocated physical channel and the selected virtual channel.
Owner:ARM LTD

Multi-channel ultrasonic transducer phased array driving method based on DMA

The invention discloses a multichannel ultrasonic transducer phased array driving method based on direct memory access. The method comprises a waveform data buffer length determination step, a data path construction step, a phase parameter calculation step, a waveform data synthesis step and a DMA driving step. An independent DMA channel is configured for each GPIO port or each group of GPIO ports corresponding to the transducer array by utilizing the DMA characteristic of a universal microcontroller, and buffer data is transmitted to the output data registers of the GPIO ports automatically and periodically by DMA hardware, so that multi-channel driving signals with accurate frequency and controllable phases are generated in parallel under the condition that CPU (Central Processing Unit) intervention is not needed. According to the invention, high-precision time sequence control comparable with an FPGA (Field Programmable Gate Array) scheme is realized with extremely low hardware cost and CPU (Central Processing Unit) resource occupation, the problem of time sequence jitter of an existing MCU (Microprogrammed Control Unit) scheme is solved, and the method has high stability, high efficiency and excellent expandability.
Owner:SOUTHEAST UNIV

Data stream processing method and device, electronic equipment and storage medium

The invention discloses a data stream processing method and device, electronic equipment and a storage medium, and relates to the technical field of distributed storage. According to the data stream processing method and device, the data stream transmission mode is pre-judged to be intra-node transmission or cross-node transmission based on a hardware topology mapping table, and read-write caches are dynamically allocated according to the pre-judged data stream transmission mode; data received by the front-end network card is directly written into a controller memory buffer area of the target equipment through point-to-point direct memory access for data streams transmitted in the nodes, and is read and forwarded from the target equipment by the rear-end network card; the sending cache and the receiving cache of the data stream transmitted across the nodes are locked to the non-uniform memory access node to which the target network card belongs, and the data are sent to the target node through the remote direct memory access register memory, so that the problem that the data stream processing is not optimized in combination with hardware topology characteristics can be solved; the technical effects of reducing data replication overhead, avoiding redundant access across hardware nodes, improving data stream transmission efficiency and realizing end-to-end zero replication transmission are achieved.
Owner:JINAN INSPUR DATA TECH CO LTD

Offloading of adaptive all reduce operations

Examples described herein relate to a network interface device that includes: a host interface; a direct memory access (DMA) circuitry; a network interface to receive, in at least one packet, time data associated with at least one of multiple layers, wherein the multiple layers provide inputs to a collective operation associated with a large language model (LLM); and circuitry. The circuitry is to based, at least in part, on the time data associated with the multiple layers, identify a first operation of a first layer of the multiple layers as a late completing process relative to times to completion of multiple first operations of other layers and based on the first operation being identified as a late completing process, perform a remedial action to adjust at least one configuration of a first device to execute a second operation of the first layer.
Owner:INTEL CORP

Dedicated direct memory access router system and method

A direct memory access (DMA) router including interrupt inputs, action groups and a DMA router engine. Each interrupt input is configured to receive a corresponding interrupt signal. Each action group is associated with a corresponding interrupt input and is configured with at least one DMA action, in which each DMA action is configured to select a DMA controller and a corresponding channel. The DMA router is configured to initiate a transfer using a selected channel of a selected DMA controller for at least one DMA action listed in an action group associated with a corresponding interrupt input triggered by assertion of a corresponding interrupt signal. The DMA actions may indicate dependencies, such that the DMA router may initiate a second DMA action only after completion of a first DMA action within the same action group based upon the indicated dependency.
Owner:NXP USA INC

Data copying methods, apparatus, computer-readable storage media and electronic devices

A data copying method, apparatus, computer-readable storage medium, and electronic device are disclosed. The method includes: generating an address translation request via a target virtual machine, the address translation request including an intermediate physical address; converting the intermediate physical address into a physical address via a memory management unit; configuring the physical address in a direct memory access controller via the target virtual machine; and controlling a target module to copy data according to the physical address via the direct memory access controller, the target module including a memory module and / or a peripheral module. Embodiments of this disclosure can reduce chip manufacturing costs.
Owner:HORIZON JOURNEY (HANGZHOU) ARTIFICIAL INTELLIGENCE TECHNOLOGY CO LTD

Shared Memory Controller with Direct Memory Access Architecture for On-Chip Memory

The present disclosure describes System on Chip (SoC) architecture that facilitates disaggregation of memory-to-memory operations. The SoC architecture includes a host interface that communicates with a host system, processor cores, and an Advanced extensible Interface (AXI) interconnect coupling the host interface with processor cores. The SoC architecture includes an on-chip memory (OCM) subsystem coupled to the AXI interconnect, where the OCM subsystem contains memory banks, a Direct Memory Access (DMA) interconnect coupled directly with respective memories of processor cores, and a shared memory controller coupled with the AXI interconnect, memory banks, and DMA interconnect. The shared memory controller includes an OCM-internal path connecting the shared memory controller directly to memory banks within the OCM subsystem and a DMA engine that executes memory-to-memory operations by transferring data directly between memory banks through the OCM-internal path or respective memories of processor cores via a DMA interconnect.
Owner:MARVELL ASIA PTE LTD

Unified instruction processor for direct memory access scatter / gather engine

A system receives, by a network interface card (NIC), inputs including an instruction to read or write a payload of a message, a tracker state indicating a round of processing, and a datatype descriptor defining organization of the message payload. The system identifies a current context and a processing state for the instruction. If the datatype descriptor indicates a first type, the system: obtains the current context associated with the first type from a host memory or a cache of the NIC; and creates direct memory access (DMA) instructions corresponding to the received instruction by executing operations in a nested loop. If the datatype descriptor indicates a second type, the system: obtains the current context associated with the second type by fetching vector entries from a buffer of the NIC; and creates the DMA instructions corresponding to the received instruction based on addresses and lengths in the vector entries.
Owner:HEWLETT PACKARD ENTERPRISE DEV LP

Big data storage optimization method based on zero-copy collaboration technology and related equipment

The embodiment of the invention provides a big data storage optimization method based on a zero-copy collaboration technology and related equipment, and belongs to the technical field of big data storage. The method comprises a zero-copy data transmission module used for directly writing a data stream into a front-end buffer area of an annular structure through a direct memory access technology, and adopting a zero-copy algorithm based on pointer offset to parallelly separate data of each channel in a multi-thread environment; the intelligent data reduction module is used for deleting duplicated data from the data and then compressing the duplicated data; and the hybrid storage management module is used for managing the hybrid partition storage architecture and dynamically adjusting the distribution of the data in the hybrid partition storage architecture according to the data access mode. The CPU copy frequency is reduced through the zero copy technology, the data transmission efficiency is remarkably improved by combining the data reduction and intelligent layering strategies, the occupied storage space is reduced, and meanwhile the data safety and compliance are guaranteed.
Owner:GUANGDONG WANZHANG JINSHU INFORMATION TECH CO LTD

Hardware implementation method and device for communication between CPUs with low time delay

The invention discloses a low-delay hardware implementation method and device for communication between CPUs, and relates to the technical field of computers. The method comprises the following steps: receiving initialization configuration information from a sending CPU (Central Processing Unit), including an initial address and a space size of a memory of the sending CPU and an initial address and a space size of a memory of a receiving CPU; and monitoring a sending tail pointer updated by the sending CPU, and determining whether to read the to-be-sent data from the internal memory of the sending CPU to the internal cache of the hardware communication module based on the initialization configuration information, the sending tail pointer and a sending head pointer maintained by the hardware communication module. And monitoring a receiving head pointer updated by the receiving CPU, and determining whether to write the to-be-written data in the internal cache into the memory of the receiving CPU based on the initialization configuration information, the receiving head pointer and a receiving tail pointer maintained by the hardware communication module. And the hardware communication module transmits data between the sending CPU memory and the receiving CPU memory in a direct memory access mode.
Owner:PENG TI STORAGE TECH (NANJING) CO LTD

System and method for primary storage write traffic management

The present invention relates to a system (101) and method for primary storage write traffic management, which can improve the overall system on chip (SoC) data traffic efficiency between the processor (103), direct memory access (DMA) channel (111) and the main memory (113), by minimizing the latency to write the data to the main memory (113). This is done by reducing the number of writes from the processor (103) to the main memory (113) without sacrificing data consistency between the processor (103) and the DMA channel (111).
Owner:EFINIX INC

System on chip supporting data communication, data communication method and equipment

The invention discloses a system-on-chip supporting data communication and a data communication method and equipment, and relates to the technical field of controller local area networks. The system comprises a data buffer, a direct memory access controller and a processor; the data buffer is used for generating a starting signal in response to the received first data frame; the direct memory access controller is used for reading a current descriptor in a plurality of descriptors from a descriptor memory area of the memory in response to the starting signal; reading the first data frame from the data buffer area according to the current descriptor, and writing the first data frame into the data buffer area of the memory; in response to the fact that the first interrupt is in the enabled state, generating a first interrupt signal after the first data frame is written into the data buffer area; and the processor is used for reading the first data frame from the data buffer area in response to the first interrupt signal. The system can solve the overflow problem of the data buffer in a high-load scene, reduces the risk of data frame loss, and improves the real-time performance and reliability of the system.
Owner:BEIJING HORIZON INFORMATION TECH CO LTD

Method for judging stock leakage of water supply pipe network based on DMA (direct memory access) night water volume

The invention discloses a method for judging stock leakage of a water supply pipe network based on DMA (Direct Memory Access) night water volume, which is characterized in that the night water volume is relatively stable, and the night flow is finely corrected by creatively integrating lower-level subareas, night water consumption of large users and temperature data, so that the estimated leakage volume closer to the real situation is calculated. The time sequence distribution of the leakage amount ratio is further analyzed, credibility evaluation is introduced, interference caused by instantaneous fluctuation can be effectively filtered out, and finally early warning is triggered when the stock leakage credibility is high enough. According to the method, the accuracy of leakage judgment and the reliability of early warning are remarkably improved, a water affair department can be helped to actively and accurately recognize and manage stock leakage, and powerful support is provided for scientific decision making.
Owner:NINGBO DONGHAI GRP CORP +1

Method and system for in-line data conversion outside of a machine learning hardware

A system includes a component configured to send data in a first data format. The system includes a direct memory access (DMA) engine configured to receive the data in the first data format and convert the first data format to a second data format, wherein the second data format is associated with a data format of a machine learning (ML) hardware, wherein the second data format is different from the first data format. The ML hardware is configured to receive the data in the second format and perform at least one ML operation on the received data in the second format. The received data in the second data format is stored on an on-chip memory (OCM) of the ML hardware.
Owner:MARVELL ASIA PTE LTD

Machine learning acceleration architecture

A machine learning accelerator includes a scalable processor with a plurality of cores that receive data from system memory via a system direct memory access (DMA) engine. Each core may include local memory, a compute sub-system, and one or more slices, each of which includes a descriptor execution engine and one or more compute engines. Each compute engine includes input data memory, one or more sub-compute engines, and partial data memory. The sub-compute engines are separately connected to the input data memory and are configured to independently perform compute operations, such as multiply-accumulate (MAC) operations, on the input data and to provide partial output data to the partial data memory. The cores, slices and sub-compute engines may be configured to operate independently to perform separate tasks in parallel that once completed are combined as part of a large artificial intelligence model.
Owner:SYNAPTICS INC

Water supply network DMA intelligent partitioning method based on GIS

The invention discloses a GIS (Geographic Information System)-based DMA (Direct Memory Access) intelligent partitioning method for a water supply network, relates to the field of data processing, and can solve the problem that a virtual partition boundary is misjudged or a boundary at a certain position frequently fluctuates due to the fact that a DMA intelligent partitioning method for the water supply network at the present stage cannot accurately distinguish sudden water taking abnormity and leakage abnormity. The method comprises the following steps: acquiring water flow state monitoring data; determining a water flow abnormal data set according to the water flow state monitoring data; according to the water flow abnormal data set, a water supply pipe network partition preliminary result is determined; wherein the water supply pipe network zoning preliminary result is used for representing a preliminary division range of each area in the water supply pipe network; and determining the DMA partition of the water supply network according to the preliminary partition result of the water supply network and the physical equipment information of the water supply network. The method is used for DMA intelligent partitioning and leakage monitoring of the water supply network.
Owner:HUNAN LINXI CONSTRUCTION ENGINEERING CO LTD

Autonomous gradient reduction in a reconfigurable processor system

ActiveUS12450167B1Interprogram communicationVirtual memory detailsDirect memory accessNetwork data
A coarse-grained reconfigurable processor (CGRP) system for implementing data-parallel training of a neural network is presented. The CGRP system includes a set of coarse-grained reconfigurable units (CGRUs) in a first CGRP configured to implement at least a portion of the neural network, to determine first and second gradients, respectively, of first and second model parameters based on a batch of training data, and to store the first and second gradients in a memory, a network interface including an external direct memory access (DMA) engine coupled between the memory and a network, and a work queue associated with the external DMA engine, wherein completion of determining the first gradient triggers a first work queue entry of the work queue that directs the external DMA engine to transfer the first gradient from the memory over the network to another memory coupled to a second CGRP for a gradient reduction operation.
Owner:SAMBANOVA SYSTEMS INC