Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

709 results about "Direct memory access" patented technology

Direct memory access (DMA) is a feature of computer systems that allows certain hardware subsystems to access main system memory (random-access memory), independent of the central processing unit (CPU).

Complex scene-oriented AI large model lightweight deployment method

The invention provides a complex scene-oriented AI large model lightweight deployment method, and relates to the technical field of edge computing, and the method comprises the steps: carrying out the structured pruning of a pre-trained Transform network based on the attention head importance score, carrying out the dynamic sparsification of the activation state of a feedforward network according to the input tensor entropy value, employing the dynamic mixing precision quantization, and carrying out the reconstruction of an AI large model. Obtaining network parameters after pruning quantization; deploying the pruned and quantized network parameters to an edge computing device, distributing a feature extraction operator to a neural network processor through a heterogeneous computing scheduler, and unloading a classification operator to a multi-core central processing unit; and managing an on-chip memory in combination with a virtual memory paging mechanism, realizing zero-copy data transmission by utilizing a direct memory access controller, and outputting a reasoning result tensor. According to the method, efficient and reliable operation of the large model at the resource-constrained edge node is realized.
Owner:XIAN XINGXUN INTELLIGENT COMM TECH CO LTD

Vehicle-mounted edge computing data recording system and method based on protocol adaptive analysis

The invention relates to the technical field of network communication, and discloses a vehicle-mounted edge computing data recording system and method based on protocol adaptive analysis, and the system comprises a multi-protocol network access controller which is used for capturing an original data frame and extracting a physical feature triple; the identification analysis engine is used for executing hash operation on the triple to generate a physical feature code and retrieving a physical logic address mapping table; the direct memory access controller is used for responding to a target memory address pointer hit by retrieval and directly writing a data load into an input buffer area of the functional operation module, and by constructing a direct addressing mechanism based on Hash mapping, thorough decoupling of vehicle-mounted heterogeneous network physical topology and edge computing logic is achieved.
Owner:SHANGHAI JUPO TECH CO LTD

Two-level context caching and eviction for scatter-gather DMA

One aspect of the instant disclosure may provide a system and method for processing scatter-gather direct memory access (S-G DMA) instructions. During operation, the system may receive an S-G DMA instruction associated with a message and gather instruction context for the S-G DMA instruction. An S-G DMA processor may process the S-G DMA instruction based on the gathered instruction context and determine whether there exists a pending S-G DMA instruction associated with the message. In response to the presence of the pending S-G DMA instruction, the system stores the instruction context in a hot context cache at an address corresponding to the pending S-G DMA instruction. In response to the absence of the pending S-G DMA instruction, the system stores the instruction context in a cold context cache.
Owner:HEWLETT PACKARD ENTERPRISE DEV LP

Direct memory access (DMA) engine with network interface capabilities

Examples described herein include one or more processors; a network interface; and a direct memory access (DMA) engine communicatively coupled to the one or more processors. In some examples, the DMA engine is to receive a DMA data access request and based on an address in the DMA data access request corresponding to a remote memory device, the DMA engine is to cause the network interface to generate at least one packet for transmission to the remote memory device. In some examples, if the source address corresponds to a local memory device and the destination address corresponds to a remote memory device, the DMA engine is to cause the network interface to generate at least one packet for transmission to the remote memory device.
Owner:INTEL CORP

Multi-thread high-throughput data flow channel separation method and system based on zero copy

The invention belongs to the technical field of data transmission and processing, and discloses a zero-copy-based multi-thread high-throughput data stream channel separation method and system, and the method comprises the steps: directly writing a mixed data stream into a front-end buffer region configured as an annular structure through a data receiving module by adopting direct memory access; then, a multi-thread processing module dynamically allocates a plurality of processing threads from a thread pool to separate channel data in parallel, each thread adopts a zero copy algorithm based on pointer offset, positions the channel data in a memory, creates pointer reference and associates the channel data to a corresponding rear-end buffer area, and logic separation is achieved without physical copy; and finally, the data storage module efficiently writes the separated data into persistent storage in an asynchronous I / O mode. According to the method, zero-copy, multi-thread parallel and two-stage dynamic buffering strategies are combined, the data separation efficiency is remarkably improved, CPU occupation and memory bandwidth are greatly reduced, and the real-time performance and stability of high-throughput data processing are guaranteed.
Owner:CHINA JILIANG UNIV

Data transmission method, device, equipment and medium

The invention discloses a data transmission method and device, equipment and a medium, and relates to the technical field of computers. By adopting a redundancy architecture of multiple direct memory access channels and combining a channel transmission mode based on a trigger condition and a target screening strategy, the data transmission efficiency is improved. The appropriate target direct memory access channel is selected from the multiple direct memory access channels to execute the data transmission task, so that when a certain direct memory access channel has an error, the system can be automatically switched to other direct memory access channels to continue to execute the data transmission task; overall interruption of the non-transparent bridge link caused by a single direct memory access channel fault is avoided, meanwhile, for the direct memory access fault, the fault channel can be accurately repaired by adopting a fault-level-based hierarchical repair strategy, the whole non-transparent bridge link and all devices do not need to be reinitialized, and the fault channel repair efficiency is improved. The fault processing time is greatly shortened, the fault recovery efficiency is effectively improved, and the continuity of cross-storage node data transmission is guaranteed.
Owner:LANGCHAO ELECTRONIC INFORMATION IND CO LTD

Systems, methods, and media for unordered input / output direct memory access operations

Mechanisms for unordered input / output direct memory access operations are provided, including: issuing using a hardware processor a back invalidate snoop request to a cache coherency control unit of a host processor; and issuing an unordered input / output direct memory access operation request to a Compute Express Link memory device. In some of these mechanisms, the unordered input / output direct memory access operation request is for a read operation. In some of these mechanisms, the mechanisms further comprise receiving a response to the unordered input / output direct memory access operation request including data from the Compute Express Link memory device. In some of these mechanisms, the data was updated in response to the back invalidate snoop request. In some of these mechanisms, the unordered input / output direct memory access operation request is for a write operation.
Owner:SK HYNIX NAND PRODUCT SOLUTIONS CORP

Method for quickly forwarding message from PCIE interface to WIA interface

The invention relates to the technical field of industrial communication networks, in particular to a method for quickly forwarding messages from a PCIE (Peripheral Component Interface Express) interface to a WIA (Wireless Interface Architecture) interface, which comprises the following steps of: simultaneously receiving PCIE messages and WIA-FA protocol messages through a hardware logic circuit; analyzing the message in a data link layer to obtain an analysis result comprising data load and address information; packaging the analysis result according to a pre-configuration rule to generate a data frame conforming to a target protocol; and constructing a sending descriptor for the generated data frame, and sending the data frame from a corresponding interface through a direct memory access mechanism. According to the method, parallel analysis and protocol conversion of messages are achieved through a hardware logic circuit, fast address conversion is achieved through a pre-configured address mapping table, and zero-copy data transmission is achieved through a direct memory access mechanism. According to the invention, software protocol stack processing is replaced by a hardware processing flow, so that communication delay and CPU resource consumption are effectively reduced.
Owner:BONCHREE (SHANGHAI) COMMUNICATION CO LTD

Data transmission method and device and storage medium

The embodiment of the invention provides a data transmission method and device and a storage medium. In the embodiment of the invention, a new direct memory access interface is provided, in implementation, an agent module is provided for network card equipment, and a task management queue and a memory management queue are provided between the agent module and an application as a direct memory access channel between the agent module and the application; in the data transmission process of the two ends, on one hand, the two ends can adopt a bypass kernel and zero copy technology to improve the data transmission efficiency based on the respective task management queues, and on the other hand, the two ends actively implement memory management for data storage based on the respective memory management queues, and interaction of memory management with the opposite end is not needed; the interactive operation of the two ends in the data transmission process can be reduced, and the transmission delay can be further reduced on the basis of achieving direct memory access.
Owner:HANGZHOU ALICLOUD FEITIAN INFORMATION TECH CO LTD

Overhead reduction using address translation in direct memory accesses

Techniques to reduce direct memory access (DMA) overhead may include retrieving an address translation descriptor from a descriptor queue of a DMA engine, and updating an address translation table in the DMA engine with address translation information obtained from the location indicated by the address translation descriptor. A set of memory descriptors is then obtained from the descriptor queue. The set of memory descriptors can be processed by determining that the addresses in the set of memory descriptors are to be translated using the address translation table, and performing memory access operations by using the address translation table to translate the addresses in the set of memory descriptors.
Owner:AMAZON TECH INC

Electromagnetic scattering-oriented sparse approximate inverse Shenwei parallel preprocessing method and system

The invention provides a sparse approximate inverse Shenwei parallel preprocessing method and system for electromagnetic scattering, and relates to the technical field of processor parallel computing, and the method comprises the steps: carrying out the performance hotspot analysis of a sparse approximate inverse preprocessor in an electromagnetic scattering simulation process, and determining a hotspot function of the sparse approximate inverse preprocessor; extracting a sparse matrix vector multiplication operation of a hotspot function, performing master-slave core parallel calculation on the sparse matrix vector multiplication operation, migrating the sparse matrix vector multiplication operation to slave core calculation, and accelerating hotspot calculation by using a heterogeneous parallel sparse approximate inverse algorithm; parallel computing of master-slave core intensive numerical computing tasks is achieved; in the parallel computing process, the slave cores communicate with the master core in a direct memory access mode, and during the period, the master core is in a waiting state to ensure data consistency and communication synchronization until all the slave cores compute distributed tasks.
Owner:QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1

Video acquisition and transmission method and device for converting 10-gigabit network to PCIe interface

The invention relates to a video acquisition and transmission method and device for converting a 10-gigabit network into a PCIe (Peripheral Component Interconnect Express) interface. The method comprises the following steps: receiving a video stream through a double 10-gigabit network optical module, carrying out protocol analysis and preprocessing through an FPGA (Field Programmable Gate Array), realizing multi-path alternate reading and writing by utilizing DDR3 ping-pong cache, dynamically distributing the video stream to multiple devices through a three-path PCIe 3.0 * 4 interface and an XDMA (Extensible Direct Memory Access), and optimizing the transmission priority in combination with a load state. The device comprises two paths of SFP + optical modules, a Xilinx Virtex7 series FPGA chip (integrated with a 10-gigabit network MAC, a PCIe XDMA IP core and a DDR3 controller), three paths of PCIe 3.0 * 4 interfaces and an external DDR3 memory. The system has the advantages that the double 10-gigabit network input and multi-PCIe channel collaborative architecture breaks through the single-channel bandwidth bottleneck; the intelligent ping-pong cache and the AXI arbitration strategy reduce the data conflict risk; and dynamic bandwidth allocation optimizes real-time performance.
Owner:58TH RES INST OF CETC

Efficient regional ocean forecasting method

The invention provides an efficient regional ocean forecasting method, and belongs to the technical field of ocean forecasting. A regional ocean mode four-dimensional variational assimilation system is constructed, and an adjoint mode calculation framework is established; after calculation modules are grouped according to dependency dimensions, a slave core parallel scheme is designed, and data transmission and calculation assembly line overlapping are realized by adopting a direct memory access step access and double-buffering technology; introducing a multi-scale time step adaptive integral algorithm and a gradient convergence acceleration model, predicting an optimal convergence path according to a historical iteration trajectory, dynamically adjusting a search direction and a step factor, completing four-dimensional variation assimilation, outputting an optimized ocean initial field, and performing forward integral forecasting to generate ocean state field variable forecasting data; the technical problem of insufficient timeliness of regional ocean forecasting caused by low calculation efficiency of the adjoint mode is solved.
Owner:青岛国实科技集团有限公司

Naked eye 3D binocular image acquisition and real-time processing system based on hardware synchronous triggering

The invention relates to a naked eye 3D binocular image acquisition and real-time processing system based on hardware synchronous triggering, and belongs to the technical field of image processing and three-dimensional display. According to the system, a hardware synchronous triggering mechanism is adopted, a time sequence synchronous control unit is used for sending a synchronous pulse signal to a binocular image sensor, and strict synchronization of left and right viewpoint image acquisition is ensured. The image signal processing unit processes collected original data, the stereo parallax correction module performs epipolar correction, and the sub-pixel interleaving module generates a composite view frame according to grating parameters of the display terminal. The system realizes high-speed data transmission through a double-buffer direct memory access DMA mechanism. The problems of visual tearing and weak stereoscopic impression caused by asynchronous binocular image acquisition in the prior art are solved, nanosecond-level synchronization precision is realized, optical crosstalk is reduced, and smoothness and comfort of stereoscopic display are ensured.
Owner:CHONGQING UNIV OF POSTS & TELECOMM

Data scheduling and control method based on mainboard memory direct connection

The invention discloses a data scheduling and control method based on mainboard memory straight-through, and relates to the technical field of computer hardware, the method comprises the following steps: pre-allocating a straight-through memory area for mounting equipment in a firmware initialization stage, establishing a mapping relation between equipment identification and a physical address of an exclusive memory segment, and through address translation and access control rules of a firmware layer, the device bypasses a system cache level and accesses the exclusive memory segment in a direct memory access mode. Constructing a firmware scheduling table based on the mapping relation to manage access requests of all devices, monitoring transmission delay in real time, and dynamically adjusting priority weight and bandwidth allocation in the scheduling table according to a monitoring result; and sequencing the requests according to the updated scheduling table to complete scheduling and control of the memory access behavior of the equipment. According to the method and the device, the delay bottleneck caused by multi-level copying and firmware scheduling lag in mainboard data forwarding is effectively solved.
Owner:HUAIAN COLLEGE OF INFORMATION TECH +1

Adaptive system probe action to minimize input / output dirty data transfers

Adaptive system probe action to minimize input / output dirty data transfers is described. In one or more implementations, a system includes a processor, a memory configured to store data, and a cache configured to store a portion of the data stored in the memory for execution by the processor. The system also includes a cache coherence controller including a cache line history. The cache coherence controller is configured detect a direct memory access request from an input / output device. The direct memory access request is associated with an input / output operation involving the data. The cache coherence controller is further configured to identify a cache line associated with the direct memory access request, and, in response to the cache line history including a dirty data transfer record corresponding to the cache line, selectively send a probe to the cache based on a state of the cache line.
Owner:ADVANCED MICRO DEVICES INC

Mapping abstract data movements into sequential and parallel direct memory access (DMA) programming

In various examples, systems and methods are disclosed relating to a system including one or more processors to generate hardware-level configurations for direct memory access (DMA) devices based on high-level descriptions of data movements. The high-level descriptions may include data flows for transferring data using the DMA device and the system may automatically generate the hardware-level configurations for the DMA device based on the data flows, simplifying the process of programming data movements and reducing the opportunity for human error.
Owner:NVIDIA CORP

Log writing control method and device, equipment and storage medium

The embodiment of the invention discloses a log write-in control method and device, equipment and a storage medium, applied to electronic equipment, the electronic equipment comprises an embedded multimedia card eMMC, the method comprises the steps that target parameters are acquired, the target parameters comprise health parameters and / or system load data of the eMMC, and the health parameters of the eMMC and / or the system load data of the eMMC are acquired; the health parameters comprise at least one of erasing times, bad block rate and current temperature, and the system load data comprise at least one of central processing unit (CPU) utilization rate and direct memory access (DMA) queue depth; calculating a target storage pressure index SPI based on the target parameter; and according to a mapping relation between a preset SPI range and a preset log level, a target log corresponding to the target log level is written into the eMMC, and the target log level is the log level corresponding to the target SPI. Loss of storage space can be reduced, and the service life of a memory device is prolonged.
Owner:东莞市步步高教育软件有限公司

High-real-time interrupt management system and method based on RISC-V architecture

The invention relates to the technical field of integrated circuit design, computer system structures and embedded systems, in particular to a high-real-time interrupt management system and method based on an RISC-V. The system comprises an improved platform-level interrupt controller and an optimized processor internal interrupt processing unit. The system and the method are in tight coupling cooperative work through a special interruption field interface and a system bus, and the system further comprises a core local interrupter, an SRAM, an APB bus matrix and other components. The improved platform-level interrupt controller is responsible for sampling, gating, arbitration, flow control and direct memory access transmission of external interrupt, and a plurality of functional modules are arranged in the improved platform-level interrupt controller; an interrupt processing unit in the processor is responsible for hardware handshake, vector jump, nested control and tail biting mechanism implementation. Through hardware field management, multi-level priority arbitration, hardware nesting and tail biting mechanisms, delay and overhead caused by software intervention in a traditional architecture are eliminated, nanosecond response is achieved, the system throughput rate is increased, the software development threshold is lowered, and the method is suitable for strong real-time scenes.
Owner:FUDAN UNIVERSITY

Virtual machine memory processing method and device, product, virtualization server and medium

The invention discloses a virtual machine memory processing method and device, a product, a virtualization server and a medium, and relates to the field of server virtualization. In the method, through a virtual input and output rear-end module arranged on a virtual machine monitor in a host and a virtual input and output front-end module on a virtual machine, it is ensured that the virtual input and output rear-end module can obtain information of a memory allocation event of direct memory access; secondly, after acquiring the information of the memory allocation event of the direct memory access, the virtual input / output rear-end module locks the address related to the memory allocation event, namely, the locked address meets the direct memory access requirement, and the memory of the virtual machine is dynamically locked, so that the memory utilization rate is improved, and the waste of resources is reduced; and only part of the memory of the virtual machine is locked, and only the locked memory part is migrated after the host is replaced, so that the amount of data to be migrated after the host is replaced is reduced.
Owner:JINAN INSPUR DATA TECH CO LTD

Data processing method and electronic device using scatter gather DMA

A data processing method using a scatter gather direct memory access (SG DMA), the method comprising: obtaining information for a rule table for specific subtasks of a SG DMA from a host, deriving the rule table for the specific subtasks based on the information, deriving source addresses, destination addresses and data sizes for the specific subtasks based on the rule table, and performing a SG DMA operation to transfer data of the data sizes located at the source addresses of a first memory to data spaces of the data sizes located at the destination addresses of a second memory.
Owner:REBELLIONS INC

Systems and Methods for Cold Boot Using Digital Twin

In one embodiment, a method may receive a request to upgrade software for a switch. The method may use an application-specific integrated circuit (ASIC) simulator to generate a digital twin to store a first image and a first configuration of the switch. The method may use the digital twin to generate a second image and a second configuration by replaying the first configuration on the first image. The method may communicate the second image and the second configuration from the digital twin to an ASIC memory of an ASIC associated with the switch by applying batch direct memory access (DMA).
Owner:CISCO TECHNOLOGY INC

CNN-oriented batch matrix multiplication parallel optimization method and system on SW architecture

The invention provides a CNN-oriented batch matrix multiplication parallel optimization method and system on a SW architecture, and belongs to the technical field of artificial intelligence parallel optimization. Comprising the following steps: respectively converting an input feature map and a convolution kernel in a convolution layer into an input matrix and a weight matrix, and processing the input matrix and the weight matrix into a plurality of groups of independent matrix multiplication tasks in batches; the main core encapsulates a matrix multiplication task into a parameter structure array, the parameter structure array is transmitted to the slave core through single DMA, and the slave core divides rows of an input matrix into row block tasks by adopting a dynamic row block division algorithm according to the total number of threads and the height of the matrix; and executing sub-matrix multiplication calculation on the distributed independent row blocks, asynchronously prefetching matrix sub-blocks by adopting a double-buffer DMA (Direct Memory Access), and executing matrix multiply-accumulate calculation. The parallel processing efficiency of batch matrix multiplication between the master core and the slave core of the SW processor can be improved, and the algorithm performance is optimized.
Owner:QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES) +1

Heterogeneous calculation and dynamic model updating method and device for brain-like chip and medium

The invention provides a brain-like chip heterogeneous calculation and dynamic model updating method, which comprises the following steps: constructing a neuron pulse coding and video feature parallel calculation model, carrying out heterogeneous integration on a brain-like chip supporting an SNN pulse neural network and a multi-core processor, and constructing a spatial-temporal feature dual-channel processing unit; developing a task load balancing algorithm based on synaptic weight dynamic allocation, including constructing a resource state matrix, designing a task allocator based on reinforcement learning, and performing decision optimization in a dynamic environment; a direct memory access channel is established between a brain-like chip and a video codec, a dedicated instruction set is expanded, and motion vector data of a video encoder is read through instruction level collaboration; carrying out binary differential coding on the key layer by adopting hierarchical parameter importance sorting, and compressing the update quantity of the key layer by combining a compression algorithm; and migrating a full-amount model to a lightweight model through dynamic knowledge distillation, and dynamically generating an adaptive model in combination with an attention migration loss function.
Owner:BEIJING ENGINEERING DIGITAL INTELLIGENCE (BEIJING) TECHNOLOGY CO LTD

GPU asynchronous direct memory access application

The invention discloses a GPU asynchronous direct memory access application. One embodiment provides a graphics processor, comprising: a base die comprising a plurality of chiplet slots; and a plurality of chiplets, the plurality of chiplets being coupled with the plurality of chiplet slots. A chiplet of the plurality of chiplets includes: a graphics core cluster, the graphics core cluster including a plurality of graphics cores; a distributed shared local memory, the distributed shared local memory including a shared local memory within each of the plurality of graphics cores; and a direct memory access engine within each of the plurality of graphics cores, the direct memory access engine configured to asynchronously copy data from a memory device to the distributed shared local memory.
Owner:INTEL CORP

Data transmission method, device, and storage medium

Embodiments of the present disclosure provide a data transmission method, a device, and a storage medium. In the embodiments of the present disclosure, a new direct memory access interface is provided. In terms of implementation, an agent module is provided for a network interface card device, and a task management queue and a memory management queue are provided between the agent module and an application to serve as a direct memory access channel between the agent module and the application; during data transmission between two ends, on one hand, the two ends can use bypass kernel and zero-copy techniques on the basis of their respective task management queues to improve data transmission efficiency; and on the other hand, the two ends actively implement memory management for data storage on the basis of their respective memory management queues, without memory management interaction with the peer end, so that the interactive operations between the two ends during data transmission can be reduced, and the transmission delay can further be reduced while implementing direct memory access.
Owner:CLOUD INTELLIGENCE ASSETS HOLDING (SINGAPORE) PTE LTD

Key-value cache compression method and system for accelerating large language model inference

Provided is an accelerator for accelerating batch large language model (LLM) inference via key-value cache compression. An accelerator, according to one embodiment, may comprise a plurality of compute cores. Here, each of the plurality of compute cores may comprise a plurality of processing units for processing an LLM inference operation on a per-token basis, and a direct memory access unit for managing operations of reading weights from a memory and writing key-value activation data back to the memory. In addition, the direct memory access unit may comprise a compression engine for managing online compression of the key-value activation data when the key-value activation data is written to the memory, a decompression engine for decompressing the compressed key-value activation data retrieved from the memory, and a memory management unit for managing reading and writing of the compressed key-value activation data in the memory.
Owner:HYPERACCEL CO LTD

Method and apparatus for hardware resource sharing in direct memory access controller

A direct memory access controller (DMAC) includes a virtual channel, a physical channel, and a context manager. The virtual channel is configured to generate a flow control signal for executing a direct memory access (DMA) command. The physical channel includes read and write control logic to initiate data transmission from the source device to the destination device in accordance with the DMA command in response to the flow control signal. In one embodiment, a context manager includes: allocation logic configured to arbitrate between allocation requests from virtual channels and to allocate physical channels to selected virtual channels; and routing logic configured to route a flow control signal from the selected virtual channel to the allocated physical channel, and to route a status signal between the allocated physical channel and the selected virtual channel.
Owner:ARM LTD

Multi-channel ultrasonic transducer phased array driving method based on DMA

The invention discloses a multichannel ultrasonic transducer phased array driving method based on direct memory access. The method comprises a waveform data buffer length determination step, a data path construction step, a phase parameter calculation step, a waveform data synthesis step and a DMA driving step. An independent DMA channel is configured for each GPIO port or each group of GPIO ports corresponding to the transducer array by utilizing the DMA characteristic of a universal microcontroller, and buffer data is transmitted to the output data registers of the GPIO ports automatically and periodically by DMA hardware, so that multi-channel driving signals with accurate frequency and controllable phases are generated in parallel under the condition that CPU (Central Processing Unit) intervention is not needed. According to the invention, high-precision time sequence control comparable with an FPGA (Field Programmable Gate Array) scheme is realized with extremely low hardware cost and CPU (Central Processing Unit) resource occupation, the problem of time sequence jitter of an existing MCU (Microprogrammed Control Unit) scheme is solved, and the method has high stability, high efficiency and excellent expandability.
Owner:SOUTHEAST UNIV