Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.
53 results about "Memory copy" patented technology
Filter
Efficacy Topic
Property
Owner
Technical Advancement
Application Domain
Technology Topic
Technology Field Word
Patent Country/Region
Patent Type
Patent Status
Application Year
Inventor
Memory Copy can be used to quickly make the tower of blocks. Start with a set of three blocks, each a component. Then make a component from the three components. (Memory Copy only works on components.)
The present application belongs to the technical field of data processing and relates to a network message processing method and apparatus, and a computer device and a storage medium. The method comprises: on the basis of a received target network message, acquiring a target memory block from a pre-constructed shared memorypool, so as to generate a message memory; on the basis of a network-interface-card driver, performing first identification processing on the target network message, so as to obtain a target network message descriptor; on the basis of the target network message descriptor, performing second identification processing on the message memory, so as to obtain a first socket buffer; and performing parsingprocessing on the first socket buffer by means of a network protocol stack, and on the basis of a parsing processing result, reading network message data from a virtual memory corresponding to the message memory, so as to complete the processing of the network message. In the present application, on the basis of a constructed shared memorypool, zero-copy network packet reception can be realized during network message processing, thereby avoiding performance loss caused by the memory copy of messages.
The invention relates to a satellite-ground collaborative carbon sink monitoring method based on multi-core heterogeneous acceleration. The satellite-ground collaborative carbon sink monitoring method comprises the steps that ground multi-source time sequence data and satellite multi-spectral images are collected to form satellite-ground data; the method comprises the following steps: preprocessing satellite-ground data to generate a C-Frame tensor with a self-description frame header; a collaborative flow adaptive bridge C-FAB framework is established and used for directly mapping a C-Frame tensor to edge ai equipment, and traditional multiple times of memory copying are avoided; according to the method, a DS-LUE model is established in edge ai equipment, a carbon sink result of an area needing to be calculated is generated through the DS-LUE model, and rapid and accurate carbon sink estimation and dynamic monitoring are realized in a low-power-consumption and high-throughput edge calculation environment by fusing satellite multispectral images with ground humidity, soil and vegetationsensing data. And the method supports the cooperative operation of complex remote sensingprocessing tasks among different computing resources, and significantly improves the performance and adaptability of remote sensingdata processing.
The invention discloses a data transmission optimization method and system in hybrid deployment of a real-time system and a non-real-time system based on Jailhouse, belongs to the technical field of data transmission, and solves the problems of low data transmission efficiency, high copy overhead, lack of universality of interfaces and the like in an existing system. Comprising the steps that a hybrid deployment platform including a real-time system and a non-real-time system is built on a multi-core processor platform, and the hybrid deployment platform runs in an isolation environment of a Jailhouse partition manager; building an abstract transmission interface layer as a unified channel for accessing a shared memory by a real-time system and a non-real-time system; establishing a plurality of shared memory channels in the abstract transmission interface layer and scheduling according to the priority of tasks; establishing a data transmission path, and compressing a memory copy operation into one time by adopting a memory mappingmultiplexing technology and a DMA (Direct Memory Access) collaboration mechanism; and a cache consistency maintenance technology is adopted, so that the data consistency is ensured, and the optimization of data transmission is completed. The method is suitable for application scenes such as industrial automation, intelligent connected automobiles and edge calculation.
A computer-implemented system for routing artificial intelligence (AI) queries. The system utilizes a zero-copy data pipeline, which processes prompts in memory-mapped buffers to eliminate at least one memory copy operation, thereby reducing latency relative to conventional serialization pipelines. The system continuously verifies AI provider compliance by injecting synthetic prompts containing invisible, Ed25519-signed Unicode watermarks. Algorithmic bias is detected by generating counterfactual “digital twin” prompts and applying Fisher exact statistical testing.Routing decisions for multi-tier autonomous systems are governed by safety-level requirements (ASIL-D, ASIL-B, QM) and may be constrained by external routing directives received via a meta-identifier. A hash-chained manifest, cryptographically signed using Ed25519 and consumed by downstream gateways, is generated for each routing decision, with its Merkle root asynchronously anchored to a blockchain to create a tamper-evident audit trail for regulatory compliance.
The invention discloses an accelerator card resource sharing system, method and device and electronic equipment, and relates to the technical field of accelerator cards, the system comprises an address mapping module and a memory copying engine module, the address mapping module is used for receiving an access request of an accelerator card video memory of a source host, querying a mapping relation of virtual addresses in a dynamic mapping table and copying the virtual addresses to the source host; the memory copy engine module generates target label information for identifying a target data packet transmitted across hosts, the dynamic mapping table reflects a mapping relation between a source host virtual address and a target host accelerator card physical address, and the memory copy engine module copies the target data packet according to the transmission priority and sequence information of the target data packet contained in the target label information. And executing cross-host data transmission to ensure that the data is efficiently transmitted to an accelerator card video memory of a target host. According to the embodiment of the invention, the problems of relatively high delay and insufficient flexibility in coping with the complex demand of acceleration card cluster sharing in the prior art are effectively solved.
The invention relates to the technical field of artificial intelligence, and provides a fusion operator execution method, electronic equipment, a storage medium and a program product, and the method comprises the steps: determining a plurality of to-be-fused target operators; according to the tensors of the multiple target operators, a pseudo address mapping table is generated, and the pseudo address mapping table is used for recording the mapping relation between the actual memory addresses of the tensors and pseudo addresses; operator fusion is carried out on the multiple target operators, a fusion operator is obtained, and the memory address of the tensor of the fusion operator is a continuous pseudo address; and executing the fusion operator, converting the pseudo address into an actual memory address according to the pseudo address mapping table in the execution process, and accessing data of the tensor according to the actual memory address. According to the method, a large amount of extra memory copy overhead caused by physical movement and rearrangement of the original tensor in a traditional operator fusion scheme is avoided, and the computing resource utilization rate and the execution efficiency are remarkably improved.
An opportunistic approach is described to accelerate certain memory copy operations to be performed by a computing engine by executing instructions. An instruction for a memory copy operation can be identified that has a first data type with a smaller number of bits per data element than supported by the computing engine to perform the copy operation. The instruction can be replaced with another instruction that has a second data type with a higher number of bits per data element based on the alignment of the memory addresses for the copy operation and the total number of data elements to be copied. The second data type may not only accelerate the copy operation but also provide better utilization of the underlying hardware of the computing engine.
The invention discloses a data transmission method and device, a storage medium and a computer program product, and relates to the technical field of computers, the method comprises the following steps: determining a time interval range in which a data transmission request needs to wait for processing of a shared memory device; determining memory copyprocessing duration for directly executing the data transmission request among different hosts; if the memory copyprocessing duration is greater than the maximum value of the time interval range and the shared memory device does not have transmission abnormality, performing data transmission among different hosts through the shared memory device; and if the memory copy processing duration is less than the minimum value of the time interval range or the shared memory equipment has transmission abnormity, performing data transmission among different hosts through the server. According to the method, the redundant transmission path is added when the transmission of the shared memory device is abnormal, so that the data accessibility is ensured, the low-delay transmission is maintained, and the reliability and the overall efficiency of data transmission between hosts are improved.
The present invention provides a GPU computing method for computational fluid dynamics simulation, comprising building a hybridprogramming framework based on a CPU / GPU heterogeneous system; establishing an OpenFOAM-GPU data structure; migrating chemical calculations to the GPU; migrating thermal physical quantity calculation tasks from the CPU to the GPU for execution; improving memory copy logic to reduce the frequency of data exchange between the GPU memory and the CPU during the calculation process; and reconstructing a two-dimensional array structure originally used on the CPU into a one-dimensional array structure suitable for GPU parallel computing. The method solves the problem that the numerical simulation method consumes a large amount of computing resources due to low parallel efficiency caused by data communication problems between the CPU and the GPU, thereby reducing the number of communications between the CPU and the GPU, lowering the communication cost, and improving the efficiency of parallel computing and data throughput.
Processing circuitry (16) and an instruction decoder (9) supports a load chunk instruction and a store chunk instruction which can be useful for implementing memory copy functions and other library functions for manipulating or comparing blocks of memory. Number of bytes to load or store in response to these instructions is determined based on an implementation specific condition. As well as loading or storing bytes of data, the load chunk instruction and (10) store chunk instruction also designated a load / store length value as data corresponding to an architecturally visible register, which provides an indication of a number of bytes loaded or stored.
The invention discloses a memory copying method, a memory copying device, a chip and electronic equipment, and belongs to the technical field of computers. The memory copying method comprises the steps that after a direct memory access DMA request submitted by a user mode application program is received, a physical page corresponding to the DMA request is locked; determining a topological relation between the physical page and the DMA channel; determining a target DMA channel according to the topological relation and the idle state of the DMA channel; mapping the locked physical page to an addressing space of the target DMA channel to obtain a mapping relationship between an address of the locked physical page and an address in the addressing space; and copying the data in the locked physical page through the target DMA channel based on the mapping relationship. Through the topological relation between the physical pages and the DMA channels and the idle states of the DMA channels, the DMA channels in the DMA channel resourcepool can be adaptively and dynamically allocated, and the data copying efficiency can be improved.
A method for rapid data exchange between real-time and time-sharing systems. Different cores on a single processor work together to share hardware resources equally, improving overall system performance and efficiency. This not only meets real-time requirements and a large number of application computing requirements, but also reduces investment costs. A simple shared memory and signal notification mechanism is used between multi-core systems to increase data transmission bandwidth and efficiency, while also taking into account real-time and large data volume requirements. Real-time dataprocessing and high-bandwidth data transmission are met. Through address bus communication, FPGAs can process large amounts of data in parallel and in real time, accelerating the data transmission process. The number of memory copies from kernel mode to user mode is reduced, eliminating unnecessary duplication during data transmission, thereby improving performance and reducing resource usage. A high-speed, large-capacity, real-time datainteraction method is designed for the entire process, from data reception to processor processing, operating from the same memory area from start to finish.
This invention discloses a data transmission method, device, storage medium, and computer program product, relating to the field of computer technology. The method includes: determining a time interval within which a data transmission request needs to wait for processing by a shared memory device; determining the memory copyprocessing time for directly executing the data transmission request between different hosts; if the memory copyprocessing time is greater than the maximum value of the time interval and the shared memory device does not exhibit any transmission anomalies, then data transmission is performed between different hosts via the shared memory device; if the memory copy processing time is less than the minimum value of the time interval or the shared memory device exhibits a transmission anomaly, then data transmission is performed between different hosts via a server. This invention adds redundant transmission paths when a shared memory device exhibits a transmission anomaly, ensuring data reachability while maintaining low-latency transmission, thereby improving the reliability and overall efficiency of data transmission between hosts.
A direct memory access (DMA) controller issuing memory copy operations on behalf of a shader at a parallel processor stops issuing copy operations upon a context switch at the shader for a wave. The DMA controller or a trap handler associated with the shader saves the incomplete copy operations to a region of global memory, from which the incomplete operations are restored upon a context resume for the wave.
The application provides a kind of multi-hardware supported deep learning model compiling method and compiler, comprising: obtaining the computational graph of the deep learning model to be compiled and the target device information;The target device information is converted into hardware attribute intermediate form based on the preset hardware model, the linear operationintermediate form is obtained based on the basic operation of deep learning model, the tensor shape intermediate form is obtained based on the tensor of deep learning model;Based on the tensor and the tensor shape intermediate form, obtain the memory intermediate form, and use the heterogeneous memory transmission folding method to optimize memory copy behavior;Based on the obtained linear operation intermediate form, obtain the loop intermediate form, use the method of fusion, stacking and vectorization to optimize the loop in the loop intermediate form, and convert into vector intermediate form, obtain the heterogeneous device data operation rule intermediate form based on the preset calculation model;Based on the plurality of intermediate forms obtained in the foregoing, obtain the executable code that can be inferred on the target device.
The application discloses a data processing method and system in a remote procedure call, and the method comprises the following steps: one party of the remote procedure call acquires to-be-transmitted data, wherein the to-be-transmitted data comprises at least one field; the one party acquires serialized data and data in a predetermined field, wherein the serialized data is obtained by serializing data in other fields of the to-be-transmitted data, the other fields are fields of the to-be-transmitted data except the predetermined field, the predetermined field is a field satisfying a predetermined condition, and the data in the predetermined field has not been subjected to serializationprocessing; and the one party sends the serialized data and the data in the predetermined field to another party of the remote procedure call. The application solves the problem that the RPC cannot meet the requirements due to the large serializationprocessing overhead in the RPC process, thereby saving the memory copy overhead in the serializationprocessing and improving the processing efficiency of the RPC.
In some implementations, an emulation system may store a set of data to a first memory copy location of an emulated environment that is associated with a first virtual host system. The emulation system may copy the set of data from the first memory copy location to a shared memory location of the emulated environment. The emulation system may copy the set of data from the shared memory location to a second memory copy location of the emulated environment that is associated with a second virtual host system. The emulation system may load the set of data from the second memory copy location.
The invention discloses a data processing method and device, computer equipment, a storage medium and a program product, and relates to the technical field of computers. The method comprises the following steps: in response to a data writing request of a producer, writing to-be-written stream data provided by the producer into a preset data stream slice space according to a current writing address, and updating the current writing address according to the to-be-written stream data; and in response to a data reading request of a consumer, reading target stream data from the preset data stream slice space according to the current reading address, providing the target stream data to the consumer, and updating the current reading address according to the target stream data. A producer and a consumer intuitively operate the same memory space based on respective operation addresses, additional memory copying operation of data between a user level and a buffer area is not needed in the process, the problem that resources such as space and time are wasted due to frequent memory copying in an existing memory buffer algorithm is solved, and the memory copying efficiency is improved. And the read-write speed of the cache data is increased.
In some implementations, an emulation system may store a set of data to a first memory copy location of an emulated environment that is associated with a first virtual host system. The emulation system may copy the set of data from the first memory copy location to a shared memory location of the emulated environment. The emulation system may copy the set of data from the shared memory location to a second memory copy location of the emulated environment that is associated with a second virtual host system. The emulation system may load the set of data from the second memory copy location.
This invention discloses a method for accelerating numerical computation software based on CPU-GPU collaboration. The method divides the program into basic blocks and uses a code analysis module to predict the runtime information of these basic blocks, including whether they are computationally intensive tasks, whether they are easily parallelized, and whether the memory copy time of the computation task is less than the CPU execution time. Based on the analysis results, basic blocks that meet the above conditions are marked as GPU modules, and the rest are CPU modules. GPU modules are compiled into GPU code for execution on the fly. If CPU modules have hotspot code, they are compiled into machine code on the fly for execution; otherwise, they are interpreted. This method employs a CPU-GPU collaborative just-in-time compilation approach, fully utilizing CPU and GPU computing resources, achieving higher execution efficiency than CPU or GPU execution alone. It also leverages the advantages of just-in-time compilation of interpreted languages, significantly improving the performance of numerical computation software.
This invention discloses a virtualized log export method, apparatus, terminal, and storage medium, belonging to the field of logging. The method includes: allocating contiguous memory at the EL2 layer as a buffer and initializing a start timestamp; collecting data when an event occurs, determining the mask position, writing the event type and description to a bitmap, writing the relative duration to a preset position, and writing process and CPU information at a fixed length to generate a fixed-length binary event record; when aggregation conditions are met, aggregating the valid data in the buffer into a composite event information structure via memory copy to form a package to be exported; writing the structure to a memory-mapped window area accessible only by EL1, allowing EL1 to read without trapping into EL2; when the window changes from empty to non-empty, EL2 sends only one notification, allowing EL1 to continue reading without additional overhead. This reduces the overhead of cross-exception level notifications and traps, achieving efficient and secure structured log transmission.
Disclosed in the present invention are a CPU capable of quickly processing a memory copy instruction, and a method using same. The CPU comprises: an instruction decoder, a general-purpose register, a memory copy controller, a bus interface, a buffer, an adder, and a comparator. The memory copy controller comprises a state machine, and the state machine comprises an idle state, a read state, and a write state, wherein in the idle state, the memory copy controller waits to receive a valid memory copy instruction; in the read state, the memory copy controller reads data of a source address by means of the bus interface, and temporarily stores the data in the buffer; and in the write state, the memory copy controller writes the data, which is temporarily stored in the buffer, to a destination address by means of the bus interface. The adder is used for updating an address. The comparator is used for determining the end of copying. The present invention can maximize the utilization of a memory bandwidth to greatly improve the memory copy efficiency, can support an arbitrary alignment mode and support interruption, can reduce the power consumption of instruction fetching, and has a simple structure.
The application discloses a kind of hardware perception type vector quantization calculation fusion acceleration instruction parallel processingsystem and method.For the memory bandwidthbottleneck problem of asymmetric distance calculation in large-scale vector retrieval, the application is inside processor register, 8-bit quantization encoding is loaded, zero extension to 32-bit floating point precision, affine correction based on quantization offset and scale factor, and distance accumulation operation for query vector are fused into single instruction multiple data stream driven in-register execution pipeline.The method eliminates the intermediate memory copy and register overflow operation in the traditional scheme, improves the throughput capacity of vector retrieval and processor resource utilization under the premise of maintaining the calculation accuracy.
The present invention discloses a hydropower LCU controller signalinteraction method, system, device, and storage medium thereof. The method includes the following steps: dividing the same signal in the application processes of different tasks into a signal input task and a signal output task according to the signal conduction direction; creating a signal database, and registering the signal data information in the signal input task and the signal output task in the signal database by associating and combining them; the signal database transmits the data information of each task to a shared memory; the signal output task copies the signal value content to the shared memory, and the signal input task reads the signal value content copied by the signal output task from the shared memory, completing the signal data transmission in one step. This method frees a large amount of data interaction during the subsequent controller operation from traditional time-consuming methods such as FIFOs and semaphores, and only requires a single memory copy to complete the signal transmission, reducing unnecessary resource loss in the controller.