Patents
Literature
Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.

40 results about "Memory barrier" patented technology

A memory barrier, also known as a membar, memory fence or fence instruction, is a type of barrier instruction that causes a central processing unit (CPU) or compiler to enforce an ordering constraint on memory operations issued before and after the barrier instruction. This typically means that operations issued prior to the barrier are guaranteed to be performed before operations issued after the barrier.

Method and apparatus for supporting distributed graphics and compute engines and synchronization in multi-dielet parallel processor architectures

This disclosure describes supporting distributed graphics and compute engines in a multi-dielet processor, such as, for example, a multi-dielet graphics processing unit (GPU), architectures and synchronization in such architectures. Each multi-dielet processor includes a hardware-implemented remapping capability and / or a hardware-implemented memory barrier capability.
Owner:NVIDIA CORP

Multi-agent collaboration method and system based on shared memory data exchange

The invention relates to the technical field of multi-agent collaboration, in particular to a multi-agent collaboration method and system based on shared memory data exchange, and the method comprises the steps: creating and initializing a shared data area which is mapped and accessed by a plurality of agent processes in a physical memory; constructing an annular buffer structure on the shared data area, and realizing a lock-free read-write pointer propulsion mechanism based on atomic operation and a memory barrier so as to support a plurality of intelligent agents to perform data read-write concurrently; asynchronous notification of data updating is carried out between agent processes through a lightweight event notification mechanism; and each agent performs autonomous decision making and task scheduling based on the shared data so as to realize distributed collaboration. A data block version control and verification mechanism is introduced into the annular buffer structure, and when it is detected that the agents are abnormal, resources of the abnormal agents are automatically recycled, and a data consistency recovery process is triggered. Multiple times of data copying between a kernel mode and a user mode are eliminated, and communication delay is reduced.
Owner:SHANDONG INSPUR SCI RES INST CO LTD

Instruction synchronization method and artificial intelligence chip

The invention provides an instruction synchronization method and an artificial intelligence chip, and relates to the technical field of artificial intelligence chips, and the method comprises the steps: after a processing core sends a calculation instruction, sending a barrier instruction with the same memory address, enabling the calculation instruction and the barrier instruction to enter a cache in an order-preserving manner, and guaranteeing that the cache receives the calculation instruction and achieves calculation; therefore, a large amount of memory fence operation is reduced, and the processing time delay of the memory access instruction is reduced. And for the plurality of processing cores bound with the same target barrier identifier, recording the number of synchronized instructions of the plurality of processing cores in real time. When the number of the synchronized instructions is equal to the expected synchronization number, returning a synchronization success message to the plurality of processing cores so as to realize instruction synchronization of the plurality of processing cores; in the process, each processing core does not need to circularly read the atomic accumulation result, so that the occupation of bandwidth resources such as bus bandwidth and direct interconnection bandwidth is greatly reduced, the instruction synchronization overhead is reduced, and the efficiency of the whole system is improved.
Owner:SHANGHAI BIREN TECH CO LTD

System memory peak bandwidth measurement method and electronic equipment

The invention discloses a system memory peak bandwidth measurement method and electronic equipment, and relates to the technical field of bandwidth measurement, the memory bandwidth measurement process is optimized through cooperation of dynamic selection of a maximum width vector instruction, forced non-cache access and NUMA perception binding, and the bandwidth measurement efficiency is improved by utilizing the vector processing capacity of a CPU (Central Processing Unit). The data throughput of a single operation is improved to 512 bits or even higher, through cooperation of a non-temporary instruction and a memory barrier instruction, a cache level is bypassed, interference of cache hit or jitter on a measurement result is eliminated, it is ensured that the measurement result truly reflects the performance of a memory system, and the measurement accuracy is improved. Through an automatic NUMA binding mechanism, delay and congestion caused by cross-node access are avoided, and the accuracy, stability and repeatability of a test result are improved.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD

Multi-module communication method and device of string type energy storage converter

The invention discloses a multi-module communication method and device for a string type energy storage converter. The multi-module communication method comprises the following steps: respectively connecting a power module and a centralized control module through a dual-port random access memory; four independent buffer areas A1, A2, B1 and B2 are divided for each storage monomer, A1 and A2 serve as double buffer areas of the power module, and B1 and B2 serve as double buffer areas of the centralized control module; the method for writing and reading the independent buffer areas A1 and A2 and the independent buffer areas B1 and B2 comprises the following steps of: initially setting a read-write pointer, and finishing a flag bit and a memory barrier. According to the main technical scheme and the main effects, high-speed and low-delay data transmission is achieved, a dual-port random access memory array is adopted to replace traditional bus communication, the data transmission delay can be reduced from the millisecond level to the microsecond level, and the real-time acquisition requirement of the power module at the interval of 250 microseconds is met.
Owner:ZHEJIANG NARADA POWER SOURCE CO LTD

Distributed task scheduling method and device based on task state switching

This invention provides a distributed task scheduling method and apparatus based on task state switching, relating to the field of data processing. The method includes: distributively scheduling multiple tasks according to system service instructions and displaying the task states of the multiple tasks in a graphical user interface; responding to a state switching operation for a target task among the multiple tasks, mapping the target TCB and target key state fields corresponding to the target task to a writable memory region within the process in shared memory according to the state switching operation; modifying the state parameters in the writable memory region within the process online using atomic operations corresponding to memory barriers based on the target TCB and target key state fields; responding to the memory modification event corresponding to the completion of the state parameter modification, continuously polling the executor corresponding to the target task and sensing the state change through notification of the memory modification event, and adjusting the executor's own behavior according to the sensed state change.
Owner:CAPITAL INFORMATION TECH DEV CO LTD

Method and apparatus for supporting distributed graphics and computation units and synchronization in parallel multi-dielet processor architectures-memory barriers

This disclosure describes support for distributed graphics and compute engines in a multi-dielet processor, such as a multi-dielet graphics processing unit (GPU), architectures, and synchronization in such architectures. Each multi-dielet processor has a hardware-implemented remapping capability and / or a hardware-implemented memory barrier capability.
Owner:NVIDIA CORP

Speculative remote memory operation tracking for efficient memory barrier

Various embodiments include techniques for performing speculative remote memory operation tracking in a multiprocessor computing system. Conventionally, transfers of data between processors and other components of a computing system require memory synchronization operations to determine that the data is valid and coherent before the data is transferred from a destination to a requesting source. Existing techniques for performing these memory synchronization operations are increasingly inefficient as the number of components in a computing system increases, particularly for remote memory operations. The disclosed techniques track remote memory operations and speculatively perform these memory synchronization operations. As a result, a given memory synchronization operation is often complete prior to the corresponding remote memory operation arrives at the destination, leading to improved efficiency and performance of remote memory operations in complex computing systems.
Owner:NVIDIA CORP

Embedded firmware compilation-free updating method and system

The invention relates to the technical field of firmware updating, in particular to an embedded firmware compilation-free updating method which comprises the following steps: S1, firmware is divided into a plurality of independent function modules through a link script, and each function module is allocated to an independent memory address interval and recorded in a section table; s2, in response to the module replacement instruction, reading an identifier and a version number of a target function module; s3, positioning a target memory address of the target function module according to the section table, and executing compatibility check; s4, if the compatibility check is passed, writing binary data atoms of the new function module into a target memory address interval under the protection of a memory barrier; and S5, updating a module version mark in the section table, and triggering soft restart to enable the new function module to take effect. The firmware is divided into independent modules, the updating efficiency and flexibility of the firmware are improved, real-time loading and replacement of the firmware modules are supported, the flexibility and response speed of the system are improved, an automatic rollback mechanism is introduced, and the stability and safety of the system are guaranteed.
Owner:SHANGHAI SHENSILICON SEMICON CO LTD

Program detection method and device and computing equipment

The embodiment of the invention provides a program detection method and device and computing equipment, and relates to the technical field of computers. The method comprises the steps of receiving a to-be-detected program, performing code detection on the program, and determining a statement where a global variable in the program is located; according to the library file and the statement where the global variable is located, the program is modified, and the modified program is obtained. According to the embodiment of the invention, the global variable is positioned through code detection, the global variable in a program does not need to be manually positioned, and the positioning efficiency and accuracy of the global variable can be improved. Moreover, the library file provided by the embodiment of the invention can provide a programming interface for realizing variable calculation operation, relation operation, logic operation or bit operation and other operation operations on the weak memory model, a memory barrier instruction does not need to be manually inserted, and the modification difficulty is reduced. Therefore, the program modification efficiency is improved. And through program modification, safe operation of the multi-thread program under the weak memory model is ensured.
Owner:HUAWEI TECH CO LTD

Method and apparatus for supporting distributed graphics and compute engines and synchronization in multi-dielet parallel processor architectures -- memory barriers

This disclosure describes supporting distributed graphics and compute engines in a multi-dielet processor, such as, for example, a multi-dielet graphics processing unit (GPU), architectures and synchronization in such architectures. Each multi-dielet processor includes a hardware-implemented remapping capability and / or a hardware-implemented memory barrier capability.
Owner:NVIDIA CORP

A Direct3D 12 resource management method based on virtual universal layout

The application discloses a Direct3D 12 resource management method based on a virtual general layout, and after VKD3D is initialized when VK_KHR_unified_image_layouts is not supported, a mapping matrix of a D3D12 resource state, a virtual general layout, a Vulkan physical layout and a pre-compiled conversion path library are constructed by traversing a Vulkan image format, when a D3D application creates a D3D12 resource, parameters are parsed to divide a sub-resource, a resource handle containing a physical layout pool and a sub-resource view array is created, binding of a Vulkan image and memory is completed, when a resource state is set, a resource handle is inquired, the virtual general layout and the Vulkan physical layout are compared in sequence, redundant requests with no change in semantics and layout are filtered to obtain a resource conversion request, the resource conversion request is combined to generate a batch memory barrier instruction, and the batch memory barrier instruction is pre-submitted to a kernel, and execution is completed by a GPU idle time window.
Owner:北京麟卓信息科技有限公司

Weak memory order risk detection method and device, electronic equipment and storage medium

The application provides a weak memory order risk detection method and device, electronic equipment and a storage medium, and belongs to the technical field of computers. The method comprises the following steps: obtaining event information of a plurality of data race events based on running information of each thread running synchronously at a plurality of detection positions, wherein the event information comprises instruction and function call information of at least two threads executed when the data race occurs; obtaining a plurality of instruction pairs located in the same out-of-order execution window based on the function call information of the instructions related to each data race event and the window size of the out-of-order execution window, wherein the out-of-order execution window is an instruction sequence for realizing instruction out-of-order execution and dynamic scheduling, and each instruction pair comprises two instructions derived from the same data race event; and determining that the detection position corresponding to the instruction pair is a position with weak memory order risk when there is no memory barrier in the out-of-order execution window where any instruction pair is located. The application improves the accuracy of the weak memory order risk detection result of the program.
Owner:T-HEAD (SHANGHAI) SEMICON CO LTD +1

A method for optimizing java code layout without interrupting application services

ActiveCN117270822BLittle impact on application performancelittle impact on performanceCode refactoringSoftware designParallel computingEngineering
This invention discloses a Java code layout optimization method that does not interrupt application services. The method includes: acquiring front-end bottleneck data of the application; when it is determined that code layout optimization is needed based on the front-end bottleneck data, collecting the dynamic control flow graph of the application service; calculating a new code layout based on the dynamic control flow graph; generating new code according to the new code layout; correcting relocation information in all code, that is, redirecting function calls and data references pointing to the old code to the new code. During the correction process, memory barriers and cache maintenance instructions are used to ensure the atomicity and consistency of the modification operations, so as to achieve uninterrupted application services; only reclaiming the memory space used by the old code that is no longer executed to store the new code generated in the next optimization, retaining the function code still on the call stack, so that most of the memory space can be reclaimed without interrupting application services. This method effectively alleviates the processor front-end bottleneck of Java applications and improves application performance.
Owner:ZHEJIANG UNIV

Real-time Simulation Performance Evaluation Method for Multi-core Processors Based on ARMv8 Architecture

The present invention discloses a real-time simulation performance evaluation method based on an ARMv8 architecture multi-core processor, specifically relating to the technical field of multi-core performance evaluation; it includes the following steps: setting up a concurrent monitoring module to collect the access instruction sequence and memory barrier instruction information of the processing core; analyzing the address cross-relationship and timing dependency between access instructions to generate an instruction reordering risk access group; performing concurrent consistency analysis on the risk access group to identify race conditions and record the synchronization method and memory barrier instruction distribution; injecting high-intensity concurrent operations in the simulation environment, continuously tracking the read and write operations of shared variables, and analyzing the instruction execution order; counting the matching degree under different simulation load conditions to generate a real-time simulation performance evaluation index for the weak memory model, which can accurately identify potential consistency problems caused by instruction reordering.
Owner:FANGXIN TECH CO LTD

Asynchronous communication management method and apparatus, electronic device, and storage medium

This application provides an asynchronous communication management method, apparatus, electronic device, and storage medium, relating to the field of chip design and manufacturing technology. The method includes: modifying the instruction count and / or data volume count in the memory barrier data block of the asynchronous communication unit based on instructions sent by the asynchronous communication unit; and modifying the status bits in the memory barrier data block of the asynchronous communication unit when the instruction count and data volume count meet preset conditions. The method and apparatus provided by this application not only manage the synchronization state between producers and consumers based on memory barriers, ensuring the consistency and correctness of data exchanged between producers and consumers, but also abstract the complex count update and status judgment logic from specific test cases, forming a general asynchronous communication management mechanism that can be implemented by a dedicated memory barrier controller, supports random testing, and improves the flexibility of test verification.
Owner:SHANGHAI BIREN TECH CO LTD

Partially interrupting a write communication channel with a hardware memory barrier device

One example provides a computing device (100) comprising a write communication channel (104) and an initiator device (120) connected to the write communication channel (104). The initiator device (120) has ordering rules for the write communication channel (104). Further, a network on a chip (NoC) device (112) is connected to the write communication channel (104). The NoC device (112) includes a reorder buffer (118). The computing device (100) also comprises a hardware memory barrier (HMB) device (102) connected to the write communication channel (104) such that at least a portion (132) of the write communication channel (104) is routed through the HMB device (102).
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

A non-intrusive Linux memory barrier modification method and information system

This application relates to the field of computer technology, providing a non-intrusive Linux memory barrier modification method and system. The method is applied to scenarios where an information system is migrated from a strong memory model chip architecture to a weak memory model chip architecture. Specifically, it includes: in the newly added adaptation layer, determining the target system calls invoked by the component layer, and constructing a rewritten function with the same name for each target system call. The rewritten function sequentially includes a first memory barrier instruction, an instruction to call the original target system call, and a second memory barrier instruction; compiling the rewritten function into a dynamic link library; and when the business process starts, preferentially loading the custom dynamic library by preloading environment variables to hijack system calls, so that the operating system's dynamic linker redirects the component layer's calls to the target system calls to the rewritten function. This application achieves complete decoupling of business logic and adaptation logic by implanting a memory barrier at the system call layer without modifying the business layer and component layer code.
Owner:SSE INFORMATION NETWORK LTD

Code optimization method and device, electronic equipment and storage medium

The invention provides a code optimization method and device, electronic equipment and a storage medium, and relates to the technical field of computers, in particular to the technical field of code optimization and code compiling. The specific implementation scheme of the method comprises the following steps: inserting a memory barrier code between at least two first memory access codes of each of a plurality of loop code blocks to be optimized in a code snippet to be optimized to obtain a plurality of optimized loop code blocks; the at least two first memory access codes are respectively located in at least two target code sequences in the to-be-optimized loop code block, and the at least two first memory access codes are dependent on each other; inserting a memory barrier code between the at least two second memory access codes to obtain an optimized code snippet for the to-be-optimized code snippet; the second memory access codes are respectively located in at least two loop code blocks in the to-be-optimized code snippets; the at least two loop code blocks comprise optimized loop code blocks; the at least two second memory access codes have a dependency relationship with each other.
Owner:KUNLUNXIN TECHNOLOGY (BEIJING) CO LTD

Data read-write method and device based on memory barrier instruction and program product

The invention discloses a data read-write method and device based on a memory barrier instruction and a program product, relates to the technical field of processors, is applied to a multi-core instruction set processor, and comprises the following steps: determining an initial memory barrier instruction in a current running user program code; the initial memory barrier instruction is a barrier instruction added by a user; checking the initial memory barrier instruction by utilizing a preset barrier instruction checking condition to obtain a checking result; the preset barrier instruction check condition is a condition which is constructed on the basis of a front-back execution sequence of the multi-core instruction and is used for controlling the read operation to obtain the latest value of the data; and adjusting the initial memory barrier instruction according to the check result, and executing subsequent data read-write operation based on the adjusted target barrier instruction. Therefore, the initial memory barrier instruction is adjusted in combination with the barrier instruction check condition, and then the data read-write operation is executed, so that the accuracy of the memory barrier function can be ensured, and the overall performance, safety and stability of the processor are improved.
Owner:SHANDONG YUNHAI GUOCHUANG CLOUD COMPUTING EQUIP IND INNOVATION CENT CO LTD

Partially interrupting a write communication channel with a hardware memory barrier device

One example provides a computing device comprising a write communication channel and an initiator device connected to the write communication channel. The initiator device has ordering rules for the write communication channel. Further, a network on a chip (NoC) device is connected to the write communication channel. The NoC device includes a reorder buffer. The computing device also comprises a hardware memory barrier (HMB) device connected to the write communication channel such that at least a portion of the write communication channel is routed through the HMB device.
Owner:MICROSOFT TECHNOLOGY LICENSING LLC

Historical debugging-oriented multi-thread race condition accurate reproduction method

The invention discloses a historical debugging-oriented multi-thread race condition accurate reproduction method, which comprises the following steps of: establishing a semantic equivalence mapping table of a synchronization instruction between different target architectures by defining three causal relationships of release acquisition, wake-up execution and memory visibility, recording a user mode and kernel mode synchronization event when a first target architecture executes a target program, and replaying the user mode and kernel mode synchronization event when the first target architecture executes the target program. A bidirectional causal chain with confidence is constructed in combination with causal conditions, a replay engine is started in a second target architecture during replay, synchronization operation is recognized and matched with an original synchronization event in the semantic equivalence mapping table and the causal chain, the semantic reproduction state of a core dependency event is verified, and the core dependency event is replayed; and forcibly reproducing a memory, a state flag, a memory barrier and a thread state according to records by neglecting hardware actual execution, finally adding subsequent events into a queue according to confidence, and completing semantic reproduction of the subsequent events according to a recorded information sequence, so that accurate reproduction of a target program execution process under a multi-thread race condition is realized.
Owner:北京麟卓信息科技有限公司

Data synchronization system, method, device and equipment and storage medium

The invention discloses a data synchronization system, method, device and equipment and a storage medium, and relates to the technical field of data transmission. The system comprises a sender device and a receiver device which are connected through a switch network. The sender equipment is used for sending a data access request and a memory barrier request; the switch network is used for transmitting the memory barrier request on all paths between the sender equipment and the receiver equipment in a multicast manner; after the memory barrier requests transmitted on all the paths are successfully converged into a converged memory barrier request, the converged memory barrier request is sent to a receiver device; and the receiver equipment is used for determining that one or more data access operations sent by the sender equipment before sending the memory barrier request are completed after receiving the aggregation memory barrier request, and the one or more data access operations are associated with the data access request.
Owner:TENCENT TECHNOLOGY (SHENZHEN) CO LTD

Method and apparatus for rectifying weak memory ordering problem

This application relates to the field of computer technologies, and discloses methods and apparatuses, for example, for rectifying a weak memory ordering problem. An example method includes: determining a read / write instruction set in to-be-repaired code; classifying instructions in the read / write instruction set to determine a target instruction; and inserting a memory barrier instruction between a previous read / write instruction of the target instruction and the target instruction. The read / write instruction set includes a read instruction and / or a write instruction in the to-be-repaired code, and an instruction in the read / write instruction set is used for memory access.
Owner:HUAWEI TECH CO LTD

Data processing method and computer equipment

The embodiment of the invention provides a data processing method and computer equipment, which can be applied to the field of data processing, and comprises the following steps: when a reader calls an interface entering a read critical zone, updating a local first variable (representing whether the reader enters the read critical zone) of the reader to a second state, the read critical zone is configured to enable the recoverer to determine an end moment of the grace period based on the first variable, and execute the read critical zone under the condition that a second variable (representing whether the recoverer is in a state waiting for all readers to leave the grace period) local to the reader is a third state. According to the method, the useless hardware memory barrier is eliminated by additionally adding two variables in each reader, and the timely recovery of the memory is realized. According to the method, the hardware memory barrier is executed only when the recoverer exists, and the read critical zone is directly executed when the recoverer does not exist, so that the execution times of the hardware memory barrier are reduced through the on-demand execution mode, and the performance overhead of the hardware memory barrier in a normal path is avoided while the correctness is ensured.
Owner:HUAWEI TECH CO LTD

Method and apparatus for supporting distributed graphics and computing units and synchronization in parallel multi-dielet processor architectures

This disclosure describes support for distributed graphics and compute engines in a multi-dielet processor, such as a multi-dielet graphics processing unit (GPU), architectures, and synchronization in such architectures. Each multi-dielet processor has a hardware-implemented remapping capability and / or a hardware-implemented memory barrier capability.
Owner:NVIDIA CORP

Method and device for repairing weak memory order problem

This application discloses a method and apparatus for repairing weak memory ordering issues, relating to the field of computer technology. The method can automatically repair weak memory ordering issues in a multi-threaded program during the compilation phase. The method comprises: determining a read / write instruction set in the code to be repaired; classifying the instructions in the read / write instruction set to determine a target instruction; and inserting a memory barrier instruction between the read / write instruction preceding the target instruction and the target instruction; wherein the read / write instruction set includes read instructions and / or write instructions in the code to be repaired, and the instructions in the read / write instruction set are used to access memory.
Owner:HUAWEI TECH CO LTD

A heterogeneous nuclear spin lock implementation system and method

This application provides a heterogeneous core spinlock implementation system and method. The system includes: a shared memory module for supporting atomic operation protocols of the architecture atomic operation unit; an ARM core module and a DSP core module, both used to complete initialization through the atomic operation unit, independently initiate lock requests through the architecture atomic operation unit and CAS logic processing unit, detect in real time whether the lock acquisition conditions are met based on the atomic operation unit, access shared resources based on the memory barrier control unit when the lock acquisition conditions are met, and independently release the lock through the atomic operation unit after completing the access to the shared resources; and an AXI bus interconnect module for realizing the interconnection between the dual cores and the shared memory module, which can adapt to the heterogeneous architecture of ARM cores and DSP cores, improve synchronization efficiency and reliability while ensuring fairness, and fully utilize the hardware performance of the dual cores.
Owner:SHANGHAI SMARTLOGIC TECHNOLOGY LTD

System memory peak bandwidth measurement method and electronic device

The application discloses a system memory peak bandwidth measurement method and electronic equipment, and relates to the technical field of bandwidth measurement. The application optimizes a memory bandwidth measurement process by cooperating with a dynamically selected maximum width vector instruction, forced non-cache access and NUMA awareness binding. The application improves single operation data throughput to 512 bits or even higher by using the vector processing capability of a CPU. The application bypasses a cache level by cooperating a non-temporal instruction with a memory barrier instruction, eliminates the interference of cache hits or jitter on measurement results, ensures that measurement results truly reflect the performance of a memory system itself, avoids delay and congestion caused by cross-node access through an automatic NUMA binding mechanism, and improves the accuracy, stability and repeatability of test results.
Owner:INSPUR SUZHOU INTELLIGENT TECH CO LTD