Patsnap Eureka AI that helps you search prior art, draft patents, and assess FTO risks, powered by patent and scientific literature data.
96 results about "Cache coherence" patented technology
Filter
Efficacy Topic
Property
Owner
Technical Advancement
Application Domain
Technology Topic
Technology Field Word
Patent Country/Region
Patent Type
Patent Status
Application Year
Inventor
In computer architecture, cache coherence is the uniformity of shared resource data that ends up stored in multiple local caches. When clients in a system maintain caches of a common memory resource, problems may arise with incoherent data, which is particularly the case with CPUs in a multiprocessing system.
The invention relates to the technical field of power monitoring systemnetwork security and real-time micropatch hot deployment, and discloses a power disaster recoverysystem-oriented micropatch non-inductive deployment engine, a resource scheduling method, a system, equipment and a medium, and the method comprises the steps: capturing system events through a kernel eBPF probe, and carrying out feature extraction and model reasoning; generating and transmitting an encrypted scheduling token; loading and verifying a patch fragment by a patch agent, inserting a jump instruction through a kernel interface to redirect an execution stream, and maintaining multi-kernel cache consistency; fusing multi-source telemetry data to carry out fusing judgment, realizing network isolation and calling a key service to cancel a key; and collecting runtime indexes and performing trend prediction, triggering a recovery or rollback operation according to a result, and storing an operation result and data through a block chain. According to the method, through combination of deep fusion of multi-source heterogeneous data, dynamic reasoning of a knowledge graph and strategy optimization of reinforcement learning, efficient perception and defense of a complex attack scene of a digital power grid are realized.
The invention discloses a data transmission optimization method and system in hybrid deployment of a real-time system and a non-real-time system based on Jailhouse, belongs to the technical field of data transmission, and solves the problems of low data transmission efficiency, high copy overhead, lack of universality of interfaces and the like in an existing system. Comprising the steps that a hybrid deployment platform including a real-time system and a non-real-time system is built on a multi-core processor platform, and the hybrid deployment platform runs in an isolation environment of a Jailhouse partition manager; building an abstract transmission interface layer as a unified channel for accessing a shared memory by a real-time system and a non-real-time system; establishing a plurality of shared memory channels in the abstract transmission interface layer and scheduling according to the priority of tasks; establishing a data transmission path, and compressing a memory copy operation into one time by adopting a memory mappingmultiplexing technology and a DMA (Direct Memory Access) collaboration mechanism; and a cache consistency maintenance technology is adopted, so that the data consistency is ensured, and the optimization of data transmission is completed. The method is suitable for application scenes such as industrial automation, intelligent connected automobiles and edge calculation.
The invention discloses a circuit, a data processing method, equipment, a medium and a program product in the technical field of computers. In the application, the three-dimensional heterogeneous computing body layer can adapt to diversified computing power requirements, and is beneficial to realizing higher routing and higher data transmission efficiency between nodes. The shared cache layer is beneficial to realizing cache consistency of different heterogeneous computing nodes, and the cache utilization rate is improved. The internal interconnection module and the external interconnection module provided by the interface layer are easy to realize internal and external efficient communication. A first channel controller provided by the control layer can realize direct connection of physical channels between the control layer and the three-dimensional heterogeneous computing body layer and direct connection of physical channels between the control layer and the shared cache layer; a second channel controller provided by the control layer can realize direct connection of physical channels between the control layer and the interface layer; therefore, the communication delay between different levels can be reduced, the data transmission efficiency between different levels can be improved, and a high-bandwidth application scene can be met.
Cache memory systems employing multiple-level hierarchy cache coherency architecture, and related methods and computer-readable media. A processor-based system includes separate dies that each have a processor and local cache memory logically forming a portion of global cache memory for a systemaddress space. To provide a single point of cache coherency in the global cache memory, the processor-based system includes a proxy cache controller circuit in each die, and a global cache controller circuit. The global cache controller circuit can communicate with the proxy cache controller circuits to maintain single point of cache coherency in the global cache memory. Thus, a cache coherency protocol based on a single point of cache coherency can be implemented. However, the proxy cache controller circuits are also capable of locally servicing memory requests solely within its die, when possible to maintain cache coherency, to provide lower latency memory transactions
The invention discloses a control system supporting concurrent access of multiple hosts, a storage device and a medium, and relates to the technical field of data storage, the system runs on a control module externally connected with the storage device, and the system comprises a communication connection management unit used for establishing connection with multiple hosts through at least two host interfaces; the access request receiving unit is used for receiving access requests sent by different hosts through an uplink port of the multi-host switching chip, and the access arbitration and scheduling unit is used for carrying out real-time arbitration and unified scheduling on the requests according to a preset strategy; and the instruction forwarding and execution unit forwards an arbitrated instruction to the storage unit for execution through a downlink port, and the system further integrates a cache consistencymanagement unit, a configurable arbitration strategy unit, a multi-mode storage space management unit and the like, so that the data consistency, the access efficiency and the system reliability during multi-host concurrent access are ensured. According to the scheme, efficient and safe parallel access of a plurality of hosts to a single storage device is realized.
The invention discloses a maintenance, recording and transmission method, device and equipment for cache consistency, and the recording method comprises the steps: responding to a first processing unit in a processing unit array to cache a first data block, and obtaining a first position coordinate of the first processing unit in the processing unit array, the first coordinate comprises a plurality of values respectively corresponding to a plurality of dimensions of the processing unit array; according to the first coordinate data, a first directory entry corresponding to the first data block in the cache consistencydirectory is updated, the first directory entry comprises a plurality of arrays corresponding to a plurality of dimensions, and in the first directory entry, the first data block corresponds to the first directory entry; the array corresponding to each dimension is used for marking dimension coordinates of the processing unit which caches the first data block on the dimension. The recording method can reduce the directory capacity.
The present invention relates to the field of computers. Disclosed are a method for adaptively and jointly using cache coherencedirectory entries, and a computer program product. The method comprises: in response to a read request of a processor core, acquiring a first entry set corresponding to a cache; when there is an available entry in the first entry set, storing ID information of the processor core in the available entry; when there is no entry in the first entry set or there is no available entry in the first entry set, applying for a new entry and storing the ID information in the new entry; in response to a write request of the processor core, acquiring a second entry set corresponding to the cache; determining corresponding processor cores on the basis of ID information stored in entries in the second entry set, and controlling each processor core to execute an operation of invalidating a copy of the cache; and enabling the second entry set to only comprise one entry, and storing the ID information in the entry. The directory capacity is fully utilized and snoop operations for all processor cores can be reduced.
Cache memory systems employing a multi-level hierarchy cache coherency architecture, and related methods and computer readable media. A processor-based system includes multiple independent dies, each die having a processor and a local cache memory that logically forms part of a global cache memory in a systemaddress space. To provide single point cache coherency in the global cache memory, the processor-based system includes a proxy cache controller circuit in each die, and a global cache controller circuit. The global cache controller circuit can communicate with the proxy cache controller circuits to maintain single point cache coherency in the global cache memory. Thus, a single point cache coherency protocol can be implemented. However, the proxy cache controller circuits can also be able to locally service memory requests within their own die only, to provide lower latency memory transactions, while still being able to maintain cache coherency.
The invention discloses a concurrent data processing method and a related device, and the method comprises the steps: responding to a deletion request for an original datarecord in a memory stable region, and enabling an Owner Entity to use a write instruction of a preset mechanism to carry out the state marking of whether the original datarecord is in a deletion state or not on the original datarecord; the Sharing Entity reads the original data record in the memory stable region by using a read instruction of a preset mechanism, and the read original data record is preprocessed and then sent to the Owner Entity; and the Owner Entity finds out an original data record corresponding to each preprocessed data record in the memory stable area according to the data record identifier, and rechecks a state marking result of the found original data record. Wherein the read-write instruction does not use any synchronous primitive, atomic operation and the like according to a preset mechanism, so that the cache consistency flow generated by the read operation can be avoided from the source, the problems of pipeline pause and blockage on a read-write path are eliminated, and the business processing efficiency and the data synchronous updating effect are improved.
Embodiments of the present application relate to the technical field of computers, and provide a remote data processing method, an apparatus, and a computing device, capable of improving the efficiency of remote data processing. The method is applied to a first device, and the first device remotely communicates with a second device. The method comprises: receiving an operation request for data to be processed sent by the second device, the operation request comprising: a read request or a write request; in response to the operation request, when a preset condition is satisfied, performing, on the data to be processed, an operation indicated by the operation request, the preset condition being used for indicating that the data to be processed is in a cache coherence state; and in response to the operation request, when the preset condition is not satisfied, performing an invalidation operation on the data to be processed, the invalidation operation being used for invalidating the data to be processed in a replica device, the replica device being a device that caches the data to be processed.
The invention discloses a graphics processor, a cache data sharing method, equipment, a medium and a program product, and relates to the technical field of processors, first-level caches in one-to-one correspondence with computing cores are arranged, and pairwise communication links between the first-level caches and a central label buffer area are constructed through an interconnection interface unit; in cooperation with a global label table recording label information of all first-level cache data blocks, when a first-level cache receives a data reading request of a connected computing core and a target data block is not stored locally, other first-level caches where the target data block is located can be positioned through a central label buffer area; direct data transmission between first-level caches is achieved through the interconnection interface unit without relying on second-level cache transfer, the access pressure of the second-level caches is reduced, the bandwidth bottleneck problem is solved, the data transmission path is shortened, the data accessdelay is reduced, complex design brought by a traditional cache consistency protocol is avoided, the cache resource utilization rate is increased, and the data transmission efficiency is improved. And the overall operation performance of the graphics processor is optimized.
The application discloses a kind of underwater acoustic communication sending control method and system, applied to the processor with multi-level storage and cache architecture.The method allocates double-buffered storage space in advance in physical storage area;Water acoustic communication data is handled to generate time domaindigital signal sequence by multi-carrier digital modulation;In response to the filling completion state of data to the first buffer, cache coherency synchronization operation is performed to write the latest data in the cache of the processor back to the physical storage area;Subsequently trigger the hardware data carrying engine independent of kernel, and send data through external interface;During data carrying, fill the subsequent generated data to the idle second buffer.The application solves the problem of memory access conflict and data transmission break under high sampling rate through multi-level storage isolation and software and hardware collaborative zero-copy pipeline mechanism, and guarantees the high quality, high continuity sending of OFDM underwater acoustic signal.
A method includes receiving a first request to allocate a line in an N-way set associative cache and, in response to a cache coherence state of a way indicating that a cache line stored in the way is invalid, allocating the way for the first request. The method also includes, in response to no ways in the set having a cache coherence state indicating that the cache line stored in the way is invalid, randomly selecting one of the ways in the set. The method also includes, in response to a cache coherence state of the selected way indicating that another request is not pending for the selected way, allocating the selected way for the first request.
In the described example, the coherent memory system includes a central processing unit (CPU) and level 1 and level 2 caches. The CPU is configured to execute program instructions (1000) to manipulate data in at least a first or second security context. Each of the first and second caches stores (e.g., 1050) a security code indicating the at least first or second security context through which data of a corresponding cache line is received. The level 1 and level 2 caches maintain coherence by comparing (1020) the security code of the corresponding cache line and performing a cache coherence operation (1030) in response.
The invention provides a multi-core processor cache coherence communication system based on FreeRTOS, and relates to the field of computer embedded systems, and the system comprises an application layer which comprises a plurality of tasks and is used for completing cross-core data interaction; the cache consistencycommunication layer is in communication connection with the application layer and the kernel layer; the cache consistencycommunication layer comprises a shared memory management module, a cache synchronization strategy scheduling module and a communication component synchronization adaptation module; the shared memory management module is used for tracking the cache state of the shared memory; the cache synchronization strategy scheduling module is used for selecting an optimal cache synchronization strategy according to the task and is linked with the task scheduler; the communication component synchronization adaptation module is used for identifying a corresponding shared memory block; the kernel layer comprises a memory manager, a task scheduler and a communication component; the communication component comprises queues, semaphores and event groups. According to the scheme, the problem of cache inconsistency is solved, and the correctness of cross-core data is ensured.
The invention discloses a data multi-level cache collaborative acceleration method and system, and mainly relates to the technical field of data cache acceleration. Comprising the steps of obtaining data access features and inputting a model fusing sliding window statistics and exponential decay to calculate a dynamic heat value; dynamically scheduling the data to a high-speed memory, a distributed sharing or a local file cache level according to the heat value; optimizing a cross-level query path based on the intelligent routing model; the multi-level cache consistency is maintained by adopting an asynchronous version control mechanism; executing hierarchical collaborative cleaning and elastic capacity expansion according to the real-time popularity and the capacity utilization rate; and realizing hot data preloading by utilizing the prediction model. The method has the beneficial effects that through intelligent scheduling and collaborative optimization, the cache hit rate and the systemthroughput are remarkably improved, the access delay and the storage cost are reduced, and the stability and the data consistency under high concurrency are enhanced.
The invention discloses a cache coherenceinterconnection method, a cache coherence node and a multiprocessor system, relates to the field of computers, and provides the cache coherenceinterconnection method for solving the problems that a traditional cache coherence interconnection module is large in hardware resource overhead and insufficient in flexibility. Analyzing the memory access request to determine a source identifier; the source identifier is a unique identifier corresponding to a source processor of the memory access request sender; generating a monitoring address frame which carries the source identifier and corresponds to the memory access request, and sending the monitoring address frame to other processors through the interconnection module; after the monitoring data load returned in response is received, the memory access request is completed according to the monitoring data load. The method can get rid of dependence on complex structures such as directories, the cache consistency function can be achieved by using fewer hardware resources, decoupling between the cache consistency function and an interconnection structure can be achieved, and application of the cache consistency function is more flexible.
The present application relates to the technical field of multi-core heterogeneous computing, and discloses a storage coherent hub chip based on core integration, a storage coherent arbitration device and an adaptive control method, aiming to solve the technical problems of high cross-core memory access latency, low coupling degree of cache coherence maintenance and memory scheduling, and slow response of operation strategy adjustment under the existing multi-core architecture. The present application takes an independently packaged storage coherent core as a globally consistent unique maintenance node, integrates a single-cycle static addressing architecture, adopts a cache coherence state machine and a memory scheduling controller with deep fusion of logic layers, realizes automatic switching between robust mode and aggressive mode through a pure hardware MHM monitoring unit, and the atomic withdrawal process is transparent to the upper layer. The present application can reduce the cross-core memory access latency by more than 40%, improve the storage access throughput by 25%, while guaranteeing 99.999% operation reliability, adapting to various heterogeneous interconnection protocols, and being applicable to various application scenarios such as servers, high-frequency financial transactions, AR / VR wearable devices, edge computing, etc.
The invention discloses a directory-type distributed data structure implementation method and system, belongs to the technical field of data storage and concurrent processing, and aims to solve the technical problems of poor expandability and low performance of a traditional data structure under a non-cache consistency / partial cache consistency multi-core architecture. Comprising the following steps: dividing a multi-core architecture into a plurality of consistent islands, connecting the islands through a high-speed packet-loss-free network, and performing cross-island communication between cores through two communication primitives; a distributed hash table DHT is constructed, each server node maintains a plurality of hash buckets, the hash buckets adopt a linked list to solve conflicts, and a target server node and the hash buckets are positioned through a hash function; a server node is selected as a synchronizer, the synchronizer is used for distributing or recycling operation key values for a client side, the client side achieves data structure operation of stacks, queues and double-end queues through combination of two short messages and one DHT primitive, and linearization points of all the operations only depend on the moment when the synchronizer returns the key values.
The invention discloses a multi-service demandadaptation method and device based on a single chip, equipment and a medium. The method comprises the following steps: dividing a plurality of processing cores in the system-on-chip into at least one processing domain based on service requirements; configuring a memory access permission for each processing domain so as to realize memory access isolation among different processing domains; configuring a hardware resource access permission for each processing domain so as to realize hardware resource access isolation between different processing domains; if a plurality of processing domains are obtained through division, an inter-domain cache consistency maintenance mechanism is configured; if a processing domain is obtained through division, an intra-domain cache consistency maintenance mechanism is configured; wherein the plurality of processing cores are connected through a consistent interconnection network; and configuring the processing core in the target processing domain into a lock step operation mode according to the function security level requirement corresponding to the service requirement. According to the method and the device, the cost and the time for developing special chips for different scenes are reduced, and multiplexing of hardware resources and improvement of research and development efficiency are realized.
A system and method for managing memory resources. In some embodiments the system includes a first server, including a stored-program processing circuit, a first network interface circuit, and a first memory module. The first memory module may include a first memory die, and a controller. The controller may be connected to the first memory die through a memory interface, to the stored-program processing circuit through a cache-coherent interface, and to the first network interface circuit.
In one embodiment, a device includes a host interface to be connected to a host device via a data communication bus; and at least one processing core to expose a device emulation function to the host device, maintain a peripheral device cache, invalidate a cache entry in peripheral device cache in response to receiving a cache invalidation from a backend function, execute a translation function to translate between a format of the device emulation function and a format of the backend function, receive the cache invalidation from the backend function, translate the cache invalidation to the format of the device emulation function, and 10 provide the translated cache invalidation function to the device emulation function being executed by a processor of the host device.
The embodiment of the invention provides a task processing method, a computing cluster and computing equipment. The method is applied to an acceleration calculation unit, the acceleration calculation unit comprises a plurality of AI accelerators, the AI accelerators comprise a first accelerator, and the accelerators are connected with a shared memorypool based on a cache coherence protocol. The first accelerator obtains historical intermediate state data corresponding to the target reasoning task from a first address field of the shared memorypool; the first accelerator calculates a new lexical element and first intermediate state data corresponding to the new lexical element based on the historical intermediate state data; the first accelerator stores the first intermediate state data into the first address field. By adopting the method, a plurality of AI accelerators can access the shared memorypool through the shared framework, so that efficient sharing and multiplexing of the intermediate state data among different AI accelerators are realized, and the cache utilization rate and the reasoning efficiency are improved.
The described technology provides a method including generating a base snoop filter (SFT) entry for a coherence granule (cogran) in agent cache, the base SFT entry comprising a tracking_information field configured to track a plurality of agent IDs, each agent ID identifying an agent that holds a copy of the cogran, determining a number of agents that hold the copy of the cogran, comparing the and number of agents that store the copy of the cogran with number of the plurality of agent IDs tracked in the tracking_information field of the base SFT entry; and in response to determining that the number of agents that hold the copy of the cogran is greater than the number of the plurality of agent IDs tracked in the tracking_information field of the base SFT entry, selecting a second SFT entry as an extra SFT entry, wherein the extra SFT entry is configured to store a portion of tracking vector wherein each bit of the tracking vector indicates cache validity state of the cogran for a related agent.
The invention relates to cache coherency. In one embodiment, a device includes a host interface to be connected to a host device via a data communication bus; and at least one processing core to expose the device emulation functionality to the host device; maintaining a peripheral equipment cache; invalidating a cache entry in the peripheral cache in response to receiving a cache invalidation from the back-end function; a conversion function is executed to convert between a format of the device emulation function and a format of the back-end function, receive a cache failure from the back-end function, convert the cache failure to the format of the device emulation function, and provide the converted cache failure function to the device emulation function being executed by the host device processor.
A storage device includes non-volatile memory and a storage controller configured to read read data from the non-volatile memory and to write the read data to a memory device. The storage controller includes a management circuit configured to cache the read data in cache memory in a host device by using a cache coherence protocol.
A single-chip-based multi-service requirement adaptation method, device, equipment and medium. The method comprises: dividing a plurality of processing cores in a system-level chip into at least one processing domain based on service requirements; configuring memory access permissions for each processing domain to realize memory access isolation between different processing domains; configuring hardware resource access permissions for each processing domain to realize hardware resource access isolation between different processing domains; if a plurality of processing domains are obtained by division, a domain-to-domain cache consistency maintenance mechanism is configured; if one processing domain is obtained by division, an intra-domain cache consistency maintenance mechanism is configured; wherein the plurality of processing cores are connected through a consistent interconnection network; and the processing cores in the target processing domain are configured to a lockstep running mode according to the functional safety level requirement corresponding to the service requirement. Through the present application, the cost and time for developing special-purpose chips for different scenarios are reduced, and the reuse of hardware resources and the improvement of research and development efficiency are realized.