Data access method and device for shared memory, equipment and storage medium

By setting up local cache and memory for each processor in a multiprocessor system, and using the core local interrupter for synchronization and sharing mechanism, the complexity of synchronization and memory access coordination among processors is solved, and efficient memory access and data consistency management is achieved.

CN120104549APending Publication Date: 2025-06-06SHANDONG YUNHAI GUOCHUANG CLOUD COMPUTING EQUIP IND INNOVATION CENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510167614.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

Synchronization and memory access coordination between processors in multiprocessor systems rely on complex cache coherence protocols and memory controller designs, resulting in complex implementation, high cost and certain delays.

Method used

By setting up local cache and local memory in each processor and using the core local interrupter for synchronous sharing mechanism, memory access and synchronization operations between processors are simplified, dependence on shared memory is reduced, and memory access latency is reduced.

Benefits of technology

It improves memory access efficiency, ensures data consistency and correctness, simplifies the synchronization mechanism, reduces the complexity of system design and implementation, and increases the scalability and flexibility of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104549A_ABST
    Figure CN120104549A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of servers, and discloses a shared memory data access method, device and equipment and a storage medium, which are applied to a first processor in a multiprocessor system. Searching whether the target data is stored in a cache layer and a local memory of the first processor or not; if the target data is not hit, an access request is sent to the first core local interrupter, whether the target data is stored in a first shared memory of the first processor or not is searched according to the address information, if yes, the target data is obtained from the first shared memory, and the ownership of the target data is registered through the first core local interrupter; therefore, direct access can be realized during subsequent reading and writing of the target data. According to the method, each processor preferentially accesses the local cache and the local memory by setting the memory separately, so that the dependence on the shared memory is reduced, the memory access delay is reduced, and the access efficiency of the whole system is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of server technology, and in particular to a method, device, equipment and storage medium for accessing data in a shared memory. Background Art

[0002] With the continuous development of server applications, the application demand for high-end servers has entered an important stage. Complex architecture implementation supports high-performance indicators of high-end server systems, namely high security, high availability, high reliability, etc. At present, with the continuous growth of computing and data processing needs, high-speed data exchange and coordinated use of computing power between multiple host systems have become an important challenge.

[0003] In modern computing systems, multi-processor system on chip (MPSoC) has become a common architecture for achieving high performance and parallel computing. In this system, multiple processors can access shared memory simultaneously, and an effective method is usually required to manage memory access to ensure data consistency and correctness. However, in this multi-processor system, synchronization and memory access coordination between processors usually rely on complex cache consistency protocols and memory controller designs, which makes this method complex to implement, costly, and has certain delays. Summary of the invention

[0004] The present application provides a shared memory data access method, device, equipment and storage medium to at least solve the problem that synchronization and memory access coordination between multiple processors are complex, costly and have certain delays.

[0005] The present application provides a shared memory data access method, which is applied to a first processor in a multi-processor system, and the method includes:

[0006] When the first processor expects to access target data in the system shared memory resource, searching in the cache layer and local memory of the first processor whether the target data is stored;

[0007] If there is no hit, an access request is sent to the first core local interrupter, and the access request carries the address information;

[0008] Searching whether the target data is stored in a first shared memory of the first processor according to the address information, the first shared memory being a part of the system shared memory resources;

[0009] If stored, the target data is obtained from the first shared memory, and the ownership of the target data is registered through the first core local interrupter, so that the target data can be directly accessed when it is subsequently read and written.

[0010] The present application provides a shared memory data access device, the device comprising:

[0011] A first search module, configured to search whether the target data is stored in the cache layer and the local memory of the first processor when the first processor expects to access the target data in the system shared memory resource;

[0012] A sending module, used for sending an access request to the first core local interrupter in case of a miss, wherein the access request carries address information;

[0013] A second search module is used to search whether the target data is stored in the first shared memory of the first processor according to the address information, and the first shared memory is a part of the system shared memory resources;

[0014] The read / write module is used to obtain the target data from the first shared memory when searching for the stored target data, and register the ownership of the target data through the first core local interrupter so as to directly access the target data when reading and writing it later.

[0015] In addition, the present application also provides a processor, including a core, a local memory and a first core local interrupter, wherein the local memory and the first core local interrupter are respectively connected to the core;

[0016] Wherein, computer instructions are stored in the local memory; when the kernel executes the computer program through the first core local interrupter, the steps of implementing any of the above-mentioned shared memory data access methods are implemented.

[0017] The present application also provides a computer-readable storage medium, on which a computer program is stored, wherein when the computer program is executed by a processor, the steps of any of the above-mentioned shared memory data access methods are implemented.

[0018] The present application also provides a multi-processor system, the system comprising at least a first processor and a second processor, the first processor and the second processor being connected via a high-speed interface;

[0019] The first processor includes: a first core, a cache layer, a local memory and a first core local interrupter, and at least part of the local memory is a first shared memory; the second processor includes: a second core, a cache layer, a local memory and a second core local interrupter, and at least part of the local memory is a second shared memory;

[0020] Among them, the first processor uses the first kernel, cache layer, local memory and the first core local interrupter to execute the steps of any of the above-mentioned shared memory data access methods; when the first kernel fails to find the target data from the first shared memory, the second processor searches for the target data through the second kernel, cache layer, local memory and the second core local interrupter, and transmits it to the first processor.

[0021] The shared memory data access method, device, equipment and storage medium provided by the present application have at least the following beneficial effects:

[0022] First, improving memory access efficiency: The present application distributes memory on different processors, for example, setting a local cache and a local memory in the first processor, and the local memory also includes a first shared memory. When each processor needs to access data, it gives priority to accessing the data in the local cache and the local memory. If the data to be accessed is stored in the local cache or the local memory, it can be accessed directly without accessing through the shared memory, thereby reducing the dependence on the shared memory and reducing the memory access latency. Therefore, the speed of directly reading data from the local cache is fast, thereby improving the access efficiency of the overall system.

[0023] Second, ensure data consistency and correctness: The present application is equipped with a core local interrupter inside each processor, such as the first core local interrupter. The synchronous sharing mechanism of the first core local interrupter can effectively manage and coordinate data access and modification requests between multiple processors. For example, when the accessed data is not hit in the cache layer and local memory of the first processor, the target data is found in the first shared memory according to the address information, and the target data is read, thereby ensuring the consistency and correctness of the data and improving the reliability of the system.

[0024] Third, simplify the synchronization mechanism: This application uses the core local interrupter to handle data synchronization and ownership registration. After the ownership is registered, if the target data needs to be accessed again, it can be read and written directly without the need to request through the core local interrupter again, thereby simplifying the synchronization mechanism in the multi-processor system, reducing the complexity of system design and implementation, and increasing the scalability and flexibility of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the specific implementation methods of the present application or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0026] Figure 1is a structural diagram of a multi-processor system provided according to an embodiment of the present application;

[0027] Figure 2 is a schematic diagram of the structure of a local memory provided according to an embodiment of the present application;

[0028] Figure 3 It is a flowchart of a shared memory data access method provided according to an embodiment of the present application;

[0029] Figure 4 It is a flowchart of another shared memory data access method provided according to an embodiment of the present application;

[0030] Figure 5 It is a flowchart of another shared memory data access method provided according to an embodiment of the present application;

[0031] Figure 6 It is a flowchart of another shared memory data access method provided according to an embodiment of the present application;

[0032] Figure 7 is a structural block diagram of a data access device provided according to an embodiment of the present application;

[0033] Figure 8 It is a schematic diagram of the hardware structure of a computer device provided according to an embodiment of the present application. DETAILED DESCRIPTION

[0034] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of this application.

[0035] It should be noted that, in the description of this application, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in this application are used to distinguish similar objects, and are not used to describe a specific order or sequence.

[0036] In order to enable technicians in this technical field to better understand the present application, first, the background content and relevant technical terms involved in this application are introduced.

[0037] (1)RiscV processor

[0038] RISC-V (Reduced Instruction Set Computer-V, the fifth generation of reduced instruction set computer architecture) is an open instruction set architecture (ISA, Industry Standard Architecture, which is an international interface system for computers and other devices) designed for computer hardware for various purposes, from embedded systems to supercomputers. The RISC-V instruction set is based on the principles of reduced instruction set computers (RISC), with a concise and clear instruction set design that is easy to understand and implement. The RISC-V processor is a central processing unit (CPU) implemented based on the RISC-V instruction set architecture. Its design can follow the open standards established by the RISC-V International Organization, and can also be customized and optimized as needed.

[0039] (2) Shared Memory

[0040] Shared memory refers to a memory area shared between multiple concurrently executed processes or threads. In the shared memory model, each process can communicate and exchange data by reading and writing memory. Shared memory is usually managed and maintained by the operating system, and each process can communicate and share data with each other by reading and writing data in the shared memory area. Shared memory is a common concurrent programming model and is widely used in parallel computing and multi-threaded programming in multi-processor systems.

[0041] (3) Multi-way Server

[0042] A multi-way server is a system consisting of multiple servers that is used to handle a large number of network requests or tasks. In a multi-way server system, multiple servers can process requests from clients at the same time, thereby improving the system's concurrent processing capabilities and throughput. Multi-way server systems usually adopt load balancing and distributed scheduling strategies to ensure that each server can fully utilize resources and achieve fast response and efficient processing of tasks. Multi-way server systems are widely used in Internet services, distributed computing, data centers and other fields.

[0043] (4) Cache and cache lines

[0044] A cache line is the smallest unit of data in a cache in a computer architecture. It is a continuous piece of memory space, usually a power of 2, usually 64 bytes or 128 bytes. The size of a cache line is determined by the computer architecture and processor design, and its purpose is to improve the efficiency of memory access. When the processor reads data from the main memory, it divides the data into cache line sizes and stores it in the cache. In this way, if the processor needs to access adjacent data again, it can reduce the number of accesses to the main memory through the locality principle provided by the cache line, thereby improving the execution efficiency of the program.

[0045] (5) Core local interrupter

[0046] Clint (Core Local Interruptor) is a key component in the RISC-V architecture, which is responsible for handling local interrupts, especially software interrupts and timer interrupts. These two functions are frequently used by the core, so Clint is designed to provide an efficient mechanism that maps directly to the kernel interrupt vector.

[0047] In addition, Clint is also used to manage and distribute core-local interrupts, including software interrupts and timer interrupts. The specific working principle is: Clint sends software interrupts or timer interrupts directly to the specified hart (hardware thread) through a fixed interrupt number and priority. There is no arbitration in this process, so interrupt requests can be responded to quickly.

[0048] The technical solution of the present application is applied to the aforementioned RiscV architecture. In a general multi-processor system, synchronization between processors and memory access coordination usually rely on complex cache consistency protocols and memory controller designs. Furthermore, the existing defects include:

[0049] 1. Traditional Cache Coherence Protocol

[0050] Traditional multi-processor systems usually use complex cache consistency protocols, such as MESI (Modified, Exclusive, Shared, Invalid) and MOESI (Modified, Owned, Exclusive, Shared, Invalid), to manage multiple processors' access to shared memory. These protocols ensure data consistency between multiple processors by tracking the state of each cache line. However, this method has high implementation complexity, requires additional hardware resources and high inter-processor communication latency. Especially in high-performance computing systems, cache consistency protocols may become a bottleneck for system performance. This method is complex to implement, has high hardware costs, has high inter-processor communication latency, and may even become a performance bottleneck for high-performance systems.

[0051] 2. Bus-based synchronization method

[0052] This synchronization method is based on the bus synchronization mechanism, and multiple processors communicate and synchronize through a shared bus. When the processor accesses the shared memory, the bus arbitration mechanism ensures the exclusivity of access. This method is simple and direct, but the bus bandwidth and access conflict problems limit the scalability of the system, especially when the number of multi-core processors increases, the bus becomes the bottleneck of the system, resulting in access conflicts and performance degradation, and it is difficult to expand to large-scale multi-core systems.

[0053] 3. NUMA (Non-Uniform Memory Access) architecture

[0054] NUMA is a common multi-processor memory architecture in which each processor has its own independent memory and can access the memory of other processors. Processors communicate through high-speed interconnect networks to achieve memory sharing and access. The NUMA architecture solves the bottleneck problem caused by the shared bus to a certain extent, but it still requires complex memory management and access control mechanisms, and there is a high latency when accessing across nodes. Inconsistency in memory access may lead to unstable performance.

[0055] 4. Multiprocessor synchronization based on message passing

[0056] In some multi-processor systems, processors are synchronized and exchange data through message passing mechanisms. Each processor executes independently and exchanges data with each other through message queues or other communication mechanisms. This method is suitable for certain specific application scenarios, but an efficient message passing mechanism is required in its implementation, and the message passing overhead may be high under high-frequency inter-processor communication requirements. High-frequency inter-processor communication requirements bring high overhead and are not suitable for all application scenarios.

[0057] In order to solve the above-mentioned defects, the present application proposes a shared memory data access method, which can be implemented based on the multi-way processor chip of the above-mentioned RiscV architecture. The method simplifies the memory access and synchronization operations between processors through the distributed memory and clint core local interrupter module synchronization sharing mechanism, avoids the complexity and high communication delay problems of traditional cache consistency protocols, and solves the bus bandwidth limitation and high delay problems in the NUMA architecture.

[0058] In addition, the method of the present application is simpler to implement and has lower cost, and can provide efficient memory access and data consistency management in a multi-processor system.

[0059] The present application is further described in detail below in conjunction with the accompanying drawings and specific implementation methods.

[0060] This embodiment provides an embodiment of a shared memory data access method. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0061] In this embodiment, a shared memory data access method is provided, which can be used in a multi-processor system, such as Figure 1 As shown, the system may include at least one processor, such as processor 1, processor 2, ..., processor n, where n≥1 and is a positive integer.

[0062] Each processor includes: core, cache layer (cache L1 / L2 / L3), local memory (LocalMemory) and core local interrupt (Clint). Among them, at least part of the resources in the local memory (Local Memory) can be used as shared memory resources, such as Figure 2 As shown, the local memory 1 includes multiple resource blocks, some of which are used to store local data and information, and the other resources can be used as shared memory to store shared data. Shared memory is a globally shared memory module that is accessed by all processors, such as processor 1, processor 2, and processor n, thereby realizing global data sharing and storage.

[0063] In this embodiment, each cache layer includes three layers of cache, such as L1, L2 and L3, and each layer of cache may include multiple cache lines for storing data or other information.

[0064] Optionally, in processor 1, the shared memory of the local memory is called a first shared memory; similarly, in processor 2, the shared memory of the local memory may be called a second shared memory.

[0065] In addition, the core of each processor is responsible for executing the basic operations of the instruction set. The core has an internal L1 (first level) cache to store frequently accessed data and instructions to reduce the access delay to the memory. Furthermore, the core, as a core or kernel, is an important part of the CPU. All CPU calculations, receiving / storing commands, processing data and other tasks are performed by the core.

[0066] Each processor is equipped with independent local memory to store temporary data and intermediate calculation results, thereby reducing dependence on shared memory.

[0067] The core local interrupter (Clint) is used to process synchronization and interrupt requests between cores, and realize data sharing and synchronization operations between multiple processors. The specific functions and working principles are described in the above embodiments, which will not be repeated here.

[0068] In this embodiment Figure 1 In the scenario, the system consists of n processors, each of which can be interconnected through a high-speed custom interface. Figure 1 Not shown, the custom interface includes communication interfaces of one or more protocols, and the number of interfaces can be one or more, which is not limited in this embodiment.

[0069] Figure 3 is a flow chart of a method for accessing data in a shared memory according to an embodiment of the present application. Figure 3 As shown, the method can be applied to any processor in the above system, such as processor 1 or the first processor, and the method includes:

[0070] Step S101: when a first processor desires to access target data in a system shared memory resource, it searches the cache layer and local memory of the first processor to see whether the target data is stored.

[0071] Specifically, the first processor expects to access the target data in the system shared memory resource when the first core (Core1) of the first processor is triggered when executing a read or write request for the target data, or when the method flow of step S101 is triggered when a processing action is executed inside Core 1.

[0072] A specific implementation method of searching whether the target data is stored in the cache layer and local memory of the first processor is to first query the cache layer, such as the L1 cache, to see whether the target data is available. The processor first checks the local L1 cache because the L1 cache is the memory level closest to the processor and has the fastest access speed.

[0073] If the target data is not found, the target data is searched in the local memory 1, such as Local Memory 1. Finding the target data is called a "hit"; correspondingly, not finding the target data is called a "miss" target data. If there is a miss, step S102 is executed.

[0074] One search method is to search whether the target data is stored in the local memory 1 through the data identifier of the target data.

[0075] Step S102: If no, that is, the target data is not hit, an access request is sent to the first core local interrupter, and the access request carries address information.

[0076] If there is no hit, the core 1 of the first processor generates an access request for the target data and sends the access request to the first core local interrupter (Clint 1), so that Clint 1 interrupts the current processing program and starts a search operation for the target data.

[0077] Step S103: searching the first shared memory of the first processor for target data according to the address information.

[0078] The first shared memory is a part of the system shared memory resources.

[0079] After receiving the access request, Clint 1 starts the Clint synchronization mechanism. The main function of the Clint synchronization mechanism is to coordinate and manage data access requests between multiple processors to ensure data consistency and correctness.

[0080] Specifically, Clint 1 searches the first shared memory (Shared Memory 1) for the target data according to the address information carried in the access request, such as the address string ID. If data with the same address string ID is found, it is determined that the target data is stored in the first shared memory. If data with the address string ID is not found, it is determined that the target data is not stored in the first shared memory. If the target data is stored, step S104 is executed.

[0081] Step S104: If yes, that is, the target data is stored in the first shared memory, the target data is obtained from the first shared memory, and the ownership of the target data is registered through the first core local interrupter.

[0082] Specifically, Core 1 obtains the target data from the first shared memory through the transmission interface. And writes the target data into the local memory, such as writing it into the free resource block of the local memory 1 of processor 1. And, register the ownership of the resource block in Clint 1. For example, the ownership is set to be owned by processor 1. After the ownership registration is completed, processor 1 can independently read and write the data of the resource block without requesting Clint 1 again, so that direct access can be achieved when the target data is read and written in the future.

[0083] In addition, when other processors need to access the same resource block, such as when processor 2 wants to access the data of the resource block in local memory 1, it needs to request data ownership from Clint 1 of processor 1 through Clint 2 of processor 2. Clint 1 checks the ownership and performs the corresponding synchronization operation.

[0084] The method provided in this embodiment reduces the dependence on shared memory and reduces memory access latency, thereby improving the access efficiency of the overall system by distributing memory, such as the first shared memory, so that each processor preferentially accesses the local cache and local memory.

[0085] In addition, the Clint synchronization sharing mechanism effectively manages and coordinates data access and modification requests among multiple processors, ensuring data consistency and correctness and improving system reliability.

[0086] Optionally, in a specific embodiment, as Figure 4 As shown, the above step S101: searching whether the target data is stored in the cache layer and the local memory of the first processor specifically includes:

[0087] Step S1011: Generate a memory address to be accessed according to an instruction or program logic, wherein the instruction or program logic may be generated by the core 1 .

[0088] Step S1012: searching the tags of each cache line in the cache layer one by one according to the memory address to see whether there is a matching tag.

[0089] Step S1013: If it exists, determine whether the target data is hit in the cache layer.

[0090] Step S1014: If it does not exist, it is determined that the target data is not hit in the cache layer, and whether the target data is stored in the local memory is searched according to the memory address.

[0091] If the target data is not stored, the above step S102 is executed.

[0092] Specifically, core 1 first searches for the target data on the L1 cache in cache1. The L1 cache is usually divided into multiple cache lines, each of which contains multiple bytes of data and one or more tags associated with the line of data (used to store address information). Processor 1 compares the generated address with the tags of each line in the L1 cache. If a matching tag is found, it means that the required data is already in the L1 cache (cache hit), and processor 1 will read the data directly from the cache.

[0093] If no match is found in the L1 cache (cache miss), processor 1 will check local memory (which could be L2 cache, L3 cache, or higher level cache, or even main memory). Similar to the L1 cache, processor 1 will compare the address with the address in local memory to determine if the data exists in local memory 1.

[0094] If the data is in local memory 1, processor 1 will read the target data and possibly load it into L1 cache for faster access in the future.

[0095] In this embodiment, the above step S03, searching whether the target data is stored in the first shared memory of the first processor according to the address information, specifically includes: the core 1 of the processor 1 searches whether there is a data address that is the same as the physical address in the first shared memory according to the physical address indicated in the address information; if so, it is determined that the target data is stored in the first shared memory; if not, it is determined that the target data is not stored.

[0096] Further, after the above step S103, if Figure 5 As shown, if the kernel 1 does not find the target data in the first shared memory, the method further includes executing steps S105 to S107.

[0097] Step S105: If the target data is not stored in the first shared memory, an access request is sent to the second processor in the multi-processor system through the first core local interrupter.

[0098] Specifically, one implementation method is that the first processor sends an access request to the second core local interrupter of the second processor through the first core local interrupter, so that the second core local interrupter searches for target data in its cache layer, local memory or second shared memory according to the address information carried in the access request.

[0099] For example, processor 1 sends a group access request to processor 2, processor 3 or processor n through Clint 1 to request other processors to obtain target data. For example, Clint 1 sends an access request to Clint 2 of processor 2. After Clint 2 receives the access request, it starts the Clint synchronization mechanism, and Clint 2 searches whether the target data is stored in local memory 2 or the cache layer.

[0100] A specific search process is the same as the aforementioned steps S101 to S103. For example, core 2 (i.e., the second core) of processor 2 first searches the local cache to see whether the target data is stored. If it does not hit, it searches the local memory 2 and the second shared memory to see whether the target data is stored. If it hits, it sends the target data to Clint 1.

[0101] Step S106: receiving target data that the second processor searches for in its cache layer, local memory, or second shared memory according to the access request, where the second shared memory is part of the system shared memory resources.

[0102] Processor 1 receives target data sent by Clint 2 of processor 2 through Clint 1.

[0103] Step S107: Store the target data in the local memory of the first processor, and register the ownership of the target data through the first core local interrupter.

[0104] Core 1 of processor 1 stores the target data in local memory 1 and registers the ownership of the target data as processor 1, so that when the target data is read and written again later, it can be directly read from local memory 1.

[0105] It should be noted that processor 1 receives the target data and records the ownership of the target data. Regardless of whether the data is returned from the shared memory or from the local memory of other processors, after receiving the data, processor 1 needs to register the ownership of the data block in Clint 1. The ownership registration process ensures that in the subsequent data access and modification, processor 1 can independently read and write the data block without requesting it again.

[0106] Processor 1 reports its ownership of the data block to Clint 1. Clint 1 updates the ownership information to ensure that processor 1 can directly access the data block in the future. Processor 1 reads and writes data in local memory. After completing the ownership registration, processor 1 can independently read and write the data block.

[0107] Optionally, in another embodiment, as Figure 6As shown, when the second processor requests the first processor for data in the first shared memory, the method further includes:

[0108] Step S601: when receiving an interrupt request sent from a second processor via a second core local interrupter, checking the usage of data in a data segment to be accessed by the second processor, the data of the data segment is stored in a first shared memory.

[0109] Step S602: If the usage situation is that the first processor completes processing the data of the data segment, a response message is sent to the second core local interrupter through the first core local interrupter, and the response message is used to notify the second processor and authorize the second processor to access the data of the data segment.

[0110] The technical scenario provided by this embodiment is that the core 2 of the processor 2 needs to read data, such as data data2, which may be the same as or different from the target data requested to be accessed by the aforementioned processor 1. This embodiment assumes that the target data is the same as the target data, and the target data is registered and stored in the local memory 1 of the processor 1. In this case, when the processor 2 does not hit the data data2 in the cache layer and the local memory 2, an interrupt is initiated through Clint 2 to notify the processor 1 that the data data2 corresponding to the shared memory 1 in the processor 1 needs to be accessed.

[0111] After receiving this interrupt request, Clint 1 of processor 1 will check the current usage of the data in the data segment accessed by processor 2 by processor 1. After processor 1 completes processing the data in this address segment, it will respond to the access request initiated by processor 2, that is, send a response message to processor 2. The specific implementation method is: notify processor 2 through Clint 1 interrupt mode, authorize processor 2 to access the data in this address segment, and enable processor 2 to read data from local memory 1 of processor 1.

[0112] The shared memory data access method provided by the present application has at least the following beneficial effects:

[0113] First, improving memory access efficiency: The present application distributes memory on different processors, for example, setting a local cache and a local memory in the first processor, and the local memory also includes a first shared memory. When each processor needs to access data, it gives priority to accessing the data in the local cache and the local memory. If the data to be accessed is stored in the local cache or the local memory, it can be accessed directly without accessing through the shared memory, thereby reducing the dependence on the shared memory and reducing the memory access latency. Therefore, the speed of directly reading data from the local cache is fast, thereby improving the access efficiency of the overall system.

[0114] Second, ensure data consistency and correctness: The present application is equipped with a core local interrupter inside each processor, such as the first core local interrupter. The synchronous sharing mechanism of the first core local interrupter can effectively manage and coordinate data access and modification requests between multiple processors. For example, when the accessed data is not hit in the cache layer and local memory of the first processor, the target data is found in the first shared memory according to the address information, and the target data is read, thereby ensuring the consistency and correctness of the data and improving the reliability of the system.

[0115] Third, simplify the synchronization mechanism: This application uses the core local interrupter to handle data synchronization and ownership registration. After the ownership is registered, if the target data needs to be accessed again, it can be read and written directly without the need to request through the core local interrupter again, thereby simplifying the synchronization mechanism in the multi-processor system, reducing the complexity of system design and implementation, and increasing the scalability and flexibility of the system.

[0116] In this embodiment, a data access device is also provided, which is used to implement the above-mentioned embodiments and preferred implementation modes, and the descriptions that have been made will not be repeated. As used below, the term "module" can implement a combination of software and / or hardware of a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceivable.

[0117] This embodiment provides a shared memory data access device, such as Figure 7 As shown, the device includes: a first search module 710, a sending module 720, a second search module 730 and a read-write module 740. In addition, the device may also include other more or fewer modules and units, which is not limited in this embodiment.

[0118] The first search module 710 is used to search whether the target data is stored in the cache layer and local memory of the first processor when the first processor expects to access the target data in the system shared memory resource.

[0119] The sending module 720 is used to send an access request to the first core local interrupter in case of a miss, wherein the access request carries address information.

[0120] The second search module 730 is used to search whether the target data is stored in the first shared memory of the first processor according to the address information, and the first shared memory is a part of the system shared memory resources.

[0121] The read / write module 740 is used to obtain the target data from the first shared memory when searching for the stored target data, and register the ownership of the target data through the first core local interrupter so as to directly access the target data when reading and writing the target data later.

[0122] In some possible implementations, the first search module 710 is specifically used to generate a memory address to be accessed according to an instruction or program logic; search the tags of each cache line in the cache layer for a matching tag according to the memory address; if so, determine that the target data is hit in the cache layer; if not, determine that the target data is not hit in the cache layer, and search whether the target data is stored in the local memory according to the memory address.

[0123] In some other possible implementations, the first search module 710 is further specifically used to search in the first shared memory whether there is a data address that is the same as the physical address based on the physical address indicated in the address information; if so, it is determined that the target data is stored in the first shared memory; if not, it is determined that the target data is not stored.

[0124] In some other possible implementations, the sending module 720 is further configured to send an access request to a second processor in the multi-processor system through the first core local interrupter when the target data is not stored in the first shared memory.

[0125] The above device also includes a receiving module and a storage module. Figure 7 Not shown in FIG.

[0126] The receiving module is used to receive target data that the second processor searches for in its cache layer, local memory, or second shared memory according to an access request, and the second shared memory is part of the system shared memory resources.

[0127] The storage module is used to store the target data into the local memory of the first processor and register the ownership of the target data through the first core local interrupter.

[0128] In some other possible implementations, the sending module 720 is specifically used to send an access request to the second core local interrupter of the second processor through the first core local interrupter, so that the second core local interrupter searches for target data in its cache layer, local memory or second shared memory according to the address information carried in the access request.

[0129] In some other possible implementations, the first search module 710 is also used to check the current usage of data in the data segment to be accessed by the second processor when the receiving module receives an interrupt request sent by the second processor through the second core local interrupter, and the data of the data segment is stored in the first shared memory.

[0130] The sending module 720 is also used to send a response message to the second core local interrupter through the first core local interrupter when the first processor completes processing of the data in the data segment. The response message is used to notify the second processor and authorize the second processor to access the data in the data segment.

[0131] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.

[0132] The data access device in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.

[0133] The present application also provides a processor having the above Figure 7 The data access device shown.

[0134] The processor structure can be Figure 1 Any processor shown is the same, including a core, a local memory and a core local interrupter, wherein the local memory and the core local interrupter are respectively connected to the core;

[0135] The local memory stores computer instructions; the kernel executes the computer instructions through the core local interrupt to achieve the above Figures 3 to 6 The steps of the shared memory data access method are shown.

[0136] Optionally, the present application embodiment further provides a computer device, such as Figure 8 As shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. The various components are connected to each other using different buses for communication, and can be installed on a common motherboard or installed in other ways as needed. The processor can process instructions executed in the computer device, including instructions stored in or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface).

[0137] The processor 10 may be a central processing unit or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be a dedicated integrated circuit, a programmable logic device or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic or any combination thereof.

[0138] Optionally, the processor 10 may be the aforementioned Figure 1 Either processor 1 or processor n is shown.

[0139] The memory 20 stores instructions executable by at least one processor 10, so that at least one processor 10 executes the shared memory data access method shown in the above embodiment.

[0140] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and an application required by at least one function; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid-state hard disk; the memory 20 may also include a combination of the above-mentioned types of memory.

[0141] The computer device also includes an input device and an output device. The processor 10, the memory 20, the input device and the output device may be connected via a bus or other means. Figure 8 The example of the connection via bus is taken in the figure. The input device can receive input digital or character information and generate key signal input related to the user settings and function control of the computer device, such as a touch screen, a keypad, a mouse, a track pad, a touch pad, a pointer, etc. The output device can include a display device, an auxiliary lighting device (e.g., an LED), a tactile feedback device (e.g., a vibration motor), etc.

[0142] In addition, the computer device further comprises a communication interface 30, which is used for the computer device to communicate with other devices or a communication network.

[0143] The computer device provided in this embodiment can be a multi-processor chip based on RiscV. The device realizes efficient data access and synchronization management through the distributed memory and Clint synchronization sharing mechanism. Each processor reduces the dependence on shared memory by independently accessing the local memory and L1 cache. When it is necessary to access the shared memory or the local memory of other processors, synchronization and coordination are performed through Clint to ensure the consistency and correctness of the data.

[0144] Among them, Clint is responsible for data access requests and ownership registration between processors, optimizes communication and data sharing between processors, and improves the overall performance and scalability of the system. Through the above process, this application simplifies data access and synchronization operations in multi-processor systems through split memory and Clint synchronization sharing mechanism. Each processor reduces its dependence on shared memory through independent local memory and cache, while Clint is responsible for data access and synchronization requests between processors to ensure data consistency and correctness. This design not only improves the performance of the system, but also increases the scalability of the system, enabling it to adapt to the parallel operation needs of more processors.

[0145] The present application also provides a multi-processor system. Figure 1 As shown, the system includes at least a first processor and a second processor, and the first processor is connected to the second processor via a high-speed interface.

[0146] The first processor (processor 1) includes: a first core (Core 1), a cache layer, a local memory and a first core local interrupter (Clint 1), and at least part of the local memory is a first shared memory. The second processor (processor 2) includes: a second core (Core 2), a cache layer, a local memory and a second core local interrupter (Clint 2), and at least part of the local memory is a second shared memory.

[0147] The first processor utilizes the first core, the cache layer, the local memory and the first core local interrupt as described above. Figures 3 to 6 The steps of the shared memory data access method are shown.

[0148] If the first core fails to find the target data from the first shared memory, the second processor searches for the target data through the second core, the cache layer, the local memory, and the second core local interrupter, and transmits the target data to the first processor. For the specific process, please refer to the above embodiment. Figures 3 to 6 The method flow will not be repeated here.

[0149] It should be noted that the processor model in the embodiment of the present application is simplified. Figure 1 As shown in the figure, each processor has not only L1 cache, but also L2, L3 and even L4 caches and the processor's own private memory. In addition, the shared memory mounting method does not necessarily mean that all processes use the same physical memory pool. The shared memory can also be mounted on different processors in physical implementation to support shared use by all processors.

[0150] In addition, the technical solution of the present application is not limited to the multi-way processor chip design based on RiscV, but can also be applied to the following other technical fields, such as: High-Performance Computing (HPC), cloud computing and data centers, Artificial Intelligence (AI) and Machine Learning (ML), embedded systems and the Internet of Things (IoT), edge computing and other application scenarios.

[0151] Specifically, in the HPC field, efficient memory access and synchronization mechanisms of multi-processor systems are crucial. Through the distributed memory architecture and Clint synchronization sharing mechanism, the data access efficiency between computing nodes can be improved, latency can be reduced, and the performance of the overall system can be enhanced.

[0152] In cloud computing and data centers, efficient communication and data consistency management between processors are the key to ensuring service quality. By using the technical solution of this application, efficient communication and data consistency management of processor nodes within a data center can be achieved, thereby improving data processing efficiency and reliability.

[0153] In AI and ML systems, data transmission and synchronization between processors (including CPU, GPU, NPU, etc.) is the basis for efficient parallel computing. Through the distributed memory and Clint synchronization mechanism of this application, data interaction between multiple processors can be optimized, and the speed of model training and reasoning can be improved.

[0154] In embedded systems and IoT devices, processor resources are limited and efficient collaboration is required. By using the technical solution of this application, efficient data synchronization and inter-processor communication can be achieved in a resource-constrained environment, thereby improving the response speed and processing capability of the device.

[0155] Edge computing emphasizes local data processing to reduce latency. Through the technical solution of this application, efficient collaborative processing of multiple processors can be achieved in edge devices, enhancing the real-time and processing capabilities of edge computing.

[0156] Through the above-mentioned expansion and optimization, the technical solution of the present application can be applied to multiple technical fields, meet the needs of efficient data synchronization and processing in different scenarios, and provide a wider scope of protection.

[0157] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above-mentioned shared memory data access method embodiments when running.

[0158] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.

[0159] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any of the above-mentioned shared memory data access method embodiments are implemented.

[0160] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium, the non-volatile computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps in any of the above-mentioned shared memory data access method embodiments are implemented.

[0161] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0162] The above is a detailed introduction to a shared memory data access method provided by the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and its core idea of ​​the present application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.

Claims

1. A shared memory data access method, characterized in that: Applied to a first processor in a multi-processor system, the method comprises: When the first processor expects to access target data in the system shared memory resource, searching in the cache layer and local memory of the first processor whether the target data is stored; If there is no hit, an access request is sent to the first core local interrupter, wherein the access request carries address information; Searching, according to the address information, whether the target data is stored in a first shared memory of the first processor, where the first shared memory is a part of the system shared memory resources; If stored, the target data is obtained from the first shared memory, and the ownership of the target data is registered through the first core local interrupter, so that the target data can be directly accessed when it is subsequently read and written.

2. The method according to claim 1, characterized in that The searching whether the target data is stored in the cache layer and the local memory of the first processor includes: Generate memory addresses to be accessed according to instructions or program logic; Searching, according to the memory address, among the tags of each cache line in the cache layer to see whether there is a matching tag; If so, determining that the target data is hit in the cache layer; If not, it is determined that the target data is not hit in the cache layer, and whether the target data is stored in the local memory is searched according to the memory address.

3. The method according to claim 1, characterized in that The searching, according to the address information, in the first shared memory of the first processor to determine whether the target data is stored includes: According to the physical address indicated in the address information, searching in the first shared memory whether there is a data with an address that is the same as the physical address; If so, determining to store the target data in the first shared memory; If not, it is determined that the target data is not stored.

4. The method according to any one of claims 1 to 3, characterized in that: The method further comprises: If the target data is not stored in the first shared memory, sending the access request to the second processor in the multi-processor system through the first core local interrupter; receiving the target data searched by the second processor in its cache layer, or local memory, or second shared memory according to the access request, where the second shared memory is part of the system shared memory resources; The target data is stored in the local memory of the first processor, and the ownership of the target data is registered through the first core local interrupter.

5. The method according to claim 4, characterized in that The sending the access request to the second processor in the multi-processor system through the first core local interrupter includes: The access request is sent to the second core local interrupter of the second processor through the first core local interrupter, so that the second core local interrupter searches for the target data in its cache layer, local memory or second shared memory according to the address information carried in the access request.

6. The method according to claim 4, characterized in that The method further comprises: When receiving an interrupt request sent by the second processor through the second core local interrupter, checking the current usage of data in the data segment to be accessed by the second processor, where the data in the data segment is stored in the first shared memory; If the usage situation is that the first processor completes processing the data of the data segment, a response message is sent to the second core local interrupter through the first core local interrupter, and the response message is used to notify the second processor and authorize the second processor to access the data of the data segment.

7. A shared memory data access device, characterized in that: The device comprises: A first search module, configured to search whether the target data is stored in the cache layer and local memory of the first processor when the first processor expects to access the target data in the system shared memory resource; A sending module, used for sending an access request to the first core local interrupter in case of a miss, wherein the access request carries address information; A second search module, configured to search, according to the address information, whether the target data is stored in a first shared memory of the first processor, wherein the first shared memory is a part of the system shared memory resources; The read / write module is used to obtain the target data from the first shared memory when searching for the target data to be stored, and to register the ownership of the target data through the first core local interrupter so as to directly access the target data when reading and writing the target data subsequently.

8. A processor, characterized in that: comprising a kernel, a local memory and a kernel local interrupter, wherein the local memory and the kernel local interrupter are respectively connected to the kernel; The local memory stores computer instructions; The kernel executes the computer instruction through the core local interrupter to implement the steps of the shared memory data access method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the shared memory data access method according to any one of claims 1 to 6.

10. A multiprocessor system, characterized in that: The system comprises at least a first processor and a second processor, wherein the first processor and the second processor are connected via a high-speed interface; The first processor comprises: a first core, a cache layer, a local memory and a first core local interrupter, at least part of the memory in the local memory being a first shared memory; The second processor comprises: a second core, a cache layer, a local memory, and a second core local interrupter, wherein at least part of the local memory is a second shared memory; Wherein, the first processor uses the first core, the cache layer, the local memory and the first core local interrupter to perform the steps of the shared memory data access method according to any one of claims 1 to 6; When the first core fails to find the target data from the first shared memory, the second processor searches for the target data through the second core, the cache layer, the local memory and the second core local interrupter, and transmits the searched data to the first processor.