A data synchronization network-on-chip application method and device
By separating computing and storage units in a multi-core system, using custom instructions and multi-section design, the data synchronization problem is solved, data synchronization and parallel performance is improved, and the communication performance of on-chip networks is improved.
Patent Information
- Application Number
- CN202510645781.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2045-05-20
AI Technical Summary
In multi-core systems, the prior art has problems with local cores and on-chip core data synchronization, resulting in execution errors and degradation of parallel performance, especially when data is read or written before data is synchronized, data is lost or contaminated when multiple cores write data simultaneously.
By separating the calculation from the storage unit, using custom global barrier instructions and load reserve instructions, the execution control of the local core is realized; for on-chip cores, custom atomic exchange instructions are used to process data access between multiple cores; the storage unit adopts a multi-section design to realize parallel reading and writing; the RISC-V instruction set is extended to define local barriers, load reserves and global atomic exchange instructions to ensure data synchronization.
It is realized that without adding additional on-chip read and write instructions, ensuring the synchronization and parallel performance of data in multi-core systems is improved, and the communication performance and execution efficiency of on-chip networks are improved.
Smart Images

Figure CN120186110B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data synchronization, and in particular provides a network-on-chip application method and device for data synchronization. Background Art
[0002] As semiconductor processing advances into the post-Moore era, multi-core processor architectures have become a key path to breaking through the performance bottlenecks of single-core processors. This trend poses significant challenges to traditional bus interconnect architectures: communication based on shared buses suffers from inherent issues such as limited bandwidth, high arbitration latency, and exponential growth in power consumption with the number of cores. Against this backdrop, on-chip networks (NoCs), as a new interconnect paradigm, significantly improve the scalability and energy efficiency of multi-core systems by introducing packet switching and distributed routing mechanisms.
[0003] In a multi-core system based on an on-chip network, data synchronization mechanism is a key technical link to ensure the correctness and performance of the system. Storage Distributed storage is always accompanied by data synchronization issues.
[0004] Data synchronization issues between local cores and on-chip cores with local storage: If a local core requires the execution results of an on-chip core but reads the data before synchronization, the old value will be read; if a local core writes data but the on-chip core writes the new value before synchronization, the old value will be overwritten. Data synchronization issues between on-chip cores and the same storage: If multiple on-chip cores write data simultaneously, only the most recently written data will be retained in the storage, resulting in data loss; if a core is interrupted by other on-chip cores while continuously writing data, data corruption will occur. These issues will lead to execution errors and affect the parallel performance of the architecture. Summary of the Invention
[0005] The present invention aims to address the above-mentioned deficiencies in the prior art and provides a highly practical data synchronization network-on-chip application method.
[0006] A further technical task of the present invention is to provide a reasonably designed, safe and applicable data synchronization network-on-chip application device.
[0007] The technical solution adopted by the present invention to solve its technical problem is:
[0008] A data synchronization network-on-chip application method separates computation from storage for local cores, achieving unified management of local and on-chip memory access requests without adding additional on-chip read and write instructions. It also implements execution control of the local cores through custom global barrier and load-reserve instructions.
[0009] For on-chip cores, customized atomic exchange instructions are used to implement atomic request processing in the on-chip network, enabling orderly data access between multiple cores.
[0010] A multi-plate design is used for the storage unit to achieve parallel reading and writing under multiple requests.
[0011] Furthermore, the on-chip network interacts with the network interface unit, the encoding unit, the storage unit and the RV core.
[0012] The network interface unit is used for request management between the local core and the on-chip network;
[0013] The encoding unit is used to encode the local request to generate a memory access request and a packet chip memory access request;
[0014] The storage unit is used to store data and perform memory access operations from the local core and the on-chip network;
[0015] The RV core is the computing core of the RISC-V instruction set architecture.
[0016] Furthermore, the network interface unit includes a global address management logic, a flow control module, a decoding module and an atomic exchange module;
[0017] The global address management logic temporarily stores the on-chip memory access address to ensure that the memory access result is correctly returned to the on-chip core;
[0018] The flow control module monitors the on-chip network, implements memory access barriers for local cores, and avoids congestion of the on-chip network;
[0019] The decoding module decodes on-chip requests, converts memory access requests and generates core control signals;
[0020] The atomic exchange module processes memory access requests and synchronizes memory access requests of the atomic exchange class.
[0021] Furthermore, the storage unit includes a block arbitration logic and a multi-block SRAM, wherein the block arbitration logic arbitrates addresses of two types of memory access operations and maps them to corresponding multi-block SRAMs for storage;
[0022] The RV core includes pipeline control, data access and task execution logic. The pipeline control manages the execution progress of the RV core and synchronizes the load and reserve class memory access instructions.
[0023] Data access initiates memory access requests to local storage and on-chip storage of the on-chip core; task execution logic performs specific computing tasks on the accessed data.
[0024] Furthermore, when performing data synchronization control, the RISC-V instruction set is extended to define the local barrier instruction l_fence, the local load reserve instruction l_lr, and the global atomic swap instruction g_swap;
[0025] In the local barrier instruction l_fence, the mode field represents the barrier type, 00 corresponds to the barrier for local memory access operations, 01 corresponds to the barrier for on-chip memory access operations, and 11 corresponds to the barrier for all memory access operations, including local and on-chip memory access operations;
[0026] In the local load hold instruction l_lr, the rs1 field represents the register number of the target address, from which the memory access unit loads data; the rd field represents the register number of the destination for storing data; the aq and rl fields are mode selection fields. Setting aq to 1 indicates that the load hold is valid and blocks the pipeline of the local core; setting rl to 1 indicates that the load hold flag is cleared and the local core is released.
[0027] In the global atomic swap instruction g_swap, the rs1 field represents the index of the register where the on-chip target address is located, and the on-chip memory access request exchanges data from the target address; the rs2 field represents the index of the register where the data to be exchanged is located, and this data will exchange the original data in the on-chip target address; the rd field represents the index of the destination register where the original data is stored;
[0028] The aq and rl fields are mode selection fields. Setting aq to 1 means exchanging data with the storage unit in the on-chip core and obtaining the only exchange permission. Setting rl to 1 means releasing the exchange permission of the on-chip core.
[0029] Furthermore, the memory access requests of the encoding unit, storage unit and RV core realize data synchronization of the local core;
[0030] (1) For ordinary memory access instructions, the data access logic sends the memory access request to the encoding unit, which determines the type of memory access based on the target address of the memory access. For memory access requests stored locally, the encoding unit directly initiates read and write operations on the storage unit. For on-chip memory access requests, the encoding unit encodes the target address, generates the network coordinates of the on-chip core, and assembles the network coordinates, the encoded target address, the memory access type, and the memory access data into a network data packet and sends it to the on-chip network.
[0031] (2) For the barrier operation l_fence, according to the on-chip network status output by the flow control module, the flow control module will monitor the communication between the network interface unit and the on-chip network, and record the network status through the request counter. When a request is sent to the network, the counter is incremented by 1. When the request is correctly responded to by the network, the counter is decremented by 1. The pipeline control logic determines whether there is a request that has been sent to the on-chip network but has not received a response; if a request is not responded to, the pipeline of the RV core is blocked, waiting for the sent request to be processed;
[0032] (3) For the load-reservation operation l_lr, the operation is a local memory access operation. The pipeline control logic marks the target address of the request. If the aq field is set to 1, the memory access request of the RV core is blocked; the load-reservation operation with the rl field set to 1 is used to resume the execution of the pipeline as a data synchronization method for the local core; or the on-chip core accesses the marked target address and also resumes the execution of the pipeline as a data synchronization method for the on-chip core;
[0033] The blocking of the memory access request of the RV core is different from directly blocking the pipeline. If the instruction after the load and reserve operation is a non-memory access instruction, execution continues until a memory access instruction is encountered.
[0034] Furthermore, on-chip memory access requests issued by the local core are routed to a designated network address via the on-chip network, and the network interface unit manages the memory access requests to achieve data synchronization among the on-chip cores;
[0035] (1) The memory access request from the chip first enters the decoding module. The decoding module determines the memory access type based on the target address of the request. If it is a normal memory access request, the memory access request is sent to the atomic exchange module. If it is a control request, the corresponding control signal is generated to control the execution of the local core. Depending on the control request address, a core restart signal and an enable signal available to the local core are generated.
[0036] (2) The atomic swap module further processes the memory access request: For read and write requests, the atomic swap module directly sends such requests to the storage unit; for the atomic swap instruction g_swap, the atomic swap module will generate two consecutive read-first-then-write requests and assign a unique swap permission based on the target address and source network address of the swap; during the acquisition of the swap permission, other on-chip memory access requests will fail until the swap request with the same target address and source network address releases the permission;
[0037] (3) The global address management logic will monitor all requests that need to be returned, temporarily store the source network address of the request, and after the storage unit returns the data, package the source network address and data and then send it to the on-chip network.
[0038] Furthermore, the memory unit responds to memory access requests from the local core and the on-chip core in parallel;
[0039] The storage unit first performs round-robin block arbitration on the two types of requests; based on the block arbitration results, the requests are sent to the corresponding blocks, and the results of the memory access are also returned correctly based on the block arbitration results; if the two types of request addresses are located in different blocks, the read and write operations are performed simultaneously; if the two types of request addresses are located in the same block, the two types of requests are responded to alternately according to the polling strategy.
[0040] A data synchronization network-on-chip application device comprises: at least one memory and at least one processor;
[0041] The at least one memory is configured to store a machine-readable program;
[0042] The at least one processor is configured to call the machine-readable program to execute an on-chip network application method for data synchronization.
[0043] Compared with the prior art, the data synchronization network-on-chip application method and device of the present invention have the following outstanding beneficial effects:
[0044] The present invention provides an effective data synchronization function for the on-chip network through the implementation of atomic exchange, load reservation and global memory access barrier. At the same time, the multi-plate memory access unit design further improves the communication performance and parallel performance of the on-chip network. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0046] Figure 1 This is an architectural diagram of an on-chip network application method for data synchronization;
[0047] Figure 2 The present invention is a flowchart of a method for applying a network-on-chip (NOC) for data synchronization;
[0048] Figure 3 The present invention is a diagram of an on-chip network architecture in an on-chip network application method for data synchronization;
[0049] Figure 4 The present invention is a schematic diagram of an embodiment of a method for applying a network-on-chip (NoC) for data synchronization. DETAILED DESCRIPTION
[0050] In order to enable those skilled in the art to better understand the solutions of the present invention, the present invention will be further described in detail below in conjunction with specific embodiments. Obviously, the embodiments described are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.
[0051] A best embodiment is given below:
[0052] like Figure 1 and Figure 3 As shown, in this embodiment, a data synchronization network-on-chip application method is used. For the local core, by separating the computing and storage units, unified management of local memory access requests and on-chip memory access requests is achieved without adding additional on-chip read and write instructions. Execution control of the local core is achieved through customized global barrier instructions and load-reserve instructions.
[0053] For on-chip cores, customized atomic exchange instructions are used to implement atomic request processing in the on-chip network, enabling orderly data access between multiple cores.
[0054] A multi-plate design is adopted for the storage unit to realize parallel reading and writing under multiple requests, thereby improving the communication performance of the on-chip network.
[0055] Among them, the on-chip network interacts with the network interface unit, encoding unit, storage unit and RV core.
[0056] The network interface unit is used to manage requests between the local core and the on-chip network, including global address management logic, flow control module, decoding module and atomic exchange module;
[0057] The global address management logic temporarily stores the on-chip memory access address to ensure that the memory access result is correctly returned to the on-chip core;
[0058] The flow control module monitors the on-chip network, implements memory access barriers for local cores, and avoids congestion in the on-chip network;
[0059] The decoding module decodes on-chip requests, converts memory access requests and generates core control signals;
[0060] The atomic swap module processes memory access requests and synchronizes memory access requests of the atomic swap class.
[0061] The encoding unit is used to encode local requests, generate memory access requests and package on-chip memory access requests.
[0062] The storage unit is used to store data and execute memory access operations from the local core and on-chip network. It includes block arbitration logic and multi-block SRAM. The block arbitration logic arbitrates the addresses of the two types of memory access operations and maps them to the corresponding multi-block SRAM for storage.
[0063] The RV core is the computing core of the RISC-V instruction set architecture, including pipeline control, data access and task execution logic. The pipeline control manages the execution progress of the RV core and synchronizes the load-reserve memory access instructions.
[0064] Data access initiates memory access requests to local storage and on-chip storage of the on-chip core; task execution logic performs specific computing tasks on the accessed data.
[0065] To implement data synchronization control in on-chip networks, this application extends the RISC-V instruction set and defines the local barrier instruction (l_fence), the local load reserve instruction (l_lr), and the global atomic swap instruction (g_swap). The extended instruction format is as follows:
[0066] Table 1:
[0067]
[0068] As shown in Table 1, in the l_fence instruction, the mode field represents the barrier type. 00 corresponds to the barrier for local memory access operations, 01 corresponds to the barrier for on-chip memory access operations, and 11 corresponds to the barrier for all memory access operations (including local and on-chip memory access operations). This instruction is based on the original RISC-V barrier instruction (fence) extension. While implementing the original function (mode=00 barrier), it also provides the memory access barrier function in the on-chip network.
[0069] Table 2:
[0070]
[0071] As shown in Table 2, in the l_lr instruction: the rs1 field represents the label of the register where the target address is located, and the memory access unit loads data from this address; the rd field represents the label of the destination register where the data is stored, and the aq and rl fields are mode selection fields. Setting aq to 1 means that the load reserve is valid and blocks the pipeline of the local core; setting rl to 1 means clearing the load reserve flag and releasing the local core; this instruction is an implementation of the load reserve instruction (lr) in the RISC-V instruction set in the on-chip network architecture, which is used to achieve local data synchronization and execution pipeline control.
[0072] Table 3:
[0073]
[0074] As shown in Table 3, in the g_swap instruction, the rs1 field represents the index of the register where the on-chip target address is located, and the on-chip memory access request swaps data from this address; the rs2 field represents the index of the register where the data to be swapped is located, and this data will swap the original data at the on-chip target address; the rd field represents the index of the destination register storing the original data; the aq and rl fields are mode selection fields. In the present invention, setting aq to 1 means exchanging data with a storage unit in the on-chip core and obtaining unique swap permissions, and setting rl to 1 means releasing the swap permissions of the on-chip core. These three instructions are based on the atomic instruction extension, conform to the RISC-V instruction set specification, and improve the application of the instruction set in on-chip networks.
[0075] like Figure 2As shown, the encoding unit and storage unit manage the memory access requests of the RV core to achieve data synchronization of the local core:
[0076] (1) For ordinary memory access instructions, the data access logic sends the memory access request to the encoding unit, which determines the type of memory access based on the target address of the memory access. For memory access requests for local storage, the encoding unit will directly initiate read and write operations on the storage unit. For on-chip memory access requests, the encoding unit encodes the target address and generates the network coordinates of the on-chip core. It then assembles the network coordinates, the encoded target address, the memory access type (read, write, atomic exchange), and the memory access data into a network data packet and sends it to the on-chip network.
[0077] (2) For the barrier operation (l_fence), based on the on-chip network status output by the flow control module (the flow control module monitors the communication between the network interface unit and the on-chip network, and records the network status through the request counter. When a request is sent to the network, the counter increases by 1, and when the request is correctly responded to by the network, the counter decreases by 1), the pipeline control logic determines whether there is a request that has been sent to the on-chip network but has not received a response; if a request is not responded to, the pipeline of the RV core is blocked, waiting for the sent request to be processed.
[0078] (3) For the load-reservation operation (l_lr), which is a local memory access operation, the pipeline control logic marks the target address of the request. If the aq field is set to 1, the memory access request of the RV core is blocked (unlike directly blocking the pipeline, if the instruction after the load-reservation operation is a non-memory access instruction, execution continues until a memory access instruction is encountered); the load-reservation operation with the rl field set to 1 can resume the execution of the pipeline as a data synchronization method for the local core; or the on-chip core can access the marked target address to resume the execution of the pipeline as a data synchronization method for the on-chip core.
[0079] The on-chip memory access request issued by the local core is routed to the specified network address through the on-chip network. The network interface unit manages the memory access request and realizes data synchronization between the on-chip cores:
[0080] A. Memory access requests from the chip first enter the decoding module, which determines the memory access type based on the target address of the request. If it is a normal memory access request, the request is sent to the atomic swap module. If it is a control request, the corresponding control signal is generated to control the execution of the local core. Depending on the control request address, a core restart signal and an enable signal available to the local core can be generated.
[0081] B. The atomic swap module further processes memory access requests: For read and write requests, the atomic swap module directly sends such requests to the storage unit; for atomic swap instructions (g_swap), the atomic swap module will generate two consecutive read-first-then-write requests (reading out the original data and writing new data to complete the data swap), and allocate unique swap permissions based on the target address and source network address of the swap; during the period of obtaining swap permissions, other on-chip memory access requests will fail until a swap request with the same target address and source network address releases the permissions.
[0082] C. The global address management logic monitors all requests that need to be returned (read operations, atomic swap operations), temporarily stores the source network address of the request, and after the storage unit returns the data, packages the source network address and data and then sends it to the on-chip network.
[0083] In addition, the storage unit can respond to memory access requests from the local core and the on-chip core in parallel: the storage unit first performs polling block arbitration for the two types of requests (through address encoding, it ensures that continuous data is evenly distributed in each block, improving parallel access performance); according to the block arbitration result, the request is sent to the corresponding block, and the memory access result will also be correctly returned according to the block arbitration result; if the two types of request addresses are located in different blocks, read and write operations can be performed simultaneously. If the two types of request addresses are located in the same block, the two types of requests are responded to alternately according to the polling strategy; the multi-block storage structure and polling arbitration strategy ensure that any memory access request can be responded to within 2 clock cycles (when the memory access delay is 1 clock cycle).
[0084] Taking block matrix multiplication as an example, general matrix multiplication is widely used in parallel computing, especially in neural networks, where very large-scale matrix multiplications are common. In a multi-core architecture based on a network-on-chip (NoC), a matrix block approach is used to split large matrices into smaller matrices and distribute them across the computing cores. Local cores can quickly access the storage of other cores through the NoC, enabling data reuse and accelerating task execution.
[0085] For example, the on-chip core stores the first row of input matrix A and the second column of input matrix B. The local core can wait for the on-chip core to send the data to the local storage, and then perform matrix multiplication to obtain the calculation results of the first row and second column of the result matrix C.
[0086] Based on this scenario, the implementation method is as follows Figure 4 shown.
[0087] (1) The local core sets a memory access barrier through the g_fence instruction to ensure that all memory access requests before this instruction have been responded to, and then initiates a load-reserved memory access request to the l-lock local lock address through the l_lr instruction;
[0088] (2) The pipeline control logic will mark the l-lock address and wait for the on-chip core to access the address to complete the unlock, blocking the execution of the local core during this period;
[0089] (3) The on-chip core initiates an atomic swap request to the g-swap swap address through the g_swap instruction. If other on-chip cores have obtained the swap permission, it waits until the core obtains the swap permission.
[0090] (4) The on-chip core obtains the exchange permission to complete data synchronization, and can safely transmit data. The block data of the input matrix AB is written to the specified address through multiple general read and write instructions;
[0091] (5) After completing the data transfer, the on-chip core initiates an atomic swap request to the g-swap swap address through the g_swap instruction, releases the swap permission, and then initiates a read request to the l-lock address in step (2) to resume the execution of the local core;
[0092] (6) The local core has completely received the data from the on-chip core and performs matrix multiplication on the input matrix AB to complete the computation task. This method efficiently completes the block matrix multiplication, and the data transmission between the local core and the on-chip core is carried out in an orderly manner, ensuring data synchronization.
[0093] The above-mentioned specific implementation methods are only specific cases of the present invention. The patent protection scope of the present invention includes but is not limited to the above-mentioned specific implementation methods. Any technical solutions that conform to the above-mentioned specific implementation methods of the present invention and any appropriate changes or substitutions made thereto by ordinary technicians in the relevant technical field shall fall within the patent protection scope of the present invention.
[0094] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. A data synchronization network-on-chip application method, characterized in that: For local cores, by separating computation from storage, unified management of local and on-chip memory access requests is achieved without adding additional on-chip read and write instructions. Execution control of local cores is achieved through custom global barrier instructions and load-reserve instructions. For on-chip cores, customized atomic exchange instructions are used to implement atomic request processing in the on-chip network, enabling orderly data access between multiple cores. The storage unit adopts a multi-plate design to achieve parallel reading and writing under multiple requests; The on-chip network interacts with the network interface unit, encoding unit, storage unit and RV core, The network interface unit is used for request management between the local core and the on-chip network; The encoding unit is used to encode the local request to generate a memory access request and a packet-based on-chip memory access request; The storage unit is used to store data and perform memory access operations from the local core and the on-chip network; The RV core is a computing core of the RISC-V instruction set architecture; The network interface unit includes a global address management logic, a flow control module, a decoding module and an atomic exchange module; The global address management logic temporarily stores the on-chip memory access address to ensure that the memory access result is correctly returned to the on-chip core; The flow control module monitors the on-chip network, implements memory access barriers for local cores, and avoids congestion of the on-chip network; The decoding module decodes on-chip requests, converts memory access requests and generates core control signals; The atomic exchange module processes memory access requests and synchronizes memory access requests of the atomic exchange class; The storage unit includes a block arbitration logic and a multi-block SRAM. The block arbitration logic arbitrates addresses of two types of memory access operations and maps them to corresponding multi-block SRAMs for storage. The RV core includes pipeline control, data access and task execution logic. The pipeline control manages the execution progress of the RV core and synchronizes the load and reserve class memory access instructions. Data access is to initiate access requests to local storage and on-chip storage of on-chip cores; task execution logic performs specific computing tasks on the accessed data; When performing data synchronization control, the RISC-V instruction set is extended to define the local barrier instruction l_fence, the local load reserve instruction l_lr, and the global atomic swap instruction g_swap; In the local barrier instruction l_fence, the mode field represents the barrier type, 00 corresponds to the barrier for local memory access operations, 01 corresponds to the barrier for on-chip memory access operations, and 11 corresponds to the barrier for all memory access operations, including local and on-chip memory access operations; In the local load reserve instruction l_lr, the rs1 field represents the label of the register where the target address is located, and the memory access unit loads data from this address; The rd field represents the destination register number for storing data, and the aq and rl fields are mode selection fields. Setting aq to 1 means that the load reserve is valid and blocks the pipeline of the local core. Setting rl to 1 clears the load hold flag and releases the local core; In the global atomic swap instruction g_swap, the rs1 field represents the index of the register where the on-chip target address is located, and the on-chip memory access request exchanges data from the target address; the rs2 field represents the index of the register where the data to be exchanged is located, and the data will be exchanged with the original data in the on-chip target address; The rd field represents the destination register number where the original data is stored; The aq and rl fields are mode selection fields. Setting aq to 1 means exchanging data with the storage unit in the on-chip core and obtaining the only exchange permission. Setting rl to 1 means releasing the exchange permission of the on-chip core. The memory access request of the encoding unit, storage unit and RV core realizes data synchronization of the local core; (1) For ordinary memory access instructions, the data access logic sends the memory access request to the encoding unit, which determines the type of memory access based on the target address of the memory access. For memory access requests stored locally, the encoding unit directly initiates read and write operations on the storage unit. For on-chip memory access requests, the encoding unit encodes the target address, generates the network coordinates of the on-chip core, and assembles the network coordinates, the encoded target address, the memory access type, and the memory access data into a network data packet and sends it to the on-chip network. (2) For the barrier operation l_fence, according to the on-chip network status output by the flow control module, the flow control module will monitor the communication between the network interface unit and the on-chip network, and record the network status through the request counter. When a request is sent to the network, the counter is incremented by 1. When the request is correctly responded to by the network, the counter is decremented by 1. The pipeline control logic determines whether there is a request that has been sent to the on-chip network but has not received a response; if a request is not responded to, the pipeline of the RV core is blocked, waiting for the sent request to be processed; (3) For the load-reservation operation l_lr, the operation is a local memory access operation. The pipeline control logic marks the target address of the request. If the aq field is set to 1, the memory access request of the RV core is blocked; Resuming pipeline execution using a load-reserve operation with the rl field set to 1 as a data synchronization method for the local core; Or the on-chip core accesses the marked target address and also resumes the execution of the pipeline as a data synchronization method for the on-chip core; The blocking of the memory access request of the RV core is different from directly blocking the pipeline. If the instruction after the load and reserve operation is a non-memory access instruction, execution continues until a memory access instruction is encountered.
2. The data synchronization network-on-chip application method according to claim 1, characterized in that: On-chip memory access requests issued by the local core are routed to the specified network address through the on-chip network. The network interface unit manages the memory access requests and realizes data synchronization between the on-chip cores. (1) The memory access request from the chip first enters the decoding module. The decoding module determines the memory access type based on the target address of the request. If it is a normal memory access request, the memory access request is sent to the atomic exchange module. If it is a control request, a corresponding control signal is generated to control the execution of the local core; depending on the control request address, a core restart signal and an enable signal available to the local core are generated; (2) The atomic exchange module further processes the memory access request: For read and write requests, the atomic exchange module directly sends such requests to the storage unit; For the atomic swap instruction g_swap, the atomic swap module generates two consecutive read-then-write requests and allocates a unique swap permission based on the target address and source network address. While obtaining the swap permission, other on-chip memory access requests will fail until the swap request with the same target address and source network address releases the permission. (3) The global address management logic will monitor all requests that need to be returned, temporarily store the source network address of the request, and after the storage unit returns the data, package the source network address and data and then send it to the on-chip network.
3. The data synchronization network-on-chip application method according to claim 2, characterized in that: The memory unit responds to memory access requests from the local core and the on-chip core in parallel; The storage unit first performs round-robin arbitration on the two types of requests. Based on the arbitration result, the request is sent to the corresponding block. At the same time, the memory access result is also returned based on the block arbitration result. If the two types of request addresses are located in different blocks, the read and write operations are performed simultaneously. If the two types of request addresses are located in the same block, the two types of requests are responded to alternately according to the polling strategy.
4. A data synchronization network-on-chip application device, characterized in that: include: at least one memory and at least one processor; The at least one memory is configured to store a machine-readable program; The at least one processor is configured to call the machine-readable program to execute the method according to any one of claims 1 to 3.
Citation Information
Patent Citations
Instruction prefetch-based multi-core shared memory control equipment
CN102207916A
Weak coupling coprocessor design method based on RISC-V extension instruction
CN118819636A