A non-volatile storage access method, device, equipment and medium

CN122838320APending Publication Date: 2026-09-29CCORE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611202765.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-10
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

非易失性存储器(Non-Volatile Memory,NVM)因掉电不丢失,常用于存储关键数据,在实际应用中,受成本、面积或先进NVM工艺选型的限制,单个NVM的数据宽度常小于总线数据宽度;同时在多核架构中,多核对NVM访问常依赖共享总线通道而产生竞争,降低访问效率

Benefits of technology

[0014]由此可见,本申请技术方案应用于多核系统,所述多核系统包括若干处理器子系统,每个所述处理器子系统包含至少两个处理器核心以及与各所述处理器核心分别对应的非易失性存储单元,首先对目标处理器核心发起的目标访问请求进行解析,确定所述目标访问请求所指向的目标非易失性存储单元;然后在所述目标非易失性存储单元为目标处理器子系统对应的非易失性存储单元时,通过所述目标处理器核心的直连访问端口,直接发起对所述目标非易失性存储单元的访问操作,同时触发对所述目标处理器子系统内的剩余非易失性存储单元的并行访问操作;所述目标处理器子系统为所述目标处理器核心对应的子系统;再将所述目标非易失性存储单元输出的第一窄宽度数据与所述剩余非易失性存储单元输出的第二窄宽度数据进行合并,得到与总线数据宽度相匹配的合并数据;所述总线数据宽度为所述目标处理器子系统对应的所有非易失性存储单元的窄宽度数据的宽度之和;最后通过所述直连访问端口将所述合并数据返回至所述目标处理器核心。这样一来,本申请通过将多核系统中的处理器核心与非易失性存储单元对应划分为若干处理器子系统,在目标处理器核心直连访问本子系统内对应的非易失性存储单元时,同步触发对子系统内剩余非易失性存储单元的并行访问,并将各窄宽度数据合并为与总线数据宽度匹配的数据后返回。这样可将原本串行多次的窄宽度读取操作转化为单次并行合并操作,显著缩短了多核场景下窄宽度非易失性存储的直连访问时间,提升了访问效率;同时无需增加非易失性存储总容量或面积即可实现总线数据宽度的匹配,可以有效控制芯片成本。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122838320A_ABST
    Figure CN122838320A_ABST
Patent Text Reader

Abstract

This application discloses a method, apparatus, device, and medium for accessing non-volatile memory, relating to the field of memory access technology and applied to multi-core systems. The multi-core system includes several processor subsystems, each containing at least two processor cores and corresponding non-volatile memory units. The method includes: parsing a target access request initiated by a target processor core; when the parsed target non-volatile memory unit corresponds to a target processor subsystem, directly initiating an access operation while simultaneously triggering parallel access operations to the remaining non-volatile memory units; merging the output first narrow-width data with second narrow-width data, and returning merged data matching the bus data width. This transforms the original serial multiple narrow-width read operations into a single parallel merging operation, shortening the direct access time of narrow-width non-volatile memory in multi-core scenarios and improving access efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of memory access technology, and in particular to a non-volatile memory access method, apparatus, device, and medium. Background Technology

[0002] With the iteration of integrated circuit technology and the increasing demand for computing power, multi-core architecture has become the mainstream direction for high-performance processors and embedded chips in scenarios such as industrial control automation or smart IoT devices. Parallel processing capability is the key to breaking through the computing power bottleneck. Non-volatile memory (NVM) is often used to store critical data because it does not lose data when power is off. In practical applications, due to limitations of cost, area, or advanced NVM technology selection, the data width of a single NVM is often smaller than the bus data width. At the same time, in multi-core architectures, multiple cores often rely on shared bus channels to access NVM, resulting in contention and reduced access efficiency. When NVM becomes narrow-width (e.g., half the bus width), there are only two existing approaches: one is a serial approach, where each NVM access must be performed serially, consuming about twice the NVM read wait time, resulting in a significant decrease in access efficiency; the other is an independent parallel approach, which configures two narrow-width NVMs separately for each core to form a dedicated memory. Although the bit width is restored, the number of NVMs doubles, the area increases, the cost rises significantly, and selection becomes difficult.

[0003] Therefore, optimizing narrow-width NVM to achieve processor direct access efficiency is a problem that needs to be solved in this field. Summary of the Invention

[0004] In view of this, the purpose of this invention is to provide a method, apparatus, device, and medium for accessing non-volatile memory, which can transform multiple serial narrow-width read operations into a single parallel merge operation, thereby shortening the direct access time of narrow-width non-volatile memory in multi-core scenarios and improving access efficiency. The specific solution is as follows: In a first aspect, this application provides a non-volatile memory access method applied to a multi-core system, the multi-core system comprising a plurality of processor subsystems, each processor subsystem comprising at least two processor cores and non-volatile memory units corresponding to each processor core, the method comprising: The target access request initiated by the target processor core is parsed to determine the target non-volatile memory unit to which the target access request points; When the target non-volatile memory unit is a non-volatile memory unit corresponding to the target processor subsystem, an access operation to the target non-volatile memory unit is directly initiated through the direct access port of the target processor core, and at the same time, parallel access operations to the remaining non-volatile memory units in the target processor subsystem are triggered; the target processor subsystem is the subsystem corresponding to the target processor core. The first narrow-width data output by the target non-volatile memory cell is merged with the second narrow-width data output by the remaining non-volatile memory cells to obtain merged data that matches the bus data width; the bus data width is the sum of the widths of the narrow-width data of all non-volatile memory cells corresponding to the target processor subsystem. The merged data is returned to the target processor core via the direct access port.

[0005] Optionally, the method further includes: When the target non-volatile memory unit is not the non-volatile memory unit corresponding to the target processor subsystem, the target access request is routed to the shared access port of the processor subsystem to which the target non-volatile memory unit belongs via a bus matrix; the bus matrix is ​​connected to each bus access port of the plurality of processor subsystems and is used to route the target access request to the corresponding non-volatile memory unit. The system initiates an access operation to the target non-volatile memory unit through the shared access port, and returns the accessed data to the target processor core via the bus matrix.

[0006] Optionally, the method further includes: When the target non-volatile storage unit simultaneously receives multiple access requests initiated via the direct access port and the shared access port, the response order of the multiple access requests is determined by a round-robin arbitration method. The access operations corresponding to the multiple access requests are completed sequentially according to the response order.

[0007] Optionally, before initiating the access operation to the target non-volatile memory unit, the method further includes: Detect whether the target access request is hit in the cache of the target processor core; If a hit occurs, the target data is read directly from the cache of the target processor core. The target data is returned to the target processor core, and the access operation to the target non-volatile memory unit is cancelled; If the target is not found, the step of initiating an access operation to the target non-volatile memory cell continues.

[0008] Optionally, initiating the access operation to the target non-volatile memory unit includes: Detect whether the target access request is hit in the storage control cache corresponding to the target non-volatile storage unit; If a hit occurs, the first narrow-width data is read from the storage control cache; If a miss occurs, the first narrow-width data is read from the target non-volatile memory cell after one read wait time cycle.

[0009] Optionally, triggering parallel access operations to the remaining non-volatile memory units within the target processor subsystem includes: Based on the access address corresponding to the target access request, determine the target access address corresponding to each of the remaining non-volatile memory units in the target processor subsystem; The target access addresses and access operations for the target non-volatile storage units are sent out simultaneously, so that the target non-volatile storage units and the remaining non-volatile storage units can complete the access operations in parallel.

[0010] Optionally, merging the first narrow-width data output by the target non-volatile memory cell with the second narrow-width data output by the remaining non-volatile memory cells to obtain merged data matching the bus data width includes: According to the preset data position correspondence between each non-volatile memory unit and the bus data width in the target processor subsystem, the data positions of the first narrow width data and the second narrow width data are determined; The first narrow-width data and the second narrow-width data are merged according to the data position to obtain merged data that matches the bus data width.

[0011] Secondly, this application provides a non-volatile memory access device applied to a multi-core system, the multi-core system including a plurality of processor subsystems, each processor subsystem including at least two processor cores and non-volatile memory units corresponding to each processor core, the device including: The access request parsing module is used to parse the target access request initiated by the target processor core and determine the target non-volatile memory unit pointed to by the target access request. The access module is used to, when the target non-volatile memory unit is a non-volatile memory unit corresponding to the target processor subsystem, directly initiate an access operation to the target non-volatile memory unit through the direct access port of the target processor core, and simultaneously trigger parallel access operations to the remaining non-volatile memory units within the target processor subsystem; the target processor subsystem is the subsystem corresponding to the target processor core. The data merging module is used to merge the first narrow-width data output by the target non-volatile memory cell with the second narrow-width data output by the remaining non-volatile memory cells to obtain merged data that matches the bus data width; the bus data width is the sum of the widths of the narrow-width data of all non-volatile memory cells corresponding to the target processor subsystem. The data return module is used to return the merged data to the target processor core through the direct access port.

[0012] Thirdly, this application provides an electronic device, comprising: Memory, used to store computer programs; A processor for executing the computer program to implement the non-volatile memory access method as described above.

[0013] Fourthly, this application provides a computer-readable storage medium for storing a computer program, which, when executed by a processor, implements the non-volatile memory access method described above.

[0014] Therefore, the technical solution of this application is applied to a multi-core system, which includes several processor subsystems. Each processor subsystem includes at least two processor cores and non-volatile memory units corresponding to each processor core. First, the target access request initiated by the target processor core is parsed to determine the target non-volatile memory unit pointed to by the target access request. Then, when the target non-volatile memory unit is the non-volatile memory unit corresponding to the target processor subsystem, an access operation to the target non-volatile memory unit is directly initiated through the direct access port of the target processor core, while simultaneously triggering parallel access operations to the remaining non-volatile memory units within the target processor subsystem. The target processor subsystem is the subsystem corresponding to the target processor core. Next, the first narrow-width data output by the target non-volatile memory unit is merged with the second narrow-width data output by the remaining non-volatile memory units to obtain merged data that matches the bus data width. The bus data width is the sum of the widths of the narrow-width data of all non-volatile memory units corresponding to the target processor subsystem. Finally, the merged data is returned to the target processor core through the direct access port. In this way, this application divides the processor cores in a multi-core system into several processor subsystems corresponding to non-volatile memory units. When the target processor core directly accesses the corresponding non-volatile memory unit within its subsystem, it simultaneously triggers parallel access to the remaining non-volatile memory units within the subsystem, and merges the narrow-width data into data matching the bus data width before returning it. This transforms the originally serial, multiple narrow-width read operations into a single parallel merging operation, significantly shortening the direct access time of narrow-width non-volatile memory in multi-core scenarios and improving access efficiency. At the same time, it achieves bus data width matching without increasing the total capacity or area of ​​non-volatile memory, effectively controlling chip costs. Attached Figure Description

[0015] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0016] Figure 1 This application discloses a multi-core system architecture diagram; Figure 2 This is a flowchart of a non-volatile memory access method disclosed in this application; Figure 3 This application discloses a flowchart for processing multiple concurrent access requests. Figure 4 This is a schematic diagram of the structure of a non-volatile memory access device disclosed in this application; Figure 5 This is a structural diagram of an electronic device disclosed in this application. Detailed Implementation

[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0018] First, the architecture of the multi-core system involved in the technical solution of this application is introduced, such as... Figure 1 As shown, the multi-core system includes two processor subsystems: Processor subsystem 1 consists of CPU0 (Central Processing Unit), CPU1, and NVM0 and NVM1 (NVM combination 1); Processor subsystem 2 consists of CPU2, CPU3, and NVM2 and NVM3 (NVM combination 2). Each processor core is configured with two access ports: P0 port is a bus access port, connected to each NVM and peripheral via a bus matrix; P1 port (CPU0, CPU2) and P2 port (CPU1, CPU3) are direct access ports, directly connected to the corresponding NVM within their respective processor subsystems. Taking processor subsystem 1 as an example, CPU0's P1 port is directly connected to NVM0, and CPU1's P2 port is directly connected to NVM1. When CPU0 accesses NVM0 through P1 port, it triggers a parallel read operation on NVM1. The first narrow-width data output by NVM0 and the second narrow-width data output by NVM1 are combined to form merged data matching the bus data width (i.e., the data width of P0), which is then directly returned to CPU0 through P1 port. When CPU1 accesses NVM1 through port P2, it also triggers a parallel read operation on NVM0, achieving the same efficient direct access. When CPU0 needs to access NVM2, since the target NVM does not belong to this subsystem, CPU0 sends the access request to the bus matrix through port P0. The bus matrix routes the request to the shared access port of subsystem 2 based on the target address, and NVM2 returns narrow-width data to CPU0. When multiple CPUs access the same NVM simultaneously, such as CPU0 and CPU1 accessing NVM0 at the same time, the response order is determined by a round-robin method using the bus matrix or an arbitrator within the subsystem, and the access operations are completed sequentially. In this embodiment, the data width of each narrow-width NVM is half the bus data width. The two NVMs are combined to precisely match the bus width, thus achieving efficient direct access to narrow-width NVMs in multi-core scenarios without increasing the NVM area.

[0019] The following examples will provide a detailed description of the process of implementing non-volatile memory access in a multi-core system. See also... Figure 2 As shown, this invention discloses a non-volatile memory access method applied to a multi-core system. The multi-core system includes several processor subsystems, each of which includes at least two processor cores and non-volatile memory units corresponding to each processor core. The method specifically includes: Step S11: Parse the target access request initiated by the target processor core to determine the target non-volatile memory unit to which the target access request points.

[0020] In this embodiment, the access request initiated by the target processor core can be parsed, that is, the address information related to the non-volatile memory unit contained in the access request can be extracted. Specifically, according to a pre-set address mapping table, the parsed address information is compared one by one with the address ranges of each non-volatile memory unit corresponding to each processor subsystem in the multi-core system. When the parsed address information falls into the address range of a certain non-volatile memory unit, that non-volatile memory unit can be identified as the target non-volatile memory unit pointed to by this access request.

[0021] It should be noted that the address mapping table can be flexibly configured according to actual needs. For example, a contiguous address mapping method can be used, which divides the entire physical address space into several contiguous address ranges, with each processor subsystem corresponding to one of these ranges. Alternatively, an interleaved mapping method can be used to better balance the access load among multiple processor subsystems. Regardless of the address mapping rule used, the purpose is to uniquely determine the target non-volatile memory unit pointed to by the target access request and the processor subsystem to which it belongs.

[0022] Step S12: When the target non-volatile memory unit is a non-volatile memory unit corresponding to the target processor subsystem, an access operation to the target non-volatile memory unit is directly initiated through the direct access port of the target processor core, and at the same time, parallel access operations to the remaining non-volatile memory units in the target processor subsystem are triggered; the target processor subsystem is the subsystem corresponding to the target processor core.

[0023] In this embodiment, after parsing the target non-volatile memory unit corresponding to the access request, it can be determined whether the target non-volatile memory unit is included in the target processor subsystem where the target processor core is located. If the target non-volatile memory unit belongs to the target processor subsystem, a direct access path can be enabled. The corresponding direct access port is independent of the bus matrix and is dedicated to accessing non-volatile memory units within the target processor subsystem. It does not require bus arbitration and can directly initiate read commands to the target non-volatile memory unit. Simultaneously, parallel read operations on other non-volatile memory units within the same subsystem (i.e., the target processor subsystem) are triggered. In a specific embodiment, when triggering parallel access operations on the remaining non-volatile memory units within the target processor subsystem, the target access addresses corresponding to the remaining non-volatile memory units within the target processor subsystem can be determined based on the access address corresponding to the target access request, through address decoding or address offset calculation. The target access addresses and the access operations on the target non-volatile memory units are then synchronously issued, enabling the target non-volatile memory unit and the remaining non-volatile memory units to complete their access operations in parallel within the same access time period. This parallel access operation enables the simultaneous acquisition of data from multiple narrow-width storage units in a single direct access, facilitating subsequent data merging.

[0024] In one specific embodiment, before initiating an access operation to the target non-volatile memory unit, the process may further include: detecting whether the target access request hits the cache of the target processor core; if a hit occurs, directly reading the hit target data from the cache of the target processor core; returning the target data to the target processor core and canceling the access operation to the target non-volatile memory unit; if a miss occurs, continuing to execute the step of initiating the access operation to the target non-volatile memory unit. It is understood that before initiating the access operation to the target non-volatile memory unit, a cache hit check can be performed on the target access request. Specifically, the access address carried by the target access request is compared with the tag entry stored in the cache of the target processor core to determine whether the access address has already cached valid data in the cache. If the comparison matches, i.e., a cache hit occurs, the target data corresponding to the access address can be directly read from the cache, and the target data can be returned to the target processor core through a direct access port. Simultaneously, a cancellation signal is generated and sent to the access control interface of the target non-volatile memory cell. This signal blocks the issuance of read commands to the target non-volatile memory cell, thereby avoiding invalid access to the storage array, saving power consumption, releasing access channel resources, and terminating the subsequent processing flow of the current access request. Conversely, if the comparison is inconsistent (i.e., a cache miss), there is no valid data copy corresponding to the access address in the cache. In this case, the steps to initiate the access operation to the target non-volatile memory cell continue, i.e., an access command is issued to the target non-volatile memory cell via the direct access port.

[0025] In another specific embodiment, when initiating an access operation to the target non-volatile memory cell, it is checked whether the access request hits the storage control cache corresponding to the target non-volatile memory cell. This storage control cache is located inside the non-volatile memory controller and is used to cache recent data to reduce the number of direct accesses to the non-volatile memory array. Specifically, the data tag carried by the access request is compared with the tags of each cache entry stored in the storage control cache. If a matching tag exists and the cache entry is valid, it is considered a hit; otherwise, it is considered a miss. Further, if a hit occurs, the first narrow-width data is read directly from the corresponding entry in the storage control cache. This reading process only requires one clock cycle and does not require access to the non-volatile memory array, thus significantly shortening the data read time. If a miss occurs, a read command is issued to the storage array of the target non-volatile memory cell, and after a read wait time cycle, the first narrow-width data is read from the storage array. The read wait time period refers to the number of clock cycles required from issuing a read command to the memory array outputting valid data on the data bus. Its specific duration is determined by the process characteristics of the target non-volatile memory cell and the current clock frequency. After the data is read, it will be returned to the target processor core, and can also be written to the memory control cache of the target non-volatile memory cell for use in subsequent access requests.

[0026] In another specific embodiment, when the target non-volatile memory unit does not belong to the processor subsystem corresponding to the target processor core, the access request can be sent to the bus matrix through the bus access port of the target processor core. The bus matrix is ​​a shared interconnect module in a multi-core system that connects the bus access ports of each processor subsystem to each non-volatile memory unit. It maintains an internal address routing table that records the mapping relationship between the address ranges of each non-volatile memory unit and the shared access port of its respective processor subsystem. After receiving the target access request, the bus matrix extracts the target address carried within it and compares it with the address ranges in the routing table to determine the target processor subsystem to which the target non-volatile memory unit belongs. Then, it routes the access request to the shared access port of that target processor subsystem. It should be noted that the shared access port is an interface in the processor subsystem used to receive cross-subsystem access requests from the bus matrix and is independent of the directly connected access port. After the access request enters the target processor subsystem via the shared access port, the memory controller within the target processor subsystem initiates an access operation on the target non-volatile memory unit to read the corresponding narrow-width data. After the data is read, the narrow-width data is returned to the bus matrix via the shared access port. The bus matrix then, based on the source information of the access request, returns the data to the target processor core that initiated the access via the bus access port of the source processor subsystem, completing the cross-subsystem access process. It should be noted that cross-subsystem access requires address resolution, routing, and arbitration processing via the bus matrix, and only returns a single narrow-width non-volatile memory cell, without triggering parallel access operations. Therefore, the access latency is longer than direct access within a processor subsystem. This path is primarily used to handle scenarios where the target processor core accesses a non-volatile memory cell outside its own subsystem, ensuring that each processor core in a multi-core system can access all non-volatile memory spaces.

[0027] Accordingly, when multiple access requests are received simultaneously via direct access ports and shared access ports, the access arbitrator within the target processor subsystem determines the response order of the multiple access requests according to a preset polling arbitration method. Polling arbitration refers to the arbitrator maintaining a polling pointer that cycles through the request source ports, sequentially indicating that the request source corresponding to each port enjoys the highest priority in the current arbitration cycle. In the case of multiple concurrent access requests, the arbitrator uses the current position of the polling pointer as a reference, combined with the valid enable signals of each request source, to arbitrate from all pending access requests the request source that currently has access privileges. After arbitration, the access request of this request source is sent to the access control interface of the target non-volatile memory unit to perform the corresponding read or write operation; when the access operation is completed, the polling pointer moves to the next request source port, and the arbitrator continues to arbitrate the remaining access requests in the next arbitration cycle based on the updated polling pointer. This cycle repeats until all pending access requests have been processed sequentially. By using a round-robin arbitration method, multiple access request sources share the access bandwidth of the same target non-volatile storage unit, and the priority of each request source to obtain access opportunities rotates over time, preventing any request source from occupying the target non-volatile storage unit for a long time.

[0028] Step S13: Merge the first narrow-width data output by the target non-volatile memory cell with the second narrow-width data output by the remaining non-volatile memory cells to obtain merged data that matches the bus data width; the bus data width is the sum of the widths of the narrow-width data of all non-volatile memory cells corresponding to the target processor subsystem.

[0029] In this embodiment, after obtaining the first narrow-width data output by the target non-volatile memory cell and the second narrow-width data output by the remaining non-volatile memory cells within the same processor subsystem, the aforementioned narrow-width data can be merged. Specifically, the data position of each narrow-width data is determined according to the preset data position correspondence of each non-volatile memory cell within the processor subsystem, thus determining the data order; then, the narrow-width data are sequentially merged to form merged data with a width equal to the bus data width. The bus data width is equal to the sum of the widths of the narrow-width data of all non-volatile memory cells within the processor subsystem; that is, the cumulative bit width of each narrow-width data is equal to the physical data bit width of the bus. This ensures that the merged data can be completely received by the target processor core within one bus transmission cycle, avoiding multiple transmissions or bit width waste due to data bit width mismatch.

[0030] Step S14: Return the merged data to the target processor core through the direct access port.

[0031] Furthermore, after merging the first and second narrow-width data, the merged data is loaded onto the data output channel of the direct access port of the target processor core. The merged data is then directly returned to the target processor core via this direct access port, without traversing the bus matrix. This direct access port is a dedicated physical channel for data transmission between the target processor core and the corresponding non-volatile memory units within the processor subsystem. Its data width is equal to the width of the merged data, and it does not participate in the arbitration and routing process of the bus matrix. After receiving the merged data, the direct access port sends the merged data to the internal data bus of the target processor core with a full width matching the bus data width, completing a full parallel merge and read operation of narrow-width data. This process consumes only one read wait cycle and a small amount of logical concatenation delay, thus ensuring that the target processor core can fully receive the merged data matching the bus data width within one bus transmission cycle.

[0032] Therefore, this application divides the processor cores and non-volatile memory units in a multi-core system into several processor subsystems. When the target processor core directly accesses the corresponding non-volatile memory unit within its subsystem, it simultaneously triggers parallel access to the remaining non-volatile memory units within the subsystem, and merges the narrow-width data into data matching the bus data width before returning it. This transforms the originally serial, multiple narrow-width read operations into a single parallel merging operation, significantly shortening the direct access time of narrow-width non-volatile memory in multi-core scenarios and improving access efficiency. Furthermore, combined with a polling arbitration mechanism, it can achieve arbitration of concurrent access to the same NVM address by multiple processor cores, avoiding access conflicts. At the same time, it can achieve bus data width matching without increasing the total capacity or area of ​​non-volatile memory, effectively controlling chip costs.

[0033] like Figure 3 As shown, this embodiment discloses a flowchart for processing multiple concurrent access requests; specifically, it combines... Figure 1 The multi-core system shown illustrates the process by which the NVM controller handles multiple concurrent access requests through a polling arbitration mechanism when multiple processor cores simultaneously access the same NVM address, using NVM0 as the target memory access unit. This includes the following:

[0034] At a certain initial moment, CPU0 and CPU1 initiate access to the NVM0 address through ports P1 and P2 respectively, while CPU2 and CPU3 also initiate access to the NVM0 address through their respective ports P0 via the bus matrix. NVM0 internally has an access arbitrator that maintains a polling pointer, which cycles through the request source ports sequentially to determine the request source currently granted access. When multiple access requests arrive at NVM0 simultaneously, the valid enable signals of each request source are latched into the arbitrator's request queue. Upon detecting multiple pending requests, the arbitrator arbitrates based on the current position of the polling pointer.

[0035] Furthermore, if the polling pointer currently points to port P1, the arbitrator prioritizes responding to access requests initiated by CPU0 via port P1. NVM0 begins processing access requests from CPU0 via P1, pausing responses to requests from other ports. Once CPU0's access operation is complete, the polling pointer moves to port P2. If the arbitrator detects pending access requests on port P2 (i.e., access requests initiated by CPU1 via port P2), it begins processing access requests from CPU1 via P2. Once this access is complete, the polling pointer moves to port P0. If the arbitrator detects access requests from CPU2 / 3 on port P0, it begins processing these requests. After all pending access requests have been processed sequentially, the arbitrator completes this round of polling arbitration.

[0036] It should be noted that when CPU0 accesses NVM0 through port P1, it simultaneously initiates a parallel access operation to NVM1 within the same subsystem. The first narrow-width data output by NVM0 and the second narrow-width data output by NVM1 are concatenated after bit width adjustment and returned to CPU0 through CPU0's direct-connect port P1. This allows CPU0 to obtain complete data matching the bus width with only one NVM read wait time. When CPU1 accesses NVM0 through port P2, it similarly triggers a parallel access to NVM1 and returns the concatenated data through CPU1's direct-connect port P2. However, for cross-subsystem access initiated by CPU2 / 3 through port P0 and the bus matrix, since the access request originates from outside the processor subsystem, NVM0 only returns a single narrow-width data and does not trigger a parallel concatenation operation with NVM1. This narrow-width data is returned to CPU2 or CPU3 through the bus matrix, resulting in a longer access latency compared to direct-connect access. Multiple cross-subsystem access requests also require corresponding arbitration processing within the bus matrix.

[0037] Therefore, in this embodiment, the polling arbitration mechanism can be used to arbitrate concurrent access to the same NVM address by multiple processor cores, thus avoiding access conflicts. At the same time, the parallel merging mechanism within the processor subsystem can ensure high efficiency in direct access scenarios and take into account the feasibility of cross-subsystem access, thereby achieving a balance between narrow-width NVM access efficiency and overall system performance in multi-core scenarios.

[0038] like Figure 4 As shown, this embodiment discloses a non-volatile memory access device applied to a multi-core system. The multi-core system includes several processor subsystems, each of which includes at least two processor cores and non-volatile memory units corresponding to each processor core. The device includes: The access request parsing module 11 is used to parse the target access request initiated by the target processor core and determine the target non-volatile memory unit pointed to by the target access request. The access module 12 is used to, when the target non-volatile memory unit is a non-volatile memory unit corresponding to the target processor subsystem, directly initiate an access operation to the target non-volatile memory unit through the direct access port of the target processor core, and simultaneously trigger parallel access operations to the remaining non-volatile memory units in the target processor subsystem; the target processor subsystem is the subsystem corresponding to the target processor core. The data merging module 13 is used to merge the first narrow-width data output by the target non-volatile memory cell with the second narrow-width data output by the remaining non-volatile memory cells to obtain merged data that matches the bus data width; the bus data width is the sum of the widths of the narrow-width data of all non-volatile memory cells corresponding to the target processor subsystem. The data return module 14 is used to return the merged data to the target processor core through the direct access port.

[0039] As can be seen, this device can locate the target storage unit through the access request parsing module. When the access module performs direct access, it simultaneously triggers the parallel reading of the remaining storage units in the same subsystem. The data merging module combines multiple narrow-width data into data that matches the bus width and returns it through the direct connection port. This transforms multiple serial reads into a single parallel merging operation, shortening the direct access time of narrow-width non-volatile storage in multi-core scenarios and improving access efficiency. At the same time, it can achieve bus width matching without increasing the total area of ​​the storage unit, effectively controlling chip costs.

[0040] In one specific embodiment, the device may further include: An access request routing module is used to route the target access request to the shared access port of the processor subsystem to which the target non-volatile memory unit belongs via a bus matrix when the target non-volatile memory unit is not a non-volatile memory unit corresponding to the target processor subsystem; the bus matrix is ​​connected to each bus access port of the plurality of processor subsystems and is used to route the target access request to the corresponding non-volatile memory unit. The first access operation initiation module is used to initiate an access operation to the target non-volatile memory unit through the shared access port, and return the accessed data to the target processor core via the bus matrix.

[0041] In another specific embodiment, the device may further include: An arbitration module is used to determine the response order of the multiple access requests in a round-robin arbitration manner when the target non-volatile storage unit simultaneously receives multiple access requests initiated via the direct access port and the shared access port. The second access operation initiation module is used to sequentially complete the access operations corresponding to the multiple access requests according to the response order.

[0042] In one specific embodiment, the device may further include: The first cache detection module is used to detect whether the target access request hits the cache of the target processor core; The first data reading module is used to directly read the target data that was hit from the cache area of ​​the target processor core when a cache hit occurs. The access operation cancellation module is used to return the target data to the target processor core and cancel the access operation to the target non-volatile memory unit; An execution module is used to continue executing the step of initiating an access operation to the target non-volatile storage unit when a cache miss occurs.

[0043] In one specific embodiment, the access module 12 may include: The second cache detection unit is used to detect whether the target access request is hit in the storage control cache corresponding to the target non-volatile storage unit; The second data reading unit is used to read the first narrow-width data from the storage control cache when a cache hit occurs; The third data reading unit is used to read the first narrow-width data from the target non-volatile storage unit after a read wait time cycle when the cache misses.

[0044] In one specific embodiment, the access module 12 may include: An access address determination unit is used to determine, based on the access address corresponding to the target access request, each target access address corresponding to the remaining non-volatile memory unit in the target processor subsystem. The access operation issuing unit is used to synchronously issue each of the target access addresses and the access operations to the target non-volatile storage units, so that the target non-volatile storage units and the remaining non-volatile storage units can complete the access operations in parallel.

[0045] In one specific embodiment, the data merging module 13 may include: The data location determination unit is used to determine the data locations of the first narrow width data and the second narrow width data according to a preset data location correspondence between each non-volatile memory unit in the target processor subsystem and the bus data width; The data merging unit is used to merge the first narrow-width data and the second narrow-width data according to the data position to obtain merged data that matches the bus data width.

[0046] Furthermore, embodiments of this application also disclose an electronic device, Figure 5 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content of the diagram should not be construed as limiting the scope of this application.

[0047] Figure 5 This is a schematic diagram of the structure of an electronic device 20 provided in an embodiment of this application. Specifically, the electronic device 20 may include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 stores a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the non-volatile memory access method disclosed in any of the foregoing embodiments. Furthermore, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0048] In this embodiment, the power supply 23 is used to provide operating voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows can be any communication protocol applicable to the technical solution of this application, and is not specifically limited here; the input / output interface 25 is used to acquire external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs, and is not specifically limited here.

[0049] In addition, the memory 22, as a carrier for resource storage, can be a read-only memory, random access memory, disk or optical disk, etc. The resources stored thereon can include operating system 221, computer program 222, etc., and the storage method can be temporary storage or permanent storage.

[0050] The operating system 221 is used to manage and control the various hardware devices on the electronic device 20 and the computer program 222, which may be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program capable of performing the non-volatile memory access method executed by the electronic device 20 as disclosed in any of the foregoing embodiments, the computer program 222 may further include a computer program capable of performing other specific tasks.

[0051] Furthermore, this application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, it implements the aforementioned non-volatile memory access method. Specific steps of this method can be found in the corresponding content disclosed in the foregoing embodiments, and will not be repeated here.

[0052] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.

[0053] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0054] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.

[0055] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0056] The technical solutions provided in this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only intended to help understand the methods and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A non-volatile memory access method, characterized in that, Applied to a multi-core system, the multi-core system comprising a plurality of processor subsystems, each processor subsystem comprising at least two processor cores and non-volatile memory units corresponding to each processor core, the method comprising: The target access request initiated by the target processor core is parsed to determine the target non-volatile memory unit to which the target access request points; When the target non-volatile memory unit is a non-volatile memory unit corresponding to the target processor subsystem, an access operation to the target non-volatile memory unit is directly initiated through the direct access port of the target processor core, and at the same time, parallel access operations to the remaining non-volatile memory units in the target processor subsystem are triggered; the target processor subsystem is the subsystem corresponding to the target processor core. The first narrow-width data output by the target non-volatile memory cell is merged with the second narrow-width data output by the remaining non-volatile memory cells to obtain merged data that matches the bus data width; the bus data width is the sum of the widths of the narrow-width data of all non-volatile memory cells corresponding to the target processor subsystem. The merged data is returned to the target processor core via the direct access port.

2. The non-volatile memory access method according to claim 1, characterized in that, Also includes: When the target non-volatile memory unit is not the non-volatile memory unit corresponding to the target processor subsystem, the target access request is routed to the shared access port of the processor subsystem to which the target non-volatile memory unit belongs via the bus matrix. The bus matrix is ​​connected to each bus access port of the plurality of processor subsystems and is used to route target access requests to the corresponding non-volatile memory units. The system initiates an access operation to the target non-volatile memory unit through the shared access port, and returns the accessed data to the target processor core via the bus matrix.

3. The non-volatile memory access method according to claim 2, characterized in that, Also includes: When the target non-volatile storage unit simultaneously receives multiple access requests initiated via the direct access port and the shared access port, the response order of the multiple access requests is determined by a round-robin arbitration method. The access operations corresponding to the multiple access requests are completed sequentially according to the response order.

4. The non-volatile memory access method according to claim 1, characterized in that, Before initiating the access operation to the target non-volatile memory unit, the method further includes: Detect whether the target access request is hit in the cache of the target processor core; If a hit occurs, the target data is read directly from the cache of the target processor core. The target data is returned to the target processor core, and the access operation to the target non-volatile memory unit is cancelled; If the target is not found, the step of initiating an access operation to the target non-volatile memory cell continues.

5. The non-volatile memory access method according to claim 1, characterized in that, The initiation of the access operation to the target non-volatile memory unit includes: Detect whether the target access request is hit in the storage control cache corresponding to the target non-volatile storage unit; If a hit occurs, the first narrow-width data is read from the storage control cache; If a miss occurs, the first narrow-width data is read from the target non-volatile memory cell after one read wait time cycle.

6. The non-volatile memory access method according to claim 1, characterized in that, The triggering of parallel access operations to the remaining non-volatile memory units within the target processor subsystem includes: Based on the access address corresponding to the target access request, determine the target access address corresponding to each of the remaining non-volatile memory units in the target processor subsystem; The target access addresses and access operations for the target non-volatile storage units are sent out simultaneously, so that the target non-volatile storage units and the remaining non-volatile storage units can complete the access operations in parallel.

7. The non-volatile memory access method according to any one of claims 1 to 6, characterized in that, The step of merging the first narrow-width data output by the target non-volatile memory cell with the second narrow-width data output by the remaining non-volatile memory cells to obtain merged data that matches the bus data width includes: According to the preset data position correspondence between each non-volatile memory unit and the bus data width in the target processor subsystem, the data positions of the first narrow width data and the second narrow width data are determined; The first narrow-width data and the second narrow-width data are merged according to the data position to obtain merged data that matches the bus data width.

8. A non-volatile memory access device, characterized in that, Applied to a multi-core system, the multi-core system comprising a plurality of processor subsystems, each processor subsystem comprising at least two processor cores and non-volatile memory units corresponding to each processor core, the device comprising: The access request parsing module is used to parse the target access request initiated by the target processor core and determine the target non-volatile memory unit pointed to by the target access request. The access module is used to, when the target non-volatile memory unit is a non-volatile memory unit corresponding to the target processor subsystem, directly initiate an access operation to the target non-volatile memory unit through the direct access port of the target processor core, and simultaneously trigger parallel access operations to the remaining non-volatile memory units within the target processor subsystem; the target processor subsystem is the subsystem corresponding to the target processor core. The data merging module is used to merge the first narrow-width data output by the target non-volatile memory cell with the second narrow-width data output by the remaining non-volatile memory cells to obtain merged data that matches the bus data width; the bus data width is the sum of the widths of the narrow-width data of all non-volatile memory cells corresponding to the target processor subsystem. The data return module is used to return the merged data to the target processor core through the direct access port.

9. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor for executing the computer program to implement the non-volatile memory access method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, Used to store a computer program, which, when executed by a processor, implements the non-volatile memory access method as described in any one of claims 1 to 7.