A dynamic configuration system, method and device of storage access path, chip and storage medium

CN122064638BActive Publication Date: 2026-09-22JINDIE SPACETIME (BEIJING) TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610068889.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-19
Publication Date
2026-09-22
Estimated Expiration
2046-01-19

AI Technical Summary

Technical Problem

[0006]本发明旨在解决现有技术中存储访问路径映射机制固定、不灵活的问题,提供一种能够自适应芯片制造差异和用户实际配置的动态配置方案,从而提高芯片良率、系统兼容性与性能

Benefits of technology

[0089]1、提高产品良率:通过动态配置机制,允许芯片内部部分存储控制组件失效后仍能正常工作,将部分不合格芯片转化为可用的降级产品,显著提升了经济效益。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122064638B_ABST
    Figure CN122064638B_ABST
Patent Text Reader

Abstract

The application discloses a kind of dynamic configuration systems, methods, devices, chips and storage medium of storage access path, belong to integrated circuit technical field.The dynamic configuration system of the storage access path includes configuration management module, first nonvolatile storage unit, second communication interface and programmable routing network.Configuration management module is based on the chip internal component state information read from first nonvolatile storage unit, and the external storage device connection information obtained by second communication interface, dynamically determine the target component set actually available, and generate routing configuration parameters accordingly.Programmable routing network responds to the parameter, and routes access request to target component.The application can automatically adapt chip manufacturing defects and user external memory configuration difference by software and hardware cooperation, dynamically optimize storage access path, so as to significantly improve chip yield, system compatibility and access performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of integrated circuits and computer architecture, and more specifically to a system, method, apparatus, chip, and storage medium for dynamically configuring memory access paths for high-performance computing chips. Background Technology

[0002] In large-scale computing chips (such as CPUs, GPUs, and AI processors), multi-level caching and multi-channel memory architectures are commonly used to improve system bandwidth and capacity. From different perspectives, a computer's storage system manifests differently: from a hardware perspective, different memory chips (memory modules) are scattered around the main chip, and the memory controller inside the main chip, as well as the cache controller acting as an intermediate layer, are implemented in different physical areas of the chip; they are independent hardware components. However, from a software perspective (including the operating system and applications), the memory space is a unified, contiguous physical address space. The software does not need to know which specific controller or memory chip the currently accessed data will flow to at the underlying hardware level; it only cares about the performance and capacity of the memory space.

[0003] Memory mapping technology refers to the process of logically unifying physically independent memory controllers and memory chips within a chip into the main memory address space. Its core is to complete the mapping transformation from a hardware-dispersed perspective to a software-unified perspective. Simultaneously, for performance considerations, chips typically add multi-level caches between the memory access initiator (such as the computing core) and the main memory. Therefore, the mapping from the access initiator to the cache controller (called cache mapping or first-level mapping) and the mapping from the cache controller to the memory controller (called memory mapping or second-level mapping) together constitute a complete memory access path. A typical software memory access request's execution flow in hardware is as follows: 1) The computing unit sends the request to the system bus; 2) The bus, based on address logic and cache mapping relationships, routes the request to the corresponding cache controller; 3) The cache controller, based on address logic and memory mapping relationships, further routes the request to the corresponding memory controller, ultimately reaching the target memory chip; 4) Data returns along the path.

[0004] To improve the parallelism of system access and overall bandwidth, the system bus and the interconnection between controllers at all levels often adopt interleaving technology, that is, distributing continuous access requests to multiple parallel cache controllers or memory controllers in a way that involves address interleaving.

[0005] However, in existing technologies, the aforementioned mapping relationships and interleaving patterns are typically fixed during the chip design phase, or only a limited number of static configuration options are provided (semi-fixed). This fixed mechanism brings two prominent practical problems: First, during chip manufacturing, some cache controllers or memory controllers may fail due to manufacturing defects. The fixed interleaving mapping will not function properly due to the failure of the dependent components, resulting in the scrapping of the entire chip and a loss of yield. Second, in real-world usage scenarios, users may not use all the memory channels supported by the chip for cost reasons (e.g., only inserting a single memory stick). The fixed mapping pattern cannot detect and adapt to such external configuration changes, leading to some controllers being idle, address space gaps, or performance not reaching its optimal level, resulting in performance loss and resource waste. Summary of the Invention

[0006] This invention aims to solve the problem of fixed and inflexible storage access path mapping mechanisms in the prior art, and provides a dynamic configuration scheme that can adapt to differences in chip manufacturing and actual user configuration, thereby improving chip yield, system compatibility and performance.

[0007] To solve the above-mentioned technical problems, the present invention adopts the following technical solution:

[0008] A dynamic configuration system for storage access paths, its core lies in achieving intelligent and dynamic configuration of storage access paths through hardware and software collaboration. The system mainly includes:

[0009] A configuration management module;

[0010] A first non-volatile memory cell configured to store first state information for characterizing the available state of a first set of memory control components within the chip;

[0011] A second communication interface configured to acquire second status information of an external storage device connected to the chip;

[0012] And a programmable routing network coupled between at least one access initiator, a first set of storage control components, and a second set of storage control components to form an access path.

[0013] In this system, the configuration management module is configured to: make a comprehensive decision based on the first state information read from the first non-volatile storage unit and the second state information obtained through the second communication interface, thereby identifying and determining the actual set of target components that can be used to process data access requests from the first set of storage control components and the second set of storage control components; and further generate corresponding routing configuration parameters based on the set of target components.

[0014] The programmable routing network is configured accordingly to dynamically establish routing logic in response to routing configuration parameters generated by the configuration management module, thereby accurately guiding and routing access requests from the access initiator to the corresponding components in the target component set.

[0015] Furthermore, the programmable routing network can be implemented using a hierarchical routing architecture, which specifically includes:

[0016] A first-level programmable routing unit is deployed on the access path between the access initiator and the first group of storage control components. Its function is to perform initial routing selection on the access request from the access initiator according to the first routing parameters, and direct it to a corresponding component in the first group of storage control components.

[0017] And a secondary programmable routing unit, which is deployed on the access path between the first group of storage control components and the second group of storage control components, and whose function is to perform secondary routing selection on access requests after the first-level routing based on the second routing parameters, and finally direct them to one or more corresponding components in the second group of storage control components.

[0018] Furthermore, the system's key information sources can be implemented using industry-standard devices and protocols:

[0019] A typical choice for the first non-volatile memory cell is an electronic fuse or a one-time programmable memory. The stored first state information is preferably organized in a bitmap format, in which one or more binary bits are used to uniquely map and represent the availability state of a first or second set of memory control components (e.g., "1" indicates availability, and "0" indicates failure or unavailability).

[0020] And / or,

[0021] A typical implementation of the second communication interface is a controller that conforms to the I2C or I3C bus standard. The second status information obtained through this interface mainly comes from the serial presence detection memory that is standard on the external storage device (such as memory module). The key information read includes the actual presence status of each memory channel (i.e., whether a memory module is installed) and the total memory capacity of each present channel.

[0022] Furthermore, as a core function of the system, the configuration management module is also responsible for performing more granular resource configuration and path organization. Its specific process can be summarized as follows:

[0023] First, the module comprehensively analyzes the first and second state information, and selects an optimal set of target components from all the first and second group storage control components that have been identified as usable, based on the preset component quantity ratio, as the hardware resource pool actually used by the system.

[0024] Subsequently, based on this set of target components, the module performs collaborative configuration of the primary and secondary programmable routing units. The core objective of the configuration is to logically construct multiple independent access groups within the programmable routing network.

[0025] In this architecture, each logical access group is assigned at least one first group storage control component and at least one second group storage control component from the target component set, thus forming a complete local access path. Importantly, after configuration, the routing process of any access request from the first group storage control component to the second group storage control component during its lifecycle will be strictly limited to the logical access group to which it belongs, without cross-routing across different logical access groups.

[0026] Furthermore, in the process of selecting the target component set as described above, the configuration management module follows a defined set of decision-making logic and algorithmic procedures. This process includes the following steps in sequence:

[0027] Step 1, Internal Availability Filtering: The module first parses the first status information, extracts and summarizes all the first group of storage control components that are identified as available, forming an initial available set (first set); at the same time, it summarizes all the second group of storage control components that are identified as available, forming another initial available set (second set).

[0028] Step two, external connectivity correction: The module then refines the second set obtained in step one based on the second status information. Specifically, it checks the physical connection status of the external storage device corresponding to each second group of storage control components in the second set one by one, and removes those components whose corresponding channels are idle (i.e., not connected to a valid storage device) from the set, thereby obtaining a corrected second set, which represents a second group of components that are internally intact and have actual external connectivity available.

[0029] Step 3, Proportion Compliance Judgment and Target Determination: The module then compares the number of components in the first set with the number of components in the corrected second set. If these two numbers exactly meet the system's preset proportion (e.g., 1:1, 1:2, etc.), the module directly determines the union of the first set and the corrected second set as the final target component set.

[0030] Step 4, Intelligent Discarding When the Proportion is Imbalanced: If Step 3 determines that the number of components in the two sets does not meet the preset proportion, the module will initiate a preset discarding strategy. The core principle of this strategy is to maximize the retention of available bandwidth and optimize access performance while meeting the proportion. To this end, the module will actively and iteratively discard one or more currently available components from the first or second set according to a predefined priority order (for example, the primary principle is to prioritize retaining controllers that are physically closer to the access initiator to reduce access latency and balance heat distribution within the chip). This discarding process will continue until the number of components in the two new sets formed by the remaining components once again meets the preset proportion. At this point, the union of these two new sets is determined as the final target component set.

[0031] Furthermore, to achieve efficient and uniform request distribution, the core routing logic of the first-level programmable routing unit is designed as a specific calculation mode based on address bits. This unit receives the physical address carried in the access request and, according to the configured first routing parameters, performs operations on the binary bit sequence of that address, thereby dynamically generating a first routing signal for selecting the first group of storage control components.

[0032] The first routing parameter mainly defines two key dimensions of the operation:

[0033] Interleaving granularity parameter (n): It determines the lowest starting bit position in the address that participates in the routing calculation, and usually corresponds to the logarithm of the smallest data block size accessed by the system (such as the cache line size).

[0034] Interleaving object number parameter (m): It represents the logarithm of the routing targets that need to be distinguished (i.e., the number of the first group of storage control components), which is the bit width of the selection signal SEL.

[0035] Specifically, for the binary bit sequence PA[...] of the input physical address PA, the first-level programmable routing unit performs the following calculations:

[0036] (1) Determine the starting index of the address bits involved in the calculation based on the parameter n.

[0037] (2) Based on parameter m, determine the m-bit selection signal SEL that needs to be generated.

[0038] (3) For each bit of SEL (index i, i from 0 to m-1), its value is not determined by a single address bit, but by performing an XOR operation on a set of bits at specific positions in the address bit sequence. The positions of this set of bits form an arithmetic sequence with the first term being n+i and the common difference being m.

[0039] (4) Therefore, the specific formula for calculating the i-th bit (SEL[i]) of the selected signal SEL is:

[0040] SEL[i]=PA[n+i]⊕PA[n+m+i]⊕PA[n+2m+i]⊕...⊕PA[n+xm+i],

[0041] in:

[0042] The symbol “⊕” represents the bitwise XOR logical operator;

[0043] i is the current signal bit index being calculated, with a value ranging from 0 to (m-1);

[0044] x is an integer coefficient determined by the total physical address width, n, i, and m. Its function is to ensure that the last term n+xm+i does not exceed the valid range of address bits.

[0045] This algorithm efficiently and pseudo-randomly maps contiguous physical address spaces to multiple parallel first-set memory control components, thereby introducing access randomness in addition to address locality, significantly improving the efficiency and load balancing effect of multi-controller parallel operation.

[0046] Furthermore, in contrast to the dynamic calculation mode of the first-level programmable routing unit, the second-level programmable routing unit adopts a relatively simplified fixed mapping mechanism to achieve efficient and deterministic path selection.

[0047] The mapping relationship of this unit is pre-defined statically by the second routing parameters. Specifically, these parameters establish a mapping table for each first group of storage control components, which explicitly specifies one or more second group of storage control components that the first group component can be routed to when it initiates a request. This mapping relationship remains fixed after the system is configured, ensuring the stability and predictability of the access path.

[0048] Building upon this, to further enhance the parallelism within the secondary path, when a first-group storage control component is mapped to multiple second-group storage control components, the secondary programmable routing unit also integrates a fine-grained selector. This selector is configured to dynamically select the final target from the mapped components based on the value of a pre-specified bit in the physical address of each specific access request. For example, the 12th bit of the physical address can be specified as the selection criterion: if the bit is '0', the first component in the mapping group is selected; if it is '1', the second component is selected, and so on.

[0049] By employing this two-tier routing strategy of "static group mapping" plus "dynamic intra-group selection," the system can ensure a simple and controllable overall mapping structure while still leveraging the locality of address space to achieve finer-grained request distribution and load balancing within access groups.

[0050] This invention also provides a dynamic configuration method for storage access paths. This method, executed by software, aims to intelligently initialize and configure storage access paths in coordination with hardware through a series of ordered steps. The method mainly includes the following stages:

[0051] Phase 1: Information Acquisition

[0052] At this stage, the method performs a dual information acquisition operation. Specifically, this includes:

[0053] (1) Internal status acquisition: Access the non-volatile storage medium (such as electronic fuse) integrated in the chip through the system bus (such as AXI or APB bus), read the device status data recorded in the bitmap format, and parse the data into first status information characterizing the availability of the first and second sets of storage control components inside the chip.

[0054] (2) External configuration detection: Through a communication interface that conforms to the I2C or I3C standard, actively detect the external storage devices connected to each channel of the chip, read the standard configuration data from its serial presence detection memory, and parse the data into second status information that characterizes the actual connected memory channels and their capacity.

[0055] Phase Two: Intelligent Decision Making.

[0056] At this stage, the method performs fusion analysis and decision-making on the collected dual information. Its core task is to intelligently filter and determine an optimized set of target components that can actually be used to process data access requests from the complete set of the first and second sets of storage control components physically present on the chip, based on the availability of internal components revealed by the first state information and the actual external connection status reflected by the second state information.

[0057] Third stage: Parameter generation.

[0058] In this stage, the method digitally models routing rules based on the target component set determined in the second stage. Specifically, it calculates and generates a set of path configuration parameters that precisely describe how access requests are routed to each specific component within that target component set. These parameters will be used to program the subsequent hardware routing network.

[0059] Phase 4: Hardware Configuration.

[0060] In this stage, the method writes the path configuration parameters generated in the third stage into the configuration register corresponding to the hardware routing network through a specific software interface, thereby completing the programming and solidification of the logical functions of the hardware routing network.

[0061] After the above four stages are completed, the hardware routing network is successfully configured. Thereafter, any access request from the initiating party will be automatically and accurately guided and routed by the network to the corresponding component in the target component set according to the established routing logic, thereby achieving dynamic and optimized storage access.

[0062] Furthermore, in the intelligent decision-making stage of the method, its core logic can be further refined into a closed-loop process that includes quantitative analysis, conditional judgment, and optimization adjustment. This process is specifically executed as follows:

[0063] First, resource quantification statistics are performed. The method parses the first and second state information to accurately count the total number of the first group of usable storage control components, denoted as N; simultaneously, it accurately counts the total number of the second group of storage control components that are both internally usable and externally actually connected and usable, denoted as M. This step transforms the ambiguous "availability" status into a clear quantitative indicator.

[0064] Next, a proportional compliance check is performed. The method compares the statistically obtained actual quantity pairs (N, M) with the system's preset or optimal component quantity ratio K to determine whether the current actual resource allocation conforms to the ideal ratio of the architecture design. The ratio K can be, for example, 1:1, 1:2, or 2:1.

[0065] Then, implement dynamic resource optimization and adjustments (if necessary), which is the key to the decision-making process:

[0066] If the ratio K is satisfied after verification, the method directly adopts the current N first group components and M second group components as the target set without adjustment.

[0067] If the ratio K is not satisfied after verification, the method initiates an optimization algorithm. This algorithm has two optimization objectives: the primary objective is to maximize the total available storage access bandwidth of the system while satisfying the ratio K; the secondary objective (provided the primary objective is satisfied) is to optimize selection based on preset component location priorities (e.g., prioritizing components physically closer to the access initiator). To achieve this objective, the algorithm strategically iteratively discards one or more components from the currently available set of N or M components. This discarding process is dynamically calculated until the new N' and M' of the remaining components strictly satisfy the preset ratio K.

[0068] Finally, parametric modeling is completed. Based on the target component set determined through the above steps and the proportional relationship K it satisfies, the method performs logical topology construction, that is, generates specific control parameters for defining multiple logical access groups in the hardware routing network. These parameters mainly include first path parameters (used to control the routing of requests to the first group of components) and second path parameters (used to control the routing between the first group of components and their mapped second group of components).

[0069] Through this series of meticulous decision-making steps, the method ensures that even in suboptimal situations such as internal component failure or incomplete external configuration, the system can still automatically calculate and configure a storage access path topology that is optimal or near-optimal under given architectural constraints.

[0070] Furthermore, in the parameter generation stage of the method, the core function of the generated first path parameters is to program the calculation logic of the first-level routing unit. Specifically, these parameters configure the unit to perform a specific XOR interleaving algorithm for each input access physical address PA, thereby calculating the routing selection signal SEL used to select the first set of storage control components.

[0071] The mathematical expression and parameter definition of the algorithm are as follows:

[0072] For the i-th bit of signal SEL (where i is an integer from 0 to m-1), its logical value is determined by a set of bits at specific positions in the physical address PA. The indices of these bits in the address bit sequence form an arithmetic sequence with the first term being n+i and a common difference of m. Therefore, the specific formula for calculating SEL[i] is as follows:

[0073] SEL[i]=PA[n+i]⊕PA[n+m+i]⊕PA[n+2m+i]⊕...⊕PA[n+xm+i],

[0074] The physical and technical meanings of the symbols in the formula are as follows:

[0075] ⊕: Represents bitwise XOR logical operation, which is the core operator of this algorithm.

[0076] n: Represents the logarithm of the interleaving granularity. This is a configuration parameter whose value determines the lowest address bit index involved in the calculation, and is usually associated with the size of the smallest interleaved access unit in the system (such as a cache line), which is 2^n bytes.

[0077] m: Represents the logarithm of the interleaving degree. This is another key configuration parameter, and its value m satisfies the condition that the number of interleaving objects = 2^m. It directly determines the bit width of the routing selection signal SEL and the "step size" of the address bits selected in the algorithm.

[0078] i: is the index of the currently calculated select signal bit, with a value range of 0≤i<m, and is used to generate each bit of SEL.

[0079] x: is a non-negative integer coefficient derived jointly from the total effective bit width of the physical address, n, i and m. Its technical effect is to ensure that the last term n+xm+i in the formula does not exceed the most significant bit of the physical address PA, thereby determining the maximum index of the address bits participating in the XOR operation.

[0080] By configuring the first path parameters (n,m) and applying the formula, the method enables the first-level routing unit to map a linear address space to a plurality of parallel first-group memory control components in a highly randomized manner. This XOR-based mapping strategy effectively breaks access hotspots that may be caused by simple continuous address mapping, and significantly improves access parallelism and overall bandwidth utilization in a multi-controller environment.

[0081] The present invention also provides a programmable routing device, which serves as the core hardware entity of the system and is specially designed to efficiently and flexibly implement multi-level routing and forwarding of memory access requests inside a chip. The device is deployed between the access path formed by the access initiator, the first group of memory control components and the second group of memory control components, and its main hardware structure and functions are as follows:

[0082] First, the device includes a configuration register group. As the software programmable interface and control center of the device, the register group is used for non-volatile or power-on loading storage of path configuration parameters generated by software. These parameters determine the behavior of all routing logic inside the device.

[0083] Second, the device includes a first routing logic unit. This unit is directly coupled to the access initiator in hardware, and is the first station for an access request to enter the device. It is designed such that whenever an access request arrives, it immediately responds to the request and synchronously performs two key operations: 1) acquiring the currently effective first parameter from the configuration register group; 2) executing a preset routing calculation based on the first parameter and the physical address carried by the access request. The final result of the calculation is to generate and output a first select signal, which uniquely determines which specific first group of memory control components the current access request should be directed to.

[0084] Finally, the device includes a second routing logic unit. This unit is connected in hardware series between the first and second sets of storage control components, forming the next hop in the access path. It is designed to listen for and respond to access requests forwarded by the first set of storage control components. For each such request, the unit determines the next destination of the request based on a second parameter obtained from the configuration register set. Specifically, the unit routes the received request precisely to a specific one of the one or more second sets of storage control components associated with it, according to the mapping defined by the second parameter.

[0085] In summary, this programmable routing device, through the coordinated operation of its configuration register group, first routing logic unit, and second routing logic unit, achieves two-level, programmable, pipelined routing of access requests. The software can dynamically redefine the entire chip's memory access topology by writing to the configuration registers, while the hardware logic ensures high performance and low latency in routing processing.

[0086] A chip is characterized by integrating a dynamic configuration system for memory access paths as described above, or integrating a programmable routing device as described above.

[0087] A computer-readable storage medium having a computer program stored thereon, characterized in that, when the computer program is executed by a processor, it is able to implement the steps of the dynamic configuration method for storage access paths as described above.

[0088] The present invention, by adopting the above-described technical solution, has the following beneficial effects:

[0089] 1. Improve product yield: Through a dynamic configuration mechanism, some memory control components inside the chip can still work normally after failure, turning some unqualified chips into usable downgraded products, which significantly improves economic benefits.

[0090] 2. Enhanced system compatibility and flexibility: Automatically identifies the user's actual installed memory configuration and optimizes the access path accordingly. Users do not need to manually configure or change firmware, thus improving the user experience.

[0091] 3. Optimize system performance: Through intelligent component selection and efficient interleaving algorithms (such as XOR interleaving), the system can maximize the use of available bandwidth and maintain system performance even when components are incomplete.

[0092] 4. Provides complete software and hardware solutions: This invention covers the entire chain from information recognition, decision-making algorithms, hardware configuration to the final chip product, providing comprehensive protection and being easy to implement and promote. Attached Figure Description

[0093] The present invention will be further described below with reference to the accompanying drawings:

[0094] Figure 1 This is a schematic diagram of an ideal two-stage interleaving system in one embodiment of the present invention;

[0095] Figure 2 This is a schematic diagram illustrating the access group division under different controller ratios according to the present invention;

[0096] Figure 3 This is a schematic diagram illustrating the mapping when some controllers fail in one embodiment of the present invention;

[0097] Figure 4 A flowchart illustrating the software configuration process of this invention;

[0098] Figure 5 This is a schematic diagram of the system configuration phase interaction in one embodiment of the present invention;

[0099] Figure 6 This is a schematic diagram of the memory access operation phase interaction in one embodiment of the present invention;

[0100] Figure 7 This is an example diagram of address mapping routing in one embodiment of the present invention. Detailed Implementation

[0101] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it.

[0102] The core of this invention lies in providing a dynamic configuration mechanism for storage access paths through hardware and software collaboration. The following description uses a typical application scenario in a high-performance processor where the cache controller is the first group of storage control components and the memory controller is the second group of storage control components as an example; however, the scope of protection of this invention is not limited thereto.

[0103] Example 1: System Architecture and Configuration Process

[0104] This embodiment describes the overall architecture and initialization configuration process of the system, corresponding to... Figure 1 , Figure 5 and Figure 6 .

[0105] refer to Figure 1 The system of the present invention includes: a configuration management module, a first non-volatile storage unit, a second communication interface, and a programmable routing network.

[0106] Configuration management module: It is usually implemented in the form of firmware (such as BIOS or BootLoader) running on the CPU or management unit.

[0107] First non-volatile memory unit: In this embodiment, an electronic fuse is used. After chip production testing, the test equipment programs the serial numbers of the failed cache controller and memory controller into the eFuse via the JTAG interface, forming a bitmap-formatted first status information. When the system powers on, the configuration management module accesses the integrated eFuse controller via the chip's internal system bus (such as the APB bus) to read this bitmap-formatted information.

[0108] Second communication interface: In this embodiment, an I2C controller is used. When the system starts up, the configuration management module accesses the SPD chip of the memory module on each memory slot through the I2C interface, reads the second status information, and thus accurately knows which memory channels actually have memory modules installed and their capacities.

[0109] Programmable routing networks: such as Figure 1 As shown, it includes a first-level programmable routing unit (located between the bus and the cache controller) and a second-level programmable routing unit (located between the cache controller and the memory controller). Each routing unit contains a set of software-configurable registers for controlling its routing logic.

[0110] The initialization and configuration process of the entire system is as follows: Figure 4 As shown:

[0111] 1. The system is powered on and the software (configuration management module) starts.

[0112] 2. Read eFuse information: Access the eFuse controller via the internal bus to obtain the available status bitmap of the chip's internal cache controller and memory controller.

[0113] 3. Read memory SPD information: Poll each memory channel through the I2C interface to read SPD data and confirm the actual number of memory channels in use and the capacity of each channel.

[0114] 4. Calculate and determine the mapping mechanism: The configuration management module integrates the information from the above two aspects and dynamically calculates the optimal storage access path configuration based on the preset optimization algorithm (see Example 2 for details).

[0115] 5. Configure hardware parameters: Write the calculated routing configuration parameters into the configuration register group of the programmable routing network (including first-level and second-level units) to complete the initialization of hardware logic.

[0116] 6. The system completes initialization and enters normal operating mode. From then on, all memory access requests will be routed according to the configured paths, such as... Figure 6 The memory access request interaction flow is shown.

[0117] Example 2: Algorithm for Target Component Set Selection and Access Group Formation

[0118] This embodiment details the intelligent decision-making logic of the configuration management module, corresponding to... Figure 2 and Figure 3 .

[0119] Assuming the design specifications are N cache controllers and M memory controllers (ideal ratio K=N:M=1:2), the manufacturing and actual usage are as follows: the eFuse bitmap shows that cache controller C2 is faulty; the I2C probe information shows that only memory modules are inserted on memory channels 0 and 2 (i.e., memory controllers M1 and M3 are not connected).

[0120] 1. Preliminary screening: Based on the eFuse bitmap, the set of internally available cache controllers N_usable={C0,C1,C3} is obtained; based on the I2C information, the set of actually available memory controllers M_usable={M0,M2} is obtained.

[0121] 2. Ratio Judgment and Intelligent Discarding: The preset ratio K = 1:2. Currently, n = 3 and m = 2, which does not satisfy the ratio. The algorithm aims to find the largest subset that satisfies the ratio, prioritizing the retention of controllers closer to the CPU (the initiating access point) to optimize latency. After trying, the algorithm decides to retain cache controllers {C0, C1} and memory controllers {M0, M2}. At this point, n' = 2 and m' = 2, satisfying the 1:1 ratio (another preset ratio supported by the system). The intact controller C3 is strategically discarded and disabled.

[0122] 3. Forming Logical Access Groups: The final determined target component sets {C0, C1} and {M0, M2} are divided into two logical access groups in a 1:1 ratio: Group0: (C0->{M0}), Group1: (C1->{M2}). Based on this, the software calculates the first routing parameter (routing the request to C0 or C1) and the second routing parameter (within Group0, requests from C0 are fixed to M0; within Group1, requests from C1 are fixed to M2), and writes them into the hardware registers.

[0123] Example 3: Level 1 routing XOR interleaving algorithm

[0124] This embodiment corresponds to Figure 7 The left side.

[0125] The first-level programmable routing unit is configured to execute an interleaving algorithm based on address bit XOR operations. Assume the system is configured to use four cache controllers (i.e., the number of interleaving objects is four, hence m=2), and the interleaving granularity is the cache line size (64 bytes, hence n=6).

[0126] For the input 32-bit physical address PA[31:0], this cell does not simply take a few bits as an index, but performs the following calculation to generate the 2-bit selection signal SEL[1:0]:

[0127] SEL[0]=PA[6]⊕PA[8]⊕PA

[10] ⊕…⊕PA

[30] ;

[0128] SEL[1]=PA[7]⊕PA[9]⊕PA

[11] ⊕…⊕PA

[31] ;

[0129] Here, ⊕ represents the XOR operation. The address bit indices involved in the operation form an arithmetic sequence with the first term n+i and a common difference of m. This algorithm can map contiguous address spaces to multiple cache controllers in a highly randomized manner, effectively improving access parallelism and load balancing.

[0130] Example 4: Two-level routing and address mapping

[0131] Following Example 3, refer to Figure 7 The right side.

[0132] Assume the first-level routing calculation result SEL[1:0]='2' (binary 10) selects cache controller C2. Further assume that, according to the configuration of Example 2, C2 is fixedly mapped to memory controllers {M4,M5}.

[0133] The secondary programmable routing unit adopts a fixed mapping and can perform fine-grained selection within a group based on a specific address. In this example, it is configured to use address PA

[10] for selection:

[0134] If PA

[10] =0, the request will be routed to memory controller M4.

[0135] If PA

[10] =1, the request will be routed to memory controller M5.

[0136] Thus, a memory access request (address PA) has completed a full dynamic route from the CPU to the final memory controller.

[0137] Example 5: A specific configuration and address mapping example

[0138] To further clarify the interleaving mapping mechanism of the present invention, a specific numerical example is provided below.

[0139] Assume a system with 4 cache controllers as the first group of storage control components and 8 memory controllers as the second group of storage control components, with a physical address width of 32 bits (y=32). The system design uses cache lines (64 bytes) as the interleaving granularity, so the interleaving granularity logarithm n=log2(64)=6. Since there are 4 cache controllers, the number of interleaving objects is 4, so its logarithm m=log2(4)=2.

[0140] At this time, the first-level routing unit performs calculations according to the algorithm of Example 3. At the same time, the 10th bit (PA

[10] ) of the second-level interleaving selection address is configured as the basis for selection within the mapping group.

[0141] The overall interleaving method and address mapping routing path under this configuration are as follows: Figure 7 As shown in the figure, the above calculation and selection process is demonstrated using a path to an example address.

[0142] Address Bit Handling Explanation: During the above interleaving process, the address bits used for routing (mainly PA[6], PA[7], ..., PA

[31] and PA

[10] in this example) need to be truncated (i.e. ignored) within the memory controller after the request is routed to the final memory controller. For example, in this case, PA[7:6] and PA

[10] usually need to be truncated. If this truncation is not performed, although it will not affect the correctness of the basic read and write functions, it will cause the upper limit of the physical address space that the system can address to be unnecessarily reduced, because the address bits that have been used for routing are still interpreted as address offsets within the memory particles. By truncating, it can be ensured that the entire physical address space of the system is effectively used to address memory units.

[0143] Example 6: Chip and Storage Medium

[0144] This embodiment provides a chip that integrates a dynamic configuration system for memory access paths as described in any of embodiments 1 to 5. The chip may be a central processing unit (CPU), a graphics processing unit (GPU), or an artificial intelligence accelerator (AI Processor).

[0145] This embodiment also provides a computer-readable storage medium, such as a UEFI firmware chip, a driver in a hard disk or flash memory, on which a computer program (instructions) are stored. When the program is executed by the processor in the chip or system, it can achieve the following: Figure 4 The method steps described in Examples 2 and 5 are used to complete the dynamic configuration of the storage access path.

[0146] The above are merely specific embodiments of the present invention, but the technical features of the present invention are not limited thereto. Any simple changes, equivalent substitutions, or modifications made based on the present invention to solve essentially the same technical problems and achieve essentially the same technical effects are all covered within the protection scope of the present invention.

Claims

1. A dynamic configuration system for storage access paths, characterized in that, include: Configuration management module; The first non-volatile memory cell is used to store first state information characterizing the availability of the first set of memory control components inside the chip; The second communication interface is used to obtain the second status information of the external storage device connected to the chip. A programmable routing network is coupled between at least one access initiator and the first group of storage control components and the second group of storage control components; in, The configuration management module is used to determine the set of target components that can actually be used for data access in the first group of storage control components and the second group of storage control components based on the first status information and the second status information, and generate routing configuration parameters accordingly. The programmable routing network, in response to the routing configuration parameters, is configured to route access requests from the access initiator to the corresponding components in the target component set.

2. The dynamic configuration system for storage access paths according to claim 1, characterized in that, The programmable routing network includes: A first-level programmable routing unit, disposed between the access initiator and the first group of storage control components, is used to perform a first routing of the access request based on the first routing parameters; and A secondary programmable routing unit is disposed between the first group of storage control components and the second group of storage control components, and is used to perform a second routing of the request based on the second routing parameters.

3. The dynamic configuration system for storage access paths according to claim 2, characterized in that: The first non-volatile storage unit is an electronic fuse or a one-time programmable memory, and the first status information is in bitmap format, wherein each bit or more bits correspond to a status identifier of the first group or the second group of storage control components; And / or, The second communication interface is an I2C bus interface or an I3C bus interface, and the second status information includes channel presence information and capacity information read from the serial presence detection memory of the external storage device.

4. The dynamic configuration system for storage access paths according to claim 2, characterized in that, The configuration management module is further used for: Based on the first state information and the second state information, from all available first and second group storage control components, according to a preset quantity ratio, available first and second group storage control components are selected to form the target component set; Configure the first-level programmable routing unit and the second-level programmable routing unit to form multiple logical access groups based on the target component set; Each logical access group includes at least one first group of storage control components and at least one second group of storage control components from the target component set, and the access request is routed from the first group to the second group of storage control components within the logical access group.

5. A dynamic configuration system for storage access paths according to claim 4, characterized in that, When the target component set is selected, the configuration management module is configured to execute: Based on the first status information, all available first group storage control components are selected to form a first set, and all available second group storage control components are selected to form a second set; Based on the second status information, remove the components that are not connected to the corresponding external storage device from the second set; If the number of components in the first set and the second set meets the preset ratio, then they are taken as the target component set; If not, one or more available components are discarded from the first set or the second set according to the preset discard priority, until the number of remaining components meets the ratio relationship, wherein the discard priority includes giving priority to retaining components whose physical location is closer to the access initiator.

6. A dynamic configuration system for storage access paths according to claim 2, characterized in that, The first-level programmable routing unit is configured to perform calculations based on the physical address of the access request to generate a first routing signal; The first routing parameter includes an interleaving granularity parameter and an interleaving object number parameter. The interleaving granularity parameter determines the lowest starting bit position in the address that participates in the routing calculation, corresponding to the logarithm of the smallest data block size accessed by the system. The interleaving object number parameter represents the number of routing targets that need to be distinguished, i.e., the bit width of the selection signal SEL. The first-level programmable routing unit is configured as follows: For the bit sequence of the input physical address PA, an m-bit selection signal SEL is calculated based on the starting bit n determined by the interleaving granularity parameter and the number m determined by the interleaving object quantity parameter, where the i-th bit of SEL is: SEL[i]=PA[n+i]⊕PA[n+m+i]⊕PA[n+2m+i]⊕...⊕PA[n+xm+i], Where "⊕" represents the XOR operator, i is an integer from 0 to m-1, and x is an integer coefficient determined by the address width.

7. A dynamic configuration system for storage access paths according to claim 6, characterized in that, The second-level programmable routing unit adopts a fixed mapping algorithm, and the second routing parameters are used to define one or more second-level storage control components mapped to each first-level storage control component. The secondary programmable routing unit is further configured to select from a plurality of second-group storage control components mapped by a first-group storage control component, based on a specified bit of the physical address of the access request.

8. A method for dynamically configuring storage access paths, characterized in that, include: Acquisition steps: Read the bitmap format information stored in the non-volatile medium inside the chip via the chip's internal system bus, as the first state information; And read the SPD information of the external storage device connected to the chip via the I2C or I3C protocol as the second status information; Decision-making steps: Based on the first state information and the second state information, determine the actual set of target components that can be used for data access from the first group of storage control components and the second group of storage control components inside the chip; Generation step: Based on the target component set, generate path configuration parameters for configuring the hardware routing network; Configuration steps: Configure the hardware routing network using the path configuration parameters; The configured hardware routing network is used to route access requests to the corresponding components in the target component set.

9. A method for dynamically configuring storage access paths according to claim 8, characterized in that, The decision-making steps include: Based on the first status information and the second status information, determine the actual number N of the first group of storage control components and the number M of the second group of storage control components that are available; Determine whether N and M satisfy the preset ratio K; If not satisfied, one or more of the available components are discarded based on maximizing available bandwidth and component location priority, so that the number of remaining components satisfies the ratio K. Based on the final determined set of target components and their proportional relationships, first path parameters and second path parameters are generated to define multiple logical access groups.

10. A method for dynamically configuring storage access paths according to claim 9, characterized in that, The first path parameter is used to configure the first-level routing unit, such that the first-level routing unit calculates the i-th bit of the routing signal SEL based on the access physical address PA in the following way: SEL[i]=PA[n+i]⊕PA[n+m+i]⊕PA[n+2m+i]⊕...⊕PA[n+xm+i], Wherein, ⊕ represents the XOR operation, n is the interleaving granularity parameter, which determines the lowest starting bit position in the address that participates in the routing calculation, corresponding to the logarithm of the smallest data block size accessed by the system, m is the logarithm of the number of interleaving objects to base 2, which represents the number of routing targets that need to be distinguished, i.e. the bit width of the selection signal SEL, i is an integer from 0 to m-1, and x is an integer coefficient determined according to the address bit width.

11. A programmable routing device for connecting an access initiator, a first set of memory control components, and a second set of memory control components in a chip, characterized in that, include: The configuration register group is used to store path configuration parameters; The first routing logic unit is coupled to the access initiator and is configured to: in response to an access request, calculate and output a first selection signal based on the first parameter in the configuration register group and the address of the access request, so as to direct the access request to the corresponding component in the first group of storage control components; The second routing logic unit, coupled between the first group of storage control components and the second group of storage control components, is configured to: in response to a request from the first group of storage control components, route the request to one or more corresponding components in the second group of storage control components according to the second parameter in the configuration register group.

12. A chip, characterized in that, The system includes a dynamic configuration system for storage access paths as described in any one of claims 1 to 7, or includes a programmable routing device as described in claim 11.

13. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the dynamic configuration method for storage access paths as described in any one of claims 8 to 10.

Citation Information

Patent Citations

  • Programmable coherent proxy for attached processor

    CN103838567A

  • Access processing device and method, processing equipment, electronic equipment and storage medium

    CN114661654A