Chip system and channel interleaving method thereof
By setting up memory channel interleaving units in the main chip to perform address translation for access requests, the problem of unbalanced access to memory units in the chip system is solved, load balancing of memory units and channels is achieved, and bandwidth utilization and performance are improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- XINGYUNJIWEI (CHENGDU) TECHNOLOGY CO LTD
- Filing Date
- 2025-12-23
- Publication Date
- 2026-04-14
AI Technical Summary
In existing technologies, access requests to memory cells in chip systems typically determine the target memory cell directly through bus routing. This results in frequent access to the same memory cell while other memory cells remain idle, leading to insufficient utilization of the overall system bandwidth and underutilization of performance potential.
A memory channel interleaving unit is set in the main chip. The address of the access request is transformed by the first interleaving parameter to determine the target storage unit. The access request is then transmitted by the main routing module to achieve balanced access to the storage units. At the same time, the interleaving module performs load balancing on the storage unit channels.
It improves the bandwidth utilization and channel load balancing of multi-memory-cell, multi-access-channel chip architecture, reduces serial blocking, and enhances the performance of the chip system.
Smart Images

Figure CN121858480A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a chip system and its channel interleaving method. Background Technology
[0002] As chip system scale continues to expand and applications such as artificial intelligence and high-performance computing become more widespread, the memory architecture within chips is gradually evolving from the traditional single-memory-cell model to a complex structure with multiple memory cells and multi-channel parallel access. Masters capable of initiating access requests, including GPUs and CPUs, can typically access multiple memory cells simultaneously through multiple access channels. This includes technologies such as DDR (Double Data Rate Synchronous Dynamic Random-Access Memory), GDDR (Graphics Double Data Rate Synchronous Dynamic Random-Access Memory), and LPDDR (Low Power Double Data Rate SDRAM) to improve overall memory bandwidth and data access efficiency.
[0003] In existing technologies, access requests typically determine the target memory cell directly through bus routing. This type of bus routing usually does not perform any form of address translation or remapping of the access address, thus a large number of consecutive access requests are assigned to the same memory cell. When a memory cell is frequently accessed, access conflicts occur, requiring serial processing, while other memory cells remain idle. This results in the overall system bandwidth not being fully utilized, thereby reducing the performance potential of the chip system.
[0004] There is an urgent need for a balanced access scheme for different memory cells in a chip system. Summary of the Invention
[0005] In view of this, embodiments of this application provide a chip system and its channel interleaving method to improve the bandwidth utilization and channel load balancing of a multi-memory-cell, multi-access-channel chip architecture in centralized response to access requests.
[0006] A first aspect of this application provides a chip system, including: a main die and a plurality of memory cells, wherein at least one memory cell channel is provided between the main die and each memory cell, and the main die includes: The main interleaving module includes a memory channel interleaving unit, which is configured to: obtain the address of the access request, and perform a first interleaving transformation on the address of the access request according to the address of the access request and a first interleaving parameter to obtain the target address. The first interleaving parameter includes the number of storage units that the access request can correspond to. The main routing module is connected to the main interleaving module. The main routing module is configured to transmit access requests to the target storage unit through the corresponding storage unit channel based on the target address.
[0007] In some implementations of the first aspect, the chip system includes a slave chip, which includes a slave routing module and a slave interleaving module. The storage unit includes multiple master storage units and multiple slave storage units. The slave chip and the master chip are connected through interconnect channels. The master chip and the master storage unit are connected through master storage unit channels that correspond one-to-one with each master storage unit. The slave chip and the corresponding slave storage unit are fully interconnected through multiple slave storage unit channels. The number of storage units includes the number of master storage units and the number of slave storage units. The interleaving module is configured to: if the target address corresponds to a slave storage unit, or the access request corresponds to a specific target slave storage unit; Then the interleaving module performs a second interleaving transformation based on the second interleaving parameters, and determines the target memory channel from the slave storage cell corresponding to the target slave storage cell. The second interleaving parameters include the number of corresponding slave storage cell channels, and the corresponding storage cell channels include the target memory channel.
[0008] In some implementations of the first aspect, the chip system includes slave chips, the master interleaving module further includes interconnect channel interleaving units, and the memory unit includes multiple master memory units and multiple slave memory units. The interconnect channel interleaving unit is configured such that each slave core is connected to the master core through multiple interconnect channels; The interconnect channel interleaving unit performs a third interleaving transformation based on the third interleaving parameters to determine the target interconnect channel from the corresponding interconnect channels. The third interleaving parameters include the number of corresponding interconnect channels, and the corresponding memory cell channel includes the target interconnect channel.
[0009] In some implementations of the first aspect, the memory channel interleaving unit includes a first address partitioning unit and a first numbering calculation unit; The first address partitioning unit divides the address of the access request into a first calculated bit and a first reserved bit. The first calculated bit is used to determine the target storage unit, and the access address of the target storage unit includes the first reserved bit. The first numbering calculation unit groups the first calculation bits, sums all the groups to obtain the feature value, and determines the number of the target storage unit based on the feature value; The target address includes the access address and the number of the target storage unit.
[0010] In some implementations of the first aspect, the interleaving parameters also include interleaving granularity, which is determined by factors including bus width and cache line size. The first address partitioning unit divides the address of the access request into a first computed bit and a first reserved bit according to the interleaving granularity; The first numbered calculation unit groups the first calculation bits according to the number of storage units that can correspond to the access request.
[0011] In some implementations of the first aspect, the first numbering calculation unit includes: The first numbered calculation subunit is configured as follows: The number of the target storage unit is obtained by performing a modulo operation between the feature value and the number of corresponding storage units.
[0012] In some implementations of the first aspect, the chip system includes: The mapping storage sub-unit, connected to the memory channel interleaving unit, is configured to: pre-store the mapping relationship between feature values and the number of the target storage unit, so that it can be accessed by the number calculation unit and the number of the target storage unit can be directly obtained based on the feature values.
[0013] In some implementations of the first aspect, the interleaving module is configured to: determine the target memory channel number based on the feature value and the number of memory channel channels; and determine the target memory channel based on the target memory channel number. Alternatively, the interleaving unit of the main interleaving module is configured as follows: determine the number of the target interleaving channel based on the characteristic value and the number of interleaving channels; determine the target interleaving channel based on the number of the target interleaving channel.
[0014] In some implementations of the first aspect, the chip system also includes: Interconnect chip: One side of the interconnect chip is connected to the master chip via an L2C channel, and the other side is connected to a slave chip via one or more interconnect channels; the number of L2C channels is the same as the number of slave storage units connected to the corresponding slave chip.
[0015] A second aspect of this application provides a channel interleaving method for a chip system. The channel interleaving method is applicable to the chip system described in the first aspect above. The chip system includes a main interleaving module and a main routing module interconnected with each other. The channel interleaving method includes: The main interleaving module obtains the address of the access request and performs a first interleaving transformation on the address of the access request based on the address of the access request and the first interleaving parameters to obtain the target address. The first interleaving parameters include the number of storage units that the access request can correspond to. The main routing module controls the transmission of access requests to the target storage unit through the corresponding storage unit channel based on the target address.
[0016] A third aspect of this application provides a computer program product including a computer program that, when run, causes the channel interleaving method of the chip system as described in the first aspect to be executed.
[0017] This embodiment of the application incorporates a memory channel interleaving unit within the main die of the chip system. This unit performs a first interleaving transformation on the received access request based on the access request address and the number of corresponding memory units, obtaining the target address. This allows an independent interleaving module, rather than a main routing module, to perform the first interleaving transformation on the access request address, thus freeing the first interleaving process from the limitations of the main routing module and making it applicable to different buses. After the memory channel interleaving unit obtains the target address through the first interleaving transformation, the main routing module determines the target memory unit to receive the access request based on the target address and sends the access request to the target memory unit, achieving balanced access to memory units. Simultaneously, the main routing module can also determine the memory unit channel for transmitting the access request based on the target address, achieving load balancing when multiple memory unit channels exist between the same memory unit and the main routing module. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a schematic diagram of a chip system provided in an embodiment of this application; Figure 2 This is a schematic diagram of another chip system provided in an embodiment of this application; Figure 3 This is a schematic diagram of an address mapping provided in an embodiment of this application; Figure 4 This is a schematic diagram of a channel interleaving method for a chip system provided in an embodiment of this application. Detailed Implementation
[0020] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0021] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.
[0022] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0023] As used in this application specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if detected [the described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once detected [the described condition or event]," or "in response to detection [the described condition or event]."
[0024] Furthermore, in the description of this application and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0025] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0026] It should be noted that the modules and units mentioned in this application may include one or more hardware circuits, including but not limited to: application-specific integrated circuits, digital signal processors, field-programmable gate arrays, discrete logic circuits, state machines, and any combination of the aforementioned circuits; these circuits are specifically designed, configured, and interconnected to perform one or more specific functions disclosed in this application.
[0027] One of the concepts in this application is to set an interleaving module in the main chip. The interleaving module can flexibly interleave based on the number of storage cells and can perform non-power-four interleaving to obtain the target address. By selecting the corresponding storage cell channel according to the target address, the access request is transmitted to the target transmission unit to be accessed, thereby achieving balanced access to the storage cells and storage cell channels.
[0028] Reference Figure 1 This illustration shows a chip system provided in an embodiment of this application. The chip system includes a main chip 100 and a plurality of memory cells 200. At least one memory cell channel 300 is provided between the main chip 100 and each memory cell 200. The main chip 100 includes a main interleaving module 110, which includes a memory channel interleaving unit 111. The memory channel interleaving unit 111 is configured to: obtain the address of an access request, and perform a first interleaving transformation on the address of the access request according to the address of the access request and a first interleaving parameter, to obtain a target address. The first interleaving parameter includes the number of memory cells 200 that can correspond to the access request. A main routing module 120 is connected to the main interleaving module 110 and is configured to: transmit the access request to the target memory cell through the corresponding memory cell channel 300 according to the target address. It is understood that these memory cells can be directly connected to the main chip or indirectly connected to the main chip.
[0029] The chip system provided in this application embodiment includes a main chip 100, which is connected to multiple memory cells 200. According to the performance requirements of the chip system, the main chip 100 can be configured to connect to the memory cells 200 through one or more memory cell channels 300.
[0030] A memory channel interleaving unit 111 is set in the main die 100 to perform address interleaving. Specifically, when the main die 100 receives an access request for the memory unit 200, the memory channel interleaving unit 111 can obtain the access request and determine the address of the access request. Alternatively, the access request can also be generated by the DMA unit within the main die 100.
[0031] When the access request corresponds to a specified storage unit 200 (for example, the access request is to read data in a specified storage unit 200), there is no need to perform the first interleaving transformation. The address of the access request can be used as the target address, and the main routing module 120 can transmit the access request to the specified storage unit 200 through the corresponding storage unit channel 300 according to the target address.
[0032] When the address of the access request does not specify a storage unit 200 (e.g., random access storage), or corresponds to multiple selectable storage units 200 (accessing a specific number of storage units 200), the channel interleaving unit in the main interleaving module 110 performs a first interleaving transformation on the address of the access request according to the number of storage units 200 that satisfy the access request to obtain the target address. The main routing module 120 then executes the request based on the data information of the access request (non-address information, such as specific data to be read or written, control data, etc.) and the target address, including determining the storage unit 200 corresponding to the target address as the target storage unit, and transmitting the access request to the target storage unit through the storage unit channel 300 corresponding to the target address.
[0033] Since the embodiments of this application perform address interleaving transformation of the storage unit 200 by setting an independent main interleaving module 110, compared with the bus (equivalent to the main routing module 120 in the embodiments of this application) directly accessing without interleaving, or the scheme of performing fixed power-two interleaving through the bus, the embodiments of this application do not require hardware-level improvements to the bus, and can be better adapted to various types of buses. At the same time, non-power-two interleaving can also be performed, which improves the flexibility of the chip system architecture.
[0034] Taking a write request as an example, the memory channel interleaving unit 111 performs a first interleaving transformation on multiple consecutive access requests according to the number of corresponding storage units 200, and obtains the target addresses corresponding to different storage units 200 in sequence. This can evenly distribute the access requests to multiple storage units 200, reduce the continuous transmission of access requests to the same storage unit 200, and even prevent serial blocking while other storage unit channels are idle.
[0035] In some scenarios, there may be multiple storage unit channels 300 between storage unit 200 and main routing module 120. In this embodiment, main routing module 120 can also determine the optional storage unit channel 300 for transmitting access requests based on the target address and select the target storage unit channel. When there are multiple storage unit channels 300 between the same storage unit 200 and main routing module 120, load balancing of storage unit channels 300 can be achieved, improving the bandwidth utilization of the chip system and reducing serial blocking.
[0036] In this embodiment, the main chip 100 includes a memory channel interleaving unit 111. The memory channel interleaving unit 111 performs a first interleaving transformation on the acquired access request based on the address of the access request and the number of storage units 200 corresponding to the access request, obtaining the target address. This allows the first interleaving transformation of the access request address to be performed by an independent memory channel interleaving unit 111 instead of the main routing module 120, thus making the first interleaving transformation process independent of the main routing module 120 and applicable to different chip system architectures and bus types. After the memory channel interleaving unit 111 obtains the target address through the first interleaving transformation, the main routing module 120 determines the target storage unit to receive the access request based on the target address and sends the access request to the target storage unit, achieving balanced access to the storage units 200. Simultaneously, the main routing module 120 can also determine the storage unit channel 300 for transmitting the access request based on the target address. When multiple storage unit channels 300 exist between the same storage unit 200 and the main routing module 120, load balancing of the storage unit channels 300 is achieved.
[0037] Reference Figure 2 This illustration shows a schematic diagram of another chip system provided in an embodiment of this application. In some implementations of this application, the chip system includes a slave chip 400, which includes a slave routing module 410 and a slave interleaving module 420. The storage unit 200 includes multiple master storage units 210 and multiple slave storage units 220. The slave chip 400 and the master chip 100 are connected through an interconnect channel 310. The master chip 100 and the master storage units 210 are connected through master storage unit channels 320 corresponding to each master storage unit 210. The slave chip 400 and the corresponding slave storage unit 220 can be fully interconnected through multiple slave storage unit channels 330 (i.e., any slave storage unit channel). Each of the 330 modules can access any of the slave storage units 220. The number of storage units 200 includes the number of master storage units 210 and the number of slave storage units 220. The interleaving module 420 is configured such that if the target address corresponds to a slave storage unit 220, or the access request corresponds to a specific target slave storage unit, the interleaving module 420 performs a second interleaving transformation according to the second interleaving parameters. The target memory channel is determined by the slave storage unit channel 330 corresponding to the target slave storage unit. The second interleaving parameters include the number of corresponding slave storage unit channels 330, and the corresponding storage unit channel 300 includes the target memory channel.
[0038] like Figure 2As shown, in some implementations of this application, the master chip 100 can be connected to one or more slave chips 400, and the slave chips 400 and the master chip 100 are connected through interconnect channels 310. Each master storage unit 210 is connected to the master chip 100 through a master storage unit channel 320. Taking the slave chip 400 on the left side of the master chip 100 as an example, it is provided with a slave routing module 410 and a slave interleaving module 420. The slave chip 400 is fully connected to the slave storage unit 220 through multiple slave storage unit channels 330, that is, the slave chip 400 can access any slave storage unit 220 through multiple slave storage unit channels 330.
[0039] In some embodiments of this application, after receiving an access request, the master chip 100 can first identify whether the access request corresponds to a specific target slave storage unit (i.e., the number of storage units that the access request can correspond to is 1). If the access request corresponds to a specific target slave storage unit, the master chip 100 can send the access request to the slave chip 400 connected to the specific target storage unit without performing a first interleaving transformation. If the access request does not correspond to a specific target slave storage unit, the master chip 100 performs a first interleaving transformation on the address of the access request according to the address of the access request and the first interleaving parameters to obtain the target address. If the target address corresponds to slave storage unit 220, or the access request corresponds to a specific target slave storage unit, the master routing module 120 needs to send the access request to the slave storage unit 220 corresponding to the target address (i.e., the target storage unit is the target slave storage unit). If the slave storage unit 220 is connected to the slave chip 400 through multiple fully interconnected slave storage unit channels 330, the access request can be sent to the target slave storage unit via one of the multiple slave storage unit channels.
[0040] If the storage cell 200 and the storage cell channel 300 are not in a one-to-one correspondence, the bandwidth balance of the storage cell channel 300 will also affect the bandwidth balance of the entire chip system. That is, if a large number of access requests are allocated to the same storage cell channel 300, even if there are multiple available storage cell channels 300 for the same storage cell 200, using only one of them will still cause an imbalance in the bandwidth of the chip system due to the bandwidth limitation of a single storage cell channel 300, resulting in low bandwidth utilization of the chip system. Therefore, if the storage cell 200 and the storage cell channel 300 are not in a one-to-one correspondence, the storage cell channels 300 can be interleaved to reduce the decrease in the bandwidth utilization of the chip system.
[0041] In this embodiment, the storage cell channel 300 includes a slave storage cell channel 330. The slave interleaving module 420 performs a second interleaving transformation on the address of the access request based on the address of the access request and a second interleaving parameter to obtain a second interleaving transformation result. Based on the second interleaving transformation result, the target memory channel is determined among the slave storage cell channels 330 connected to the target slave storage cell. When the slave interleaving module 420 receives an access request sent by the master routing module 120, it transmits the access request to the target slave storage cell via the target memory channel, thereby achieving load balancing among the multiple slave storage cell channels 330. The multiple slave storage cell channels 330 can transmit different access requests in parallel, reducing the degradation of chip system bandwidth utilization.
[0042] For example, when the master chip 100 receives multiple access requests, all of which are read operations targeting a specific slave storage unit 220, the slave interleaving module 420 performs a second interleaving transformation on the address of the access request based on the address and second interleaving parameters. This allows different storage unit channels 300 to be determined for the multiple access requests. The access requests are then transmitted and executed to the same slave storage unit 220 through these different storage unit channels 300, reading data and achieving load balancing across the storage unit channels 300. Simultaneously, since access requests can be sent to the same slave storage unit 220 via different storage unit channels 300, it is not necessary to use a single storage unit channel 300 to transmit multiple access requests, enabling parallel transmission of access requests to the same slave storage unit 220.
[0043] It should be noted that, as Figure 2 As shown, if some of the slave memory cells 220 are connected to the slave core 400 only through a single slave memory cell channel 330 (e.g.) Figure 2 The right-hand slave chip 400 is connected to the slave memory cell 220 through a one-to-one corresponding slave memory cell channel 330. If the target memory cell is connected to the slave memory cell 220 through a unique slave memory cell channel 330, then the slave memory cell channel 330 connecting the target memory cell is unique and no second interleaving transformation is required.
[0044] In other implementations of this application, the slave chip 400 may also be provided with an external communication I / O interface for directly receiving external access requests. In specific cases, the slave chip 400 obtains the access request and forwards it to the master chip 100, or actively initiates an access request to the master chip 100 or an external source. When the slave chip 400 obtains an external access request through its own I / O interface, if the access request corresponds to the slave storage unit 220 connected to the slave chip 400 that currently receives the access request, the slave chip 400 that receives the access request can directly perform the second interleaving transformation and transmit the access request to the corresponding slave storage unit 220 through the slave storage unit channel 330 determined by the second interleaving transformation, without having to send the access request to the master chip 100.
[0045] In some implementations of the embodiments of this application, the chip system includes a slave chip 400, the master interleaving module 110 further includes an interconnect channel interleaving unit 112, and the storage unit 200 includes a plurality of master storage units 210 and a plurality of slave storage units 220. The interconnect channel interleaving unit 112 is configured such that: if the slave chip 400 and the master chip 100 are connected through a plurality of interconnect channels 310; then the interconnect channel interleaving unit 112 performs a third interleaving transformation according to a third interleaving parameter, and determines a target interconnect channel from the corresponding interconnect channels 310. The third interleaving parameter includes the number of corresponding interconnect channels 310, and the corresponding storage unit channel 300 includes the target interconnect channel.
[0046] like Figure 2 As shown, the main interleaving module 110 is also provided with an interconnect channel interleaving unit 112. In order to ensure that the transmission bandwidth between the slave core 400 and the main core 100 is not less than the bandwidth provided by the multiple slave storage units 220 connected to the slave core 400, the slave core 400 can be connected to the main core 100 through multiple interconnect channels 310.
[0047] Depend on Figure 2 It is understood that the storage cell channel 300 in this embodiment may also include an interconnect channel 310. Similar to the storage cell channel 330 described above, the interconnect channel 310 can be interleaved to balance the bandwidth of the interconnect channel and reduce the decrease in the bandwidth utilization of the chip system.
[0048] In this embodiment, if multiple interconnect channels 310 exist between slave chip 400 and master chip 100, and the address of the access request corresponds to slave storage unit 220, the interconnect channels 310 and the target storage unit are not in a one-to-one correspondence. After the request is transmitted to slave chip 400 through any interconnect channel 310, it can access any slave storage unit connected to the corresponding slave chip 400. The interconnect channel interleaving unit 112 performs a third interleaving transformation on the address of the access request based on the address of the access request and the third interleaving parameters, thereby obtaining the third interleaving transformation result, and determines the target interconnect channel from the interconnect channels 310 corresponding to the target address based on the third interleaving transformation result. The master routing module 120 can transmit the access request to the slave chip 400 connected to the target interconnect channel, and the target slave storage unit channel of the slave chip 400 can transmit it to the target storage unit, thereby achieving load balancing of multiple interconnect channels 310. Different interconnect channels can transmit different access requests in parallel, and reduce the decrease in chip system bandwidth utilization.
[0049] In practical implementation, when the chip system receives an access request, it can first perform a first interleaving transformation on the address of the access request based on the address information of the access request and the first interleaving parameters to obtain the target address. If the slave memory cell corresponding to the target address corresponds to multiple slave memory cell channels 330, the target memory channel can be determined by performing a second interleaving transformation; if the slave memory cell corresponding to the target address corresponds to one slave memory cell channel 330, then no second interleaving transformation is required. If the slave memory cell corresponding to the target address corresponds to multiple interconnect channels 310, the target interconnect channel can be determined by performing a third interleaving transformation; if the slave chip where the slave memory cell corresponding to the target address is located corresponds to one interconnect channel 310, then no third interleaving transformation is required. Figure 2 For example, when the target access request can select all primary storage units 210 and all secondary storage units 220, the target storage unit can be obtained according to the first interleaving transformation. If the target storage unit is a primary storage unit 210, it can be accessed directly. If the target storage unit is a storage unit of the left secondary core, the corresponding target storage unit number can be obtained according to the first interleaving transformation. Based on the number, the secondary core number it belongs to can be determined, and then it can be determined that there are multiple interconnect channels 310 between the secondary core 400 and the primary core 400, and multiple secondary storage unit channels 330 between the secondary core 400 and the target storage unit. In this case, the second and third interleaving transformations are required. If the target storage unit is a secondary storage unit of the right secondary core, only the third interleaving transformation is required.
[0050] In some implementations of this application, the memory channel interleaving unit 111 includes a first address partitioning unit and a first number calculation unit; the first address partitioning unit divides the address of the access request into a first calculated bit and a first reserved bit, wherein the first calculated bit is used to determine the target storage unit, and the first reserved bit is used to represent a part of the access address of the target storage unit (i.e., the first reserved bit of the access address of the target storage unit); the first number calculation unit groups the first calculated bit, performs summation on all groups to obtain a feature value, and determines the number of the target storage unit based on the feature value; wherein the target address includes the access address and the number of the target storage unit.
[0051] Each memory cell 200 corresponding to the first interleaving transformation can be pre-numbered. The memory channel interleaving unit 111 includes a first address partitioning unit and a first number calculation unit. The first address partitioning unit divides the address of the access request into first calculation bits and first reserved bits. The first number calculation unit groups the first calculation bits according to a preset method, with each group including several bits. It performs a summation operation on all groups to obtain a feature value, determines the number of the target memory cell based on the feature value, and determines the target address based on the number of the target memory cell and the access address. Each slave chip can also be pre-numbered, and the interconnection channel between each slave chip and the master chip can also be pre-numbered. As an example, performing a summation operation on all groups specifically involves converting the bits of each group into numerical values, for example, converting them into decimal values, and then summing the converted values to obtain the characteristic value. For example, after grouping the first calculation bit, four groups are obtained, namely
[0111] ,
[1111] ,
[0011] , and
[1011] . The values obtained by converting each group into decimal values are 7, 15, 3, and 11, respectively, and the characteristic value is 7+15+3+11=36.
[0052] For example, a chip system includes 17 memory cells 200. The 17 memory cells 200 can be assigned corresponding numbers 00 to 16. If the target memory cell is determined to be numbered 10 based on the feature value, then the memory cell 200 with the number 10 is determined as the target memory cell.
[0053] To ensure efficient processing of the access request address and target address at the hardware level, and to avoid information loss or conflict between the access request address and the target address, the access request address can have the same number of bits as the target address. Since the first reserved bit is obtained from the access request address, its width can be preset, while the width of the feature value and the first calculated bit are not necessarily identical. If only the feature value and the first reserved bit are used as the target address, it may result in a mismatch between the widths of the access request address and the target address.
[0054] Reference Figure 3This illustrates an address mapping diagram provided in an embodiment of this application, such as... Figure 3 As shown, the first calculation bit can be further divided into a first part and a second part, and the number of bits in the first part of the first calculation bit is consistent with the number of bits in each group of feature value calculations. The second part of the calculation bit can participate in the group calculation of feature values and also serve as a second reserved bit, passed through to the target address. The target address is obtained based on the target memory cell number determined by the feature value, the second part of the first calculation bit, and the first reserved bit. This allows the chip system to generate the target address efficiently and reliably, avoiding information loss or conflict between the address of the access request and the target address. By ensuring that the first reserved bit does not participate in the calculation, continuous data can be located in different memory cells as much as possible, while data stored in the same memory cell can be kept as continuous as possible. The second part of the computation bit participates in the grouping calculation of feature values and also serves as part of the target address. This allows it to adapt to various large-step access patterns. High-bit flipping also ensures balanced access, which is significantly better than using a power of two for direct interleaving of low bits. This is because when high bits are flipped and low bits remain unchanged, the same memory unit is easily accessed frequently (e.g., the memory unit to be accessed is determined by certain fixed address bits, such as 64B interleaving granularity, 4-bit interleaving, which only uses bits 6-7 of the address to determine which memory unit to go to. When the service memory access characteristics show a large step size within a certain period of time, such as greater than 256B, all access requests will go to the same memory unit).
[0055] In some implementations of this application, each interleaving parameter further includes interleaving granularity, and the influencing factors of interleaving granularity include bus width and cache line size; the first address partitioning unit divides the address of the access request into first computed bits and first reserved bits according to the interleaving granularity, wherein the number of bits of the first computed bits is greater than or equal to log2 (interleaving granularity), and the number of bits of the first reserved bits is less than or equal to log2 (interleaving granularity); the first numbering calculation unit groups the first computed bits according to the number of storage units that the access request can correspond to. Interleaving granularity is the smallest unit of data allocated to different memory cells for reading, fetching, and writing, as defined by the interleaving method. In some cases, if no interleaving transformation is performed, the interleaving granularity can be the capacity of a single memory cell. In some embodiments, the smallest data unit is represented as... Byte. For example, if m is 7 and the interleaving granularity is 128 bytes, then each 128 bytes of data is allocated to a different storage unit, and 128 bytes of data allocated to the same storage unit are not contiguous, but the addresses of 128 bytes of data that are close to each other are as contiguous as possible within the same storage unit. For example, if there are 17 storage units, then the first and 18th data items can both be allocated to storage unit numbered 00 according to the interleaving method, and their addresses within that storage unit are contiguous.
[0056] The interleaving granularity is determined by the bus width and the cache line size. The bus width refers to the bit width of the data path between the main chip 100 and the storage unit 210 or the number of bytes that can be transferred each time. The cache line size refers to the size of the data block transferred from the main storage unit 210 to the cache at one time.
[0057] For example, the interleaving granularity can be greater than or equal to the bus width, ensuring that a single data transfer does not require crossing memory cells. A cache line must cover at least one interleaving granularity. The cache line size is an integer multiple of the bus size, allowing a cache line of data to be transferred completely once or multiple times via the bus. Assuming G = interleaving granularity (bytes), W = bus width (bytes per transfer), and L = cache line size (bytes), then the following constraints apply: G = k1 × W = k2 × L, where k1 and k2 are positive integers. In one example, the interleaving granularity = 64 bytes, the bus width = 32 bytes, and the cache line size = 64 bytes.
[0058] For example, factors influencing interleaving granularity also include memory access characteristics. Memory access characteristics refer to the behavior patterns of accessing storage units, such as large-step or small-step access. If large-step access is used, the interleaving granularity value should be increased to reduce the probability of access requests continuously falling on the same storage unit / channel. As an example, the selectable range of interleaving granularity can be determined based on the address width of the data access requests issued upstream and the bus width downstream.
[0059] The address width used to select storage unit 200 in the data access request sent from upstream starts at the address bit plus the address range, and cannot cross the interleaving granularity boundary. The corresponding address width cannot exceed the interleaving granularity; if it does, multiple data packets need to be generated.
[0060] Taking writing as an example, if the calculation is only based on the predefined number of bits, it might cause two data packets that should be written to different memory units to be written to consecutive addresses in the same memory unit. This would lead to errors when reading these two data packets later. As an example, the interleaving granularity can be determined to be consistent with the corresponding address bit width. For example, if the interleaving granularity is 64B, then the number of bits in the first reserved bit can be 6 bits, where 64 = 2 to the power of 6.
[0061] The smaller the interleaving granularity, the better it is for parallel storage cell channels, but it cannot be smaller than the storage cell channel bit width, otherwise the storage cell channel bandwidth utilization will be insufficient. At the same time, for storage cells, the more continuous the address, the higher the bandwidth utilization, and scattered addresses will lead to a decrease in the bandwidth utilization of the storage cell.
[0062] As an example, suppose the width of a single bus bit in the downstream bus is X, and the interleaving granularity can be an integer multiple of X. For example, the interleaving granularity is 2X (i.e., m=2X). If the interleaving granularity is less than 2X, the memory cell latency hiding capability is poor and the bus utilization is low; if it is greater than 2X, the memory cell response latency perceived by the master die 100 to access requests increases.
[0063] Based on the determined interleaving granularity, the bit length constraints for the first computed bit and the first reserved bit can be determined. Specifically, the number of bits in the first computed bit is not less than m, and the number of bits in the first reserved bit is not greater than m. As an example, the number of bits in the first reserved bit can be the same as m.
[0064] When grouping the first calculation bit as described above, the number of groups is obtained by taking the logarithm of 2 based on the number of storage units corresponding to the access request and rounding up. The first calculation bit can be a multiple of the number of groups.
[0065] For example, if the number of storage units corresponding to an access request is 17, then the number of groups is log2(17) rounded up, resulting in 5 groups, meaning each group consists of 5 bits. If the number of groups is less than the value determined in this way (logarithm of the number of storage units corresponding to the access request divided by 2 and rounded up), it may cause data that should have been in different storage units to be allocated to the same data. If the number of groups is greater than the value determined in this way, the computational pressure on the hardware circuit is relatively high, increasing costs. By determining the number of groups in this way, and then performing group summation to obtain the characteristic value, and determining the number of the target storage unit based on the characteristic value, the interleaving effect is balanced, which can adapt to access modes with arbitrarily large step sizes. High-bit flips can also achieve balanced access, rather than being concentrated in the same storage unit.
[0066] The embodiments of this application determine the interleaving granularity by using bus width and cache line size, thereby enabling the division of the address of the access request into first computed bits and first reserved bits based on the interleaving granularity. This can avoid data access errors when passing through the target storage unit and improve the utilization rate of storage unit channel bandwidth.
[0067] In some implementations of the embodiments of this application, the first number calculation unit includes: The first numbering calculation subunit is configured to: perform a modulo operation between the feature value and the number of corresponding storage units to obtain the number of the target storage unit.
[0068] The modulo operation refers to determining the remainder after dividing the characteristic value by the number of storage units. For example, if the characteristic value is 19 and the corresponding number of storage units is 17, then the result of the modulo operation is 2.
[0069] The first calculation bit is divided into multiple groups, each group is converted into a value, the values of all groups are summed to obtain the feature value, the number of corresponding storage units is used as the interleaving number, and then the feature value is modulo the interleaving number to obtain the hash value, which is used as the number of the target storage unit.
[0070] For example: After grouping the first calculation bit, four groups are obtained:
[0111] ,
[1111] ,
[0011] , and
[1011] . Each group is converted into a value based on the corresponding bit. Taking decimal as an example, the converted values for each group are 7, 15, 3, and 11. Summing all the group values yields a characteristic value of 36. Since the corresponding number of storage units is 16, taking the modulo of 16 with 26 yields a hash value of 10, thus determining the target storage unit number as 10.
[0071] This method allows for the accurate determination of the target storage unit number while using minimal computing resources, thereby enabling balanced access to the storage units.
[0072] For example, if the access request address is 26 bits, the interleaving granularity is 64 bytes, and there are 17 corresponding storage units, then the number of groups is 5. The first reserved bit is 6 bits, and the high-order 20 bits are divided into 4 groups of 5 bits each. These groups are then summed to obtain a feature value. Taking the feature value modulo 17 yields a hash value. This hash value can be used as the corresponding target storage unit number, represented by 5 bits in the high-order bits of the target address. The middle 15 bits are the original address data, which is directly passed through. The first reserved bit is the last 6 bits, which are also directly passed through. In other words, the first 5 bits of the target address are the number part, and the last 21 bits are the address part corresponding to the target storage unit, i.e., the access address.
[0073] In some implementations of this application, the main interleaving module 110 includes a mapping storage subunit 113, which is connected to the memory channel interleaving unit 111. The mapping storage subunit 113 is configured to pre-store the mapping relationship between feature values and the number of the target storage unit, so that it can be accessed by the first number calculation unit and the number of the target storage unit can be directly obtained according to the feature values.
[0074] For example, the mapping relationship between feature values and target storage unit numbers can be stored in the mapping storage sub-unit 113 in advance. When a new access request is received, after calculating the corresponding feature value, the target storage unit number can be determined based on the feature value and the mapping relationship, without having to recalculate the target storage unit number based on the feature value. This embodiment of the application can reduce the amount of calculation and quickly obtain the corresponding target storage unit number by storing mapping information.
[0075] For example, if the summation of the address blocks in the access request yields a feature value of 32, corresponding to 5 storage units, and the modulo operation result is 2 (the target storage unit's number), with storage unit numbers ranging from 0 to 4, then the corresponding storage unit is 2 (the target storage unit). The mapping information can be a mapping table. By pre-stored the mapping relationship between the feature value and the target storage unit's number—that is, the mapping relationship from 5 to 32—it can be stored in the mapping table to quickly obtain the target storage unit's number.
[0076] In some implementations of the first aspect, the interleaving module 420 is configured to: determine the number of the target memory cell channel 330 based on the feature value and the number of memory cell channels 330; and determine the target memory channel based on the number of the target memory cell channel 330; or, the interconnect channel interleaving unit 112 is configured to: determine the number of the target interconnect channel based on the feature value and the number of interconnect channels; and determine the target interconnect channel based on the number of the target interconnect channel. Each slave memory cell channel 330 can be pre-assigned a unique corresponding number to the current slave core. During the second interleaving transformation, the slave interleaving module 420 can calculate the feature value in the manner described above, and then perform a modulo operation based on the feature value and the number of slave memory cell channels 330 to generate the number of the target slave memory cell channel 330, and determine the target memory channel from the slave memory cell channels 330 corresponding to the number of the target slave memory cell channel 330.
[0077] Each interconnect channel can be pre-assigned a unique number. During the third interleaving transformation, the interconnect channel interleaving unit 112 can calculate the characteristic value as described above, and then perform a modulo operation based on the characteristic value and the number of interconnect channels to generate the target interconnect channel number, thus determining the interconnect channel corresponding to the target interconnect channel number as the target interconnect channel. For example, the first, second, and third interleaving transformations can be performed relatively independently; that is, for the same access request, one or more of the three interleaving transformations may occur. Taking the CPU writing data to an unspecified storage unit as an example, the main routing module 120 can first perform the first interleaving transformation to obtain the target storage unit number as 7. If target storage unit 7 is a slave storage unit, then the main routing module 120 can perform the third interleaving transformation, and the slave interleaving module can perform the second interleaving transformation, respectively determining the target interconnect channel and the target memory channel. The access request is transmitted and executed according to the target interconnect channel → target memory channel → target storage unit.
[0078] To reduce computational load, when performing the second and third interleaving transformations, if the first interleaving transformation has already been performed, the feature value calculated during the first interleaving transformation for the same access request can be read. The modulus value can then be adjusted based on the number of available memory channels / interconnect channels, eliminating the need to repeatedly calculate the feature value and improving the processing efficiency of the second and third interleaving transformations. The first interleaving transformation aims to obtain the target memory unit and its access address, while the second and third interleaving transformations are both aimed at determining the corresponding access channels. Therefore, the second and third interleaving transformations only require adjusting the modulus value based on the first interleaving transformation.
[0079] For example, in embodiments of this application, the number of bits K1 of the first portion of the first computation bit can be determined based on the number N of storage units / channels selected from the first interleaving transformation, the second interleaving transformation, or the third interleaving transformation, where K1 = log2(N) , · This indicates rounding up, and then determining the number of bits K3 in the first reserved bit based on the interleaving granularity. As shown above, K3 ≤ log2 (interleaving granularity). The sum of the number of bits K2 in the second part of the first computed bit and the number of bits K3 in the first reserved bit can be obtained by taking the logarithm of the single storage capacity and rounding up. That is, for the number of bits K2 in the second part of the first computed bit and the number of bits K3 in the first reserved bit, (K2 + K3) = log2(C) C represents the capacity of a single storage unit (if the capacities of the storage units are inconsistent, then C represents the capacity of the largest storage unit). Therefore, the number of bits in the second part of the first calculation bit can be obtained as K2 = K2 = log2(C) -K3, thus obtaining the address of the access request (one address corresponds to 1 byte of address space), and the total number of bits in the target address S = K1 + K2 + K3 = log2(N) + log2(C) For example, if a single storage unit has a capacity of 2GB ( B), then (K2+K3)= log2(C) =31 bits.
[0080] In this embodiment, by performing a modulo operation based on the feature value and the number of storage cell channels 330, the target storage cell channel 330 number is generated. This allows for accurate determination of the target storage cell number with minimal computational resources, thereby determining the target memory channel to achieve balanced access to the target memory channel. Similarly, by performing a modulo operation based on the feature value and the number of interconnect channels, the target interconnect channel number is generated. This allows for accurate determination of the target interconnect channel number with minimal computational resources, thereby determining the target interconnect channel to achieve balanced access to the target memory channel.
[0081] The interleaving module 420 can also pre-store the mapping relationship between feature values and the channel numbers of the memory cells on the slave chip, thereby enabling the target memory channel number to be quickly determined without performing modulo operations when the feature values are obtained, reducing the amount of computation. The master interleaving module 110 can also pre-store the mapping relationship between feature values and the target interconnect channel numbers, thereby enabling the target interconnect channel number to be quickly determined without performing modulo operations when the feature values are obtained, reducing the amount of computation required to determine the transformation of the target interconnect channel. That is, the mapping storage sub-unit can store the mapping relationship required for interconnect channel interleaving, so that it can be accessed by the interconnect channel interleaving unit and the target interconnect channel number can be directly obtained based on the feature values; the mapping storage sub-unit can be set in the slave interleaving module, storing the mapping relationship required for memory channel interleaving, so that it can be accessed by the number calculation unit for memory channel interleaving calculation and the target memory channel number can be directly obtained based on the feature values.
[0082] In some implementations of the embodiments of this application, the chip system further includes: an interconnect chip, one side of which is connected to the main chip 100 through an L2C (L2 Cache) channel, and the other side is connected to a slave chip 400 through one or more interconnect channels. The number of L2C channels is the same as the number of slave storage units 220 connected to the corresponding slave chip 400, which is beneficial to the global bandwidth balance of the entire chip system. At this time, if necessary, a fourth interleaving transformation can be performed according to the number of L2C channels.
[0083] The L2C channel is a connection channel used to access L2C. In some implementations of this application, interconnect chips can be set in the chip system, and the number of L2C channels is the same as the number of slave memory units 220. This reduces the bandwidth between the master chip 100 and the slave chip 400, and sets bandwidth redundancy between the slave chips 400. Even if the slave memory unit 220 connected to the slave chip 400 fails, the master chip 100 can still operate with the original bandwidth, improving the stability of the chip system.
[0084] In some embodiments, the chip system may be a computing card including a GPGPU, consisting of a GPU master chip and multiple slave chips, used for large model computations.
[0085] Reference Figure 4 This illustration shows a schematic diagram of another channel interleaving method for a chip system provided in an embodiment of this application. The channel interleaving method is applicable to the chip system as described in the above embodiment. The chip system includes a main interleaving module and a main routing module interconnected with each other. The channel interleaving method includes: Step 401: The main interleaving module obtains the address of the access request and performs a first interleaving transformation on the address of the access request based on the address of the access request and the first interleaving parameters to obtain the target address.
[0086] The first interleaving parameter includes the number of storage units that can correspond to the access request; Step 402: The main routing module transmits an access request to the target storage unit through the corresponding storage unit channel based on the target address.
[0087] In some implementations of the embodiments of this application, the chip system includes a slave chip, the slave chip includes a slave interleaving module, the memory unit includes multiple master memory units and multiple slave memory units, the slave chip and the master chip are connected through interconnect channels, the master chip and the master memory unit are connected through master memory unit channels corresponding one-to-one with each master memory unit, the slave chip and the corresponding slave memory unit are fully interconnected through multiple slave memory unit channels, and the number of memory units includes the number of master memory units and the number of slave memory units; the method further includes: When the interleaving module performs a second interleaving transformation based on the second interleaving parameters when the target address corresponds to a slave memory cell or the access request corresponds to a specific target slave memory cell, the target memory channel is determined from the slave memory cell channel corresponding to the target slave memory cell. The second interleaving parameters include the number of corresponding slave memory cell channels, and the corresponding memory cell channels include the target memory channel.
[0088] In some implementations of the embodiments of this application, the chip system includes slave chips, the master interleaving module further includes interconnect channel interleaving units, the memory unit includes multiple master memory units and multiple slave memory units, and the method further includes: In the case where each slave chip and master chip are connected through multiple interconnect channels, the interconnect channel interleaving unit performs a third interleaving transformation according to the third interleaving parameter to determine the target interconnect channel from the corresponding interconnect channels. The third interleaving parameter includes the number of corresponding interconnect channels, and the corresponding memory cell channel includes the target interconnect channel.
[0089] In some implementations of this application, the target address includes the access address of the target storage unit and the number of the target storage unit. The memory channel interleaving unit includes a first address partitioning unit and a first number calculation unit; the number of the target storage unit is determined through the following steps: The address of the access request is divided into a first calculated bit and a first reserved bit; wherein, the first calculated bit is used to determine the target storage unit, and the access address of the target storage unit includes the first reserved bit.
[0090] The first calculation bits are grouped, and bitwise summation is performed on all groups to obtain the feature value. The target memory cell number is then determined based on the feature value.
[0091] In some implementations of the embodiments of this application, each interleaving parameter also includes interleaving granularity, and the influencing factors of interleaving granularity include bus width and cache line size; The address of the access request is divided into a first computed bit and a first reserved bit, including: dividing the address of the access request into a first computed bit and a first reserved bit according to the interleaving granularity; Grouping the first computation bit includes: grouping the first computation bit according to the number of storage units that can correspond to the access request. In some implementations of this application, determining the number of the target storage unit based on feature values includes: The number of the target storage unit is obtained by performing a modulo operation between the feature value and the number of corresponding storage units.
[0092] In some implementations of the embodiments of this application, the method further includes: The mapping relationship between pre-stored feature values and the number of the target storage unit is accessed by the first number calculation unit and the number of the target storage unit is obtained directly based on the feature values.
[0093] In some implementations of this application, performing a second interleaving transformation based on the second interleaving parameters to determine the target memory channel from the slave storage cell corresponding to the target slave storage cell includes: determining the number of the target slave storage cell based on the feature value and the number of slave storage cell channels; and determining the target memory channel based on the number of the target slave storage cell channel. Alternatively, a third interleaving transformation can be performed based on the third interleaving parameters to determine the target interconnecting channel from the corresponding interconnecting channels, including: determining the number of the target interconnecting channel based on the characteristic value and the number of interconnecting channels; and determining the target interconnecting channel based on the number of the target interconnecting channel. In some implementations of the embodiments of this application, the chip system further includes: Interconnect chip: One side of the interconnect chip is connected to the master chip via an L2C channel, and the other side is connected to a slave chip via one or more interconnect channels; the number of L2C channels is the same as the number of slave storage units connected to the corresponding slave chip.
[0094] It should be noted that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0095] This application also discloses a computer program product, including a computer program that, when run, causes the channel interleaving method of the chip system as described in the foregoing embodiments to be executed.
[0096] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit it. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A chip system, characterized in that, include: A main die and multiple memory cells, wherein at least one memory cell channel is provided between the main die and each memory cell, and the main die includes: The main interleaving module includes a memory channel interleaving unit, which is configured to: obtain the address of the access request, and perform a first interleaving transformation on the address of the access request according to the address of the access request and a first interleaving parameter to obtain a target address, wherein the first interleaving parameter includes the number of storage units that the access request can correspond to; The main routing module is connected to the main interleaving module, and the main routing module is configured to transmit the access request to the target storage unit through the corresponding storage unit channel according to the target address.
2. The chip system according to claim 1, characterized in that, The chip system includes a slave chip, which includes a slave routing module and a slave interleaving module. The storage unit includes multiple master storage units and multiple slave storage units. The slave chip and the master chip are connected through interconnect channels. The master chip and the master storage unit are connected through master storage unit channels corresponding to each master storage unit. The slave chip and the corresponding slave storage unit are fully interconnected through multiple slave storage unit channels. The number of storage units includes the number of master storage units and the number of slave storage units. The interleaving module is configured to: if the target address corresponds to a slave storage unit, or the access request corresponds to a specific target slave storage unit; The interleaving module then performs a second interleaving transformation based on the second interleaving parameters, and determines the target memory channel from the slave storage cell channel corresponding to the target slave storage cell. The second interleaving parameters include the number of corresponding slave storage cell channels, and the corresponding storage cell channels include the target memory channel.
3. The chip system according to claim 1, characterized in that, The chip system includes slave chips, the master interleaving module further includes interconnect channel interleaving units, and the memory unit includes multiple master memory units and multiple slave memory units. The interconnect channel interleaving unit is configured such that: if the slave core and the master core are connected through multiple interconnect channels; The interconnect channel interleaving unit then performs a third interleaving transformation based on the third interleaving parameters to determine the target interconnect channel from the corresponding interconnect channels. The third interleaving parameters include the number of corresponding interconnect channels, and the corresponding storage unit channel includes the target interconnect channel.
4. The chip system according to any one of claims 1-3, characterized in that, The memory channel interleaving unit includes a first address partitioning unit and a first numbering calculation unit; The first address partitioning unit divides the address of the access request into a first calculated bit and a first reserved bit, wherein the first calculated bit is used to determine the target storage unit, and the access address of the target storage unit includes the first reserved bit; The first numbering calculation unit groups the first calculation bits, sums all groups to obtain feature values, and determines the number of the target storage unit based on the feature values; The target address includes the access address and the number of the target storage unit.
5. The chip system according to claim 4, characterized in that, Each interleaving parameter also includes interleaving granularity, which is influenced by factors such as bus width and cache line size; The first address partitioning unit divides the address of the access request into a first computed bit and a first reserved bit according to the interleaving granularity; The first numbering calculation unit groups the first calculation bits according to the number of storage units that can correspond to the access request.
6. The chip system according to claim 4, characterized in that, The first numbering calculation unit includes: The first numbered calculation subunit is configured as follows: The number of the target storage unit is obtained by performing a modulo operation between the feature value and the number of corresponding storage units.
7. The chip system according to claim 4, characterized in that, The main interleaving module includes: The mapping storage subunit, connected to the memory channel interleaving unit, is configured to: pre-store the mapping relationship between feature values and the number of the target storage unit, so that it can be accessed by the first number calculation unit and the number of the target storage unit can be directly obtained according to the feature values.
8. The chip system according to claim 2 or 3, characterized in that, The interleaving module is configured to: determine the target slave storage cell channel number based on the feature value and the number of slave storage cell channels; and determine the target memory channel based on the target slave storage cell channel number. Alternatively, the interconnection channel interleaving unit of the main interleaving module is configured to: determine the number of the target interconnection channel based on the feature value and the number of interconnection channels; and determine the target interconnection channel based on the number of the target interconnection channel.
9. The chip system according to claim 3, characterized in that, The chip system also includes: An interconnect chip has one side connected to a master chip via an L2C channel, and the other side connected to a slave chip via one or more interconnect channels; the number of L2C channels is the same as the number of slave storage cells connected to the corresponding slave chip.
10. A channel interleaving method for a chip system, characterized in that, The channel interleaving method is applicable to the chip system as described in any one of claims 1-9, wherein the chip system includes a main interleaving module and a main routing module interconnected with each other; The channel interleaving method includes: The main interleaving module obtains the address of the access request, and performs a first interleaving transformation on the address of the access request according to the address of the access request and the first interleaving parameters to obtain the target address. The first interleaving parameters include the number of storage units that the access request can correspond to. The main routing module transmits the access request to the target storage unit through the corresponding storage unit channel based on the target address.
Citation Information
Patent Citations
Interleaver mapping and dynamic memory management system and method
CN108845958A
Chip address reconstruction method and device, electronic equipment and storage medium
CN115314438A
Chip access method and device, storage medium and electronic equipment
CN115658591A
Nonlinear multi-storage channel data interleaving method and interleaving module
CN115964310A
Data storage system and method and electronic equipment
CN118035137A
Cited By
Memory access method and device of multi-core particle system, electronic equipment and storage medium
CN122195893A