Request access method and system for avoiding access conflict in one period
By introducing front-end processing module and cross switches into the request access system of the graphics processor, deduplication of request processing and division of non-overlapping address intervals, the access conflicts and request stagnation caused by overlapping access by multiple threads to the same cache line are solved, and the effect of reducing the number of cache blocks and improving system efficiency is achieved.
Patent Information
- Application Number
- CN202311698316.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-12
- Publication Date
- 2025-06-20
AI Technical Summary
In a graph processor, multiple threads may perform overlapping access to the same cache line, resulting in access violations and request stalling.
By introducing a front-end processing module and a cross switch in the request access system, deduplication and processing requests from multiple request modules within a period, and divide the address space requested to access into non-overlapping address intervals to avoid access conflicts.
It significantly reduces the number of accesses to cache blocks, reduces the number of cache blocks, avoids access conflicts and request stagnation, and improves the efficiency of the system.
Smart Images

Figure CN120179375A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of signal processing, and more specifically, relates to a request access method and system for avoiding access conflicts within one cycle. Background Art
[0002] As a parallel processor, a Graphics Processing Unit (GPU) includes multiple cores that can execute multiple threads of the same program in parallel. In this design, multiple threads may request data from the same memory hierarchy, so there may be overlapping accesses to the same cache line. Summary of the Invention
[0003] In view of the above-mentioned defects or improvement requirements of the prior art, the present invention provides a request access method and system for avoiding access conflicts within one cycle. When there is a situation of repeated access to the same address or address range within one cycle, it can significantly reduce the number of accesses to cache blocks, and while correspondingly reducing the number of cache blocks, it can avoid access conflicts and request stalls. It can be used in graphics technology and is widely used in scenarios where multiple threads request texture data or general computer buffer data.
[0004] To achieve the above object, according to one aspect of the present invention, a request access system is provided, including: a plurality of request modules, a front-end processing module, a crossbar switch, and one or more cache blocks; The plurality of request modules are used to send access requests; The front-end processing module is used to perform duplicate removal processing on requests from the plurality of request modules within one cycle; The crossbar switch is used to divide the address space to be accessed by requests into non-overlapping address ranges, and output requests corresponding to different address ranges respectively according to the addresses to be accessed by the requests after duplicate removal processing; Each cache block corresponds to a different address range and is used to respond to requests for its corresponding address range.
[0005] In some embodiments, the front-end processing module is used to: when there are no other requests accessing the same address as the current request, pass the current request to the crossbar switch; when there are other requests accessing the same address as the current request, select one request from all requests accessing the same address as the valid request, exclude other duplicate requests, and pass the valid request to the crossbar switch.
[0006] In some embodiments, the crossbar switch further includes a multi-write first-in-first-out memory, and the crossbar switch is used to output requests corresponding to different address ranges respectively through the multi-write first-in-first-out memory according to the addresses to be accessed by the requests after duplicate removal processing.
[0007] In some embodiments, the request access system further includes a plurality of request processing modules and one or more arbitration modules; The plurality of request processing modules correspond one-to-one with non-overlapping address ranges partitioned by a crossbar switch; the crossbar switch is configured to send requests corresponding to different address ranges to corresponding request processing modules respectively; The plurality of request processing modules are configured to collect requests for their corresponding address ranges, and after partitioning the collected requests according to the address information accessed by the requests, pack the requests corresponding to the same address information into a uniquified request for output; One or more arbitration modules correspond one-to-one with one or more cache blocks, and each arbitration module corresponds to a plurality of request processing modules, and is configured to send the uniquified request output by one of its corresponding plurality of request processing modules to its corresponding cache block; One or more cache blocks are configured to make a response according to the received uniquified request.
[0008] In some embodiments, the request processing module is configured to store the address information to be accessed by the request and the request information corresponding to the address information; the request processing module can store multiple pieces of address information, and for each piece of address information, the request processing module can store multiple pieces of request information.
[0009] In some embodiments, each cache block is configured to, after receiving the uniquified request, return response data according to the storage order of the request information in the request processing module in the uniquified request.
[0010] In some embodiments, the request processing module is further configured to number the positions where the request information is stored, and each uniquified request includes all the request information corresponding to a single address information and its position number; each cache block is configured to, after receiving the uniquified request, reorder the response according to the position number of the request information in the uniquified request.
[0011] In some embodiments, the front-end processing module is configured to: for the request after deduplication processing, when there are other duplicate requests, add flag information and duplicate identity information to the request information of the request after deduplication processing; wherein, the flag information is used to indicate whether the request is unique, and the duplicate identity information is the identity information of the request module from which the excluded duplicate request comes.
[0012] In some embodiments, each cache block is further configured to: according to the corresponding request module identity information and duplicate identity information included in the request information in the uniquified request, return response data to the corresponding request module and the duplicate request module.
[0013] According to another aspect of the present invention, there is provided a request access method, including: Divide the address space requested to be accessed into non-overlapping address ranges; Multiple request modules send access requests; Receive requests and perform duplicate removal processing on the requests from the multiple request modules within one cycle; According to the addresses to be accessed by the requests after duplicate removal processing, output the requests corresponding to different address ranges separately; Respond to the requests corresponding to different address ranges separately.
[0014] In some embodiments, performing duplicate removal processing on the requests from the multiple request modules within one cycle includes: when there are no other requests accessing the same address as the current request, passing the current request to the crossbar switch; when there are other requests accessing the same address as the current request, selecting one request from all the requests accessing the same address as the valid request, excluding other duplicate requests, and passing the valid request to the crossbar switch.
[0015] In some embodiments, performing duplicate removal processing on the requests from the multiple request modules within one cycle further includes: for the requests after duplicate removal processing, when there are other duplicate requests, adding flag information and duplicate identity information to the request information of the requests after duplicate removal processing; wherein, the flag information is used to indicate whether the request is unique, and the duplicate identity information is the identity information of the request module from which the excluded duplicate requests come.
[0016] In some embodiments, responding to the requests corresponding to different address ranges separately includes: according to the corresponding request module identity information and duplicate identity information included in the request information of the requests, after the request processing module receives the response to the requests, returning the response data to the corresponding request module and the duplicate request modules.
[0017] In some embodiments, outputting the requests corresponding to different address ranges separately according to the addresses to be accessed by the requests after duplicate removal processing includes: outputting the requests corresponding to different address ranges separately through a multi-write first-in-first-out memory according to the addresses to be accessed by the requests after duplicate removal processing.
[0018] In some embodiments, responding to the requests corresponding to different address ranges separately includes: Collect the requests for the address range, and divide the collected requests according to the address information accessed by the requests; Pack the requests corresponding to the same address information into a unique request; Respond according to the received unique request.
[0019] In some embodiments, responding according to the received unique request includes: returning the response data according to the order of collecting the requests corresponding to the same address information.
[0020] In some embodiments, separately responding to requests corresponding to different address ranges further includes: storing the address information to be accessed by the request and the request information corresponding to the address information; numbering the positions where the request information is stored; and enabling each uniquified request to include all the request information corresponding to a single address information and its position number.
[0021] In some embodiments, responding according to the received uniquified request includes: reordering the response according to the position number of the request information in the uniquified request.
[0022] Generally speaking, compared with the prior art, the above technical solution conceived by the present invention has the following beneficial effects: on the one hand, duplicate requests for accessing the same address within a period are removed; on the other hand, requests for accessing the same address range within a period are buffered by a multi-write first-in-first-out memory, alleviating the conflicts and stalls of requests in the crossbar switch, enabling multiple requests to be written into the crossbar switch simultaneously within the same period. In addition, fundamentally speaking, through the deduplication process by the front-end processing module, the number and bandwidth of requests can be reduced, so the number of required cache blocks can be reduced. The present invention can be used in graphics technology. For example, during ray tracing, many threads need to repeatedly request data from the acceleration structure for traversal. The present invention can also be widely used in scenarios where multi-threaded requests texture data or general-purpose computer buffers data. Description of the Drawings
[0023] Figure 1 is a schematic structural diagram of a request access system; Figure 2 is a schematic structural diagram of a request access system for avoiding access conflicts within a period according to an embodiment of the present invention; Figure 3 is a schematic flowchart of a request access method for avoiding access conflicts within a period according to an embodiment of the present invention. Detailed Embodiments
[0024] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. As those skilled in the art can recognize, the described embodiments can be modified in various different ways without departing from the spirit or scope of the present application. Therefore, the drawings and descriptions are considered to be exemplary in nature and not restrictive.
[0025] Figure 1The request access system shown includes a request module, a crossbar, and cache blocks. The request module is connected to the cache blocks through the crossbar. To complete multiple requests in each cycle, the address space is divided into non-overlapping address ranges through hashing, and the cache is divided into multiple cache blocks. Each cache block processes requests for one address range of the address space. Multiple request modules continuously issue requests to read data from the cache hierarchy. Since any request module can access any address in the memory, similarly, any request module can access any cache block. Therefore, the crossbar is used to effectively enable any request module to access any cache block.
[0026] There are some defects in this structure. First, at least as many cache blocks as the number of request modules are required to ensure the efficiency of request processing. Although the cache space of the cache blocks is divided, there is duplicate management logic for each cache block, which will increase the area cost, and the cost will be higher with more cache blocks. Second, when multiple request modules access the same cache block, cache block conflicts will occur in the crossbar. Specifically, two request modules accessing the same cache block will cause a conflict, and only one of the request modules will be allowed to continue accessing. Even with a good hashing function, this situation cannot be avoided when the number of cache blocks is close to the number of request modules. When all request modules access the same data, these request modules will queue up to access a cache block in sequence, and the problem of access conflict is particularly prominent.
[0027] As Figure 2 shown, the request access system for avoiding access conflicts within one cycle according to an embodiment of the present invention includes multiple request modules, a front-end processing module, a crossbar, multiple request processing modules, multiple arbitration modules, and multiple cache blocks. The request module is used to send access requests. The front-end processing module is used to de-duplicate the requests from multiple request modules within one cycle and then pass them down to the crossbar.
[0028] In some embodiments, the de-duplication processing includes: when there are no other requests accessing the same address as the current request, passing the current request down to the crossbar; when there are other requests accessing the same address as the current request, selecting one request from all the requests accessing the same address as the valid request, excluding other duplicate requests, and passing the valid request down to the crossbar. In some embodiments, the information of the excluded duplicate requests is retained in the information of the valid request so that when the request data is returned from the cache block, the request data can be returned to all the request modules that requested to access the same address.
[0029] In some embodiments, the front-end processing module associates each request with a first direction of the request (e.g. Figure 2Compare all requests (on the left side in []) and if there are requests that are duplicates of this request, mark the duplicate requests and avoid passing the duplicate requests down to the crossbar. In some embodiments, in the case of duplicate requests, the request on the far left is passed down as a valid request to the crossbar.
[0030] In some embodiments, the request information includes the request content, the request address, and the corresponding request module identity information, and the front-end processing module passes the request information down to the crossbar.
[0031] In some embodiments, the front-end processing module adds the identity information of the request module from which the excluded duplicate request comes (referred to as duplicate identity information) to the requests being sent down, that is, after adding the duplicate identity information to the request information, the front-end processing module then passes the request information down to the crossbar. At this time, the request information includes the request content, the request address, the corresponding request module identity information, and the duplicate identity information.
[0032] In some embodiments, the front-end processing module adds information indicating whether the request is unique (referred to as flag information) to the requests being sent down, that is, after adding the flag information to the request information, the front-end processing module then passes the request information down to the crossbar. At this time, the request information includes the request content, the request address, the flag information, and the corresponding request module identity information. For example, add a bit. When this bit is 1, it indicates that this request is unique and there are no excluded requests that are duplicates of this request.
[0033] In some embodiments, the front-end processing module adds flag information and duplicate identity information to the requests being sent down, that is, after adding the flag information and the duplicate identity information to the request information, the front-end processing module then passes the request information down to the crossbar. At this time, the request information includes the request content, the request address, the flag information, the corresponding request module identity information, and the duplicate identity information. For example, add a bit. When this bit is 0, it indicates that this request is not unique and there are excluded requests that are duplicates of this request, and further add the identity information of the request module from which the excluded duplicate requests come.
[0034] The crossbar is used to divide the address space accessed by the requests into non-overlapping address ranges, and based on the address to be accessed by the request sent by the front-end processing module, send the request sent by the front-end processing module to the corresponding request processing module.
[0035] In some embodiments, the crossbar switch further includes multiple write first-in first-out (FIFO) memories. The crossbar switch sends the requests sent by the front-end processing module to the corresponding request processing modules through the multiple write first-in first-out memories according to the addresses to be accessed in the requests sent by the front-end processing module. That is, by adding multiple write first-in first-out memories in the crossbar switch, each write first-in first-out memory corresponds one-to-one to a non-overlapping address range divided by the crossbar switch, and the requests corresponding to different address ranges are respectively output through the corresponding write first-in first-out memories. Therefore, requests with access conflicts are allowed to exist within one cycle, that is, the requests accessing the same address range within one cycle are appended to one output. In some embodiments, the write first-in first-out memories are implemented by registers.
[0036] The request processing modules correspond one-to-one to the non-overlapping address ranges divided by the crossbar switch, and are used to collect the requests for their corresponding address ranges over time. After dividing the collected requests according to the address information accessed by the requests, the requests corresponding to the same address information are packed into a uniquified request (i.e., one request) and passed downwards. Multiple request processing modules correspond to one arbitration module, and the arbitration module corresponds one-to-one to the cache blocks, and is used to send the uniquified request output by one of its corresponding multiple request processing modules to the cache block; the cache block is used to process the requests for the corresponding address ranges of its corresponding multiple request processing modules and make responses according to the received uniquified requests. In some embodiments, the response data comes from the cache block, or the response data comes from the cache block at the upper level of the cache block, or the response data comes from the memory.
[0037] It should be understood that although multiple requests can be written to the same cache block, the write first-in first-out memories alone cannot fundamentally solve the problem, and the processing speed of these requests is still limited by the speed of the cache block. Taking 6 request modules as an example, if there is no deduplication by the front-end processing module and all the request modules continuously request one address, since a single cache block can only process 1 request per cycle, which is significantly less than the 6 requests input per cycle, the write first-in first-out memories will be quickly filled up and stagnate. Therefore, the main function of the write first-in first-out memories is to provide buffering to make the peak and valley values of the number of requests accessing the same address range smoother.
[0038] In some embodiments, the request processing module includes a first storage module and a second storage module. The first storage module is used to store the address information to be accessed by the request, and the second storage module is used to store the request information corresponding to the address information. In some embodiments, the maximum number of address information that the first storage module can store is M. In some embodiments, for each address information, the maximum number of request information that the second storage module can store is N. In some embodiments, all the request information corresponding to one address information is called a group, and one address information and its corresponding group are called an entry.
[0039] In some embodiments, the first storage module is a Context Addressable Memory (CAM), and the second storage module is a Random Access Memory (RAM).
[0040] In some embodiments, after receiving a new request, the request processing module searches in the first storage module to find if there is a matching address information. If there is, and the number of request information corresponding to this address information has not reached N, the request information is stored in the position corresponding to this address information in the second storage module. In some embodiments, the request processing module is used to count the request information corresponding to each address information and store the count value in the second storage module. In some embodiments, the request information is stored in the first empty position in the position corresponding to this address information in the second storage module in sequence, and the count value of the corresponding request information is incremented by 1.
[0041] In some embodiments, after receiving a new request, the request processing module searches in the first storage module to find if there is a matching address information. If there is, and the number of request information corresponding to this address information has reached N, the entry where this address information is located is deleted, and a new entry is created at the positions where the deleted entry is located in the first storage module and the second storage module.
[0042] In some embodiments, after receiving a new request, the request processing module searches in the first storage module to find if there is a matching address information. If there is no such matching address information, and there is still an empty position in the area of the first storage module where the address information is stored, a new entry is created. In some embodiments, after receiving a new request, the request processing module searches in the first storage module to find if there is a matching address information. If there is no such matching address information, and there is no empty position in the area of the first storage module where the address information is stored, an entry is deleted, and a new entry is created at the positions where the deleted entry is located in the first storage module and the second storage module.
[0043] In some embodiments, establishing a new table entry includes: storing the address information of the new request in an empty position in the area of the first storage module that stores address information, storing the request information of the new request in the first empty position corresponding to the address information in the second storage module, and incrementing the count value of the corresponding request information by 1, that is, at this time the count value of the request information is 1.
[0044] In some embodiments, according to the order of table entry establishment, a table entry is deleted. For example, the earliest established table entry is preferentially deleted. In some embodiments, the table entry with the request information reaching the maximum quantity is preferentially deleted. In some embodiments, when there is no table entry with the request information reaching the maximum quantity, the table entry with the most request information is preferentially deleted. In some embodiments, the table entry with the most request information is preferentially deleted; when there are multiple table entries with the most request information, the earliest established table entry among the table entries with the most request information is preferentially deleted.
[0045] In some embodiments, the request information included in a table entry has different request addresses. In some embodiments, the address information stored in the first storage module is the request address, and the matching address information refers to the address information with a predetermined number of high-order address bits being the same as those of the new request. In some embodiments, the address information stored in the first storage module is a predetermined number of high-order address bits of the request address, and the matching address information refers to the address information with the same predetermined number of high-order address bits as the new request.
[0046] In some embodiments, the deleted table entry is sent to the downstream arbitration module for storage for future use when reading the request data.
[0047] In some embodiments, the request processing module further includes a third storage module, and the group in the deleted table entry is stored in the third storage module. In some embodiments, taking the group in the third storage module as a unit, a uniquification request is sent downstream, that is, the uniquification request includes all the request information constituting a group.
[0048] In some embodiments, when the arbitration module receives uniquification requests from multiple corresponding request processing modules simultaneously, according to a predetermined rule, it selects to send the uniquification request sent by one of the request processing modules to its corresponding cache block; otherwise, in the case where they are not sent simultaneously, the arbitration module directly sends the received uniquification request to its corresponding cache block.
[0049] In some embodiments, the third storage module is a First Input First Output (FIFO) memory. All the request information in the uniquification request maintains the original storage order in the second storage module before. After the cache block receives the uniquification request, it returns the response data according to the original storage order of the request information in the second storage module in the uniquification request.
[0050] In some embodiments, the third storage module is a Random Access Memory (RAM). The positions storing the request information in the third storage module are numbered (ID). The request information deleted from the second storage block is stored in the third storage module, so that each uniquification request contains all the request information corresponding to a single address information in the third storage module and its position number. After the cache block receives the uniquification request, it re-orders the response according to the position number of the request information in the uniquification request.
[0051] In some embodiments, after the cache block returns the response data, the request processing module returns the response data to the corresponding request module according to the corresponding request module identity information included in the request information in the uniquification request. In some embodiments, after the cache block returns the response data, the request processing module returns the response data to the corresponding request module and the duplicate request modules according to the corresponding request module identity information and the duplicate identity information included in the request information in the uniquification request. In some embodiments, after the cache block returns the response data, the request processing module confirms whether the request corresponding to the request information is unique according to the flag information included in the request information in the uniquification request. When it is confirmed that the corresponding request is unique, the request processing module returns the response data to the corresponding request module according to the corresponding request module identity information included in the request information in the uniquification request; when it is confirmed that the corresponding request is not unique, the request processing module returns the response data to the corresponding request module and the duplicate request modules according to the corresponding request module identity information and the duplicate identity information included in the request information in the uniquification request.
[0052] In some embodiments, after the cache block returns the response data, the request processing module issues the uniquification request again to re-expand the response.
[0053] In some embodiments, there is one or more arbitration modules. Correspondingly, there is one or more cache blocks.
[0054] Such as Figure 2As shown in the figure, taking six request modules as an example, assuming that the average effective rate of the front-end processing module is 66%, after deduplicating the six requests sent by the six request modules within one cycle, four requests are obtained and passed down to the crossbar switch; further, taking four request processing modules as an example, assuming that each request processing module collects two requests from the front-end processing module on average, the number of cache blocks is halved, each cache block corresponds to two request processing modules, and an arbitration module is added. The crossbar switch divides the address space accessed by the requests into four non-overlapping address ranges, and the four request processing modules correspond one by one to these four non-overlapping address ranges.
[0055] Table 1
[0056] Taking the request processing module 201 as an example, the request processing module 201 includes an up-down addressable memory (CAM), a first random access memory (the first RAM), and a second random access memory (the second RAM). As time goes by, the request processing module 201 collects the requests in its corresponding address range. The address information to be accessed by the requests is stored in the CAM, and the request information corresponding to the address information is stored in the first RAM. Referring to Table 1, the maximum number of address information that the CAM can store is 8, and the maximum number of request information that can be stored in the first RAM is 4, that is, one address information can correspond to at most 4 request information.
[0057] After receiving a new request, the request processing module 201 will look up in the CAM to see if there is a matching address information. Assuming that there is a matching address information 0x11111111, the request information is stored in the first empty position corresponding to this address information 0x11111111 in the first RAM, that is, the position corresponding to ID1, and the count value of the corresponding request information is incremented by 1, so that the value of Count is 2. Assuming that there is a matching address information 0xFEFEFEFE, since the number of request information corresponding to this address information has reached the maximum value of 4, the entry where this address information 0xFEFEFEFE is located is deleted, and then the address information and request information of the new request are stored in the position where this entry is located. Among them, the request information is stored in the position corresponding to ID0, and the value of Count is 1. Assuming that there is no matching address information and the number of address information stored in the CAM has reached the maximum value of 8, then one entry is deleted from these 8 entries, for example, the entry with the most request information and established earlier is deleted, for example, the entry where the address information 0xFEFEFEFE is located is deleted, and then the address information and request information of the new request are stored in the position where this entry is located.
[0058] Number the positions in the second RAM where the request information is stored, and store the groups in the deleted table entries into the second RAM. For example, store the 4 request information corresponding to the address information 0xFEFEFEFE into the second RAM, so that the uniquified request includes these 4 request information and their position numbers in the second RAM, and send the uniquified request to the downstream arbitration module 203. If the request processing module 205 does not send a uniquified request to the arbitration module 203 at the same time, the arbitration module 203 will send the uniquified request from the request processing module 201 to its corresponding cache block 207. After reading the requested data from the cache or memory, the cache block 207 reorders the response according to the position numbers of the request information in the uniquified request.
[0059] In an embodiment of the present invention, the front-end processing module performs deduplication processing on the same requests accessing the same address within one cycle, which can effectively reduce the number of requests and the number of downstream request processing modules. Combining with the multi-write first-in-first-out memory in the crossbar switch to buffer the requests accessing the same address range within one cycle after deduplication processing can avoid access conflicts caused by multiple requests accessing the same cache block, especially the same address within one cycle, and at the same time avoid requests from stalling in the crossbar switch. In addition, by the request processing module combining the requests accessing the same cache block within a short period of time and accessing the cache block once, the number of cache blocks can be further reduced.
[0060] It should be understood that the request processing module and the arbitration module are not necessary. Similar to Figure 1 the situation shown, in some embodiments, the cache blocks correspond one-to-one to non-overlapping address ranges divided by the crossbar switch. The crossbar switch sends the requests sent by the front-end processing module to the cache blocks corresponding to the address ranges, or the crossbar switch sends the requests sent by the front-end processing module to the cache blocks corresponding to the address ranges through the multi-write first-in-first-out memory.
[0061] In some embodiments, the cache block returns response data to the corresponding request module and the duplicate request module according to the corresponding request module identity information and duplicate identity information included in the request information in the received request. In some embodiments, the cache block confirms whether the request corresponding to the request information is unique according to the flag information included in the request information in the received request. When confirming that the corresponding request is unique, the cache block returns response data to the corresponding request module according to the corresponding request module identity information included in the request information in the received request; when confirming that the corresponding request is not unique, the cache block returns response data to the corresponding request module and the duplicate request module according to the corresponding request module identity information and duplicate identity information included in the request information in the received request.
[0062] As Figure 3As shown in the figure, the request access method for avoiding access conflicts within one cycle according to the embodiments of the present invention includes: Step S301: Divide the address space to be requested for access into non-overlapping address ranges; Step S303: Multiple request modules send access requests; Step S305: Receive the requests and perform duplicate removal processing on the requests from the multiple request modules within one cycle.
[0063] In some embodiments, performing duplicate removal processing on the requests from the multiple request modules within one cycle includes: when there are no other requests accessing the same address as the current request, passing the current request to the crossbar switch; when there are other requests accessing the same address as the current request, selecting one request from all the requests accessing the same address as the valid request, excluding other duplicate requests, and passing the valid request to the crossbar switch.
[0064] In some embodiments, performing duplicate removal processing on the requests from the multiple request modules within one cycle further includes: for the requests after duplicate removal processing, when there are other duplicate requests, adding flag information and duplicate identity information to the request information of the requests after duplicate removal processing; wherein, the flag information is used to indicate whether the request is unique, and the duplicate identity information is the identity information of the request module from which the excluded duplicate requests come.
[0065] Step S307: Output the requests corresponding to different address ranges respectively according to the addresses to be accessed by the requests after duplicate removal processing; In some embodiments, outputting the requests corresponding to different address ranges respectively according to the addresses to be accessed by the requests after duplicate removal processing includes: outputting the requests corresponding to different address ranges respectively through a multi-write first-in-first-out memory according to the addresses to be accessed by the requests after duplicate removal processing.
[0066] Step S309: Respond to the requests corresponding to different address ranges respectively.
[0067] In some embodiments, responding to the requests corresponding to different address ranges respectively includes: returning response data to the corresponding request module and the duplicate request modules according to the corresponding request module identity information and duplicate identity information included in the request information of the requests.
[0068] In some embodiments, responding to the requests corresponding to different address ranges respectively further includes: Collect the requests in the address range, and divide the collected requests according to the address information to be accessed by the requests; Pack the requests corresponding to the same address information into a unique request; Respond according to the received unique request.
[0069] In some embodiments, making a response according to the received uniquification request includes: returning response data according to the order of collecting requests corresponding to the same address information.
[0070] In some embodiments, making responses to requests corresponding to different address ranges respectively further includes: storing the address information to be accessed by the request and the request information corresponding to the address information; numbering the positions where the request information is stored; and enabling each uniquification request to include all the request information corresponding to a single address information and its position number. In some embodiments, making a response according to the received uniquification request includes: reordering the response according to the position number of the request information in the uniquification request.
[0071] When the request access method of the embodiment of the present invention is further implemented, reference may be made to the description of the request access system in the foregoing embodiments, and it has the same beneficial effects, which will not be elaborated herein.
[0072] In the description of this specification, the descriptions with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. mean that the specific features, structures, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. Moreover, the specific features, structures, materials or characteristics described may be combined in a suitable manner in any one or more embodiments or examples. In addition, without conflict, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples.
[0073] In addition, the terms "first" and "second" are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present application, "a plurality of" means two or more, unless otherwise specifically defined.
[0074] Any process or method description shown in the flowchart or described in other ways herein may be understood as representing a module, segment, or part of code including one or more (two or more) executable instructions for implementing a specific logical function or process. And the scope of the preferred embodiments of the present application includes additional implementations, where the functions may be executed in a substantially simultaneous manner or in a reverse order according to the functions involved, rather than in the order shown or discussed.
[0075] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or used in combination with these instruction execution systems, apparatus, or devices.
[0076] It should be understood that the various parts of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. All or part of the steps of the method in the above embodiments can be completed by a program instructing relevant hardware, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.
[0077] In addition, each functional unit in the various embodiments of the present application can be integrated into a processing module, or each unit can exist physically alone, or two or more units can be integrated into one module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. When the above integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. The storage medium can be a read-only memory, a magnetic disk, an optical disc, etc.
[0078] As mentioned above, the above are only the specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed in the present application can easily think of various changes or substitutions, and these should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A request access system, characterized in that, Including: Multiple request modules, a front-end processing module, a crossbar switch, and one or more cache blocks; The multiple request modules are used to send access requests; The front-end processing module is used to de-duplicate the requests from the multiple request modules within one cycle; The crossbar switch is used to divide the address space to be accessed by requests into non-overlapping address ranges, and output the requests corresponding to different address ranges respectively according to the addresses to be accessed by the requests after de-duplication processing; Each of the cache blocks corresponds to a different address range and is used to respond to the requests for its corresponding address range.
2. The request access system according to claim 1, characterized in that, The front-end processing module is used to: when there is no other request accessing the same address as the current request, pass the current request to the crossbar switch; when there is another request accessing the same address as the current request, select one request from all the requests accessing the same address as the valid request, exclude other duplicate requests, and pass the valid request to the crossbar switch.
3. The request access system according to claim 1, characterized in that, The crossbar switch further includes a multi-write first-in first-out memory, and the crossbar switch is used to output the requests corresponding to different address ranges respectively through the multi-write first-in first-out memory according to the addresses to be accessed by the requests after de-duplication processing.
4. The request access system according to any one of claims 1 to 3, characterized in that, It further includes multiple request processing modules and one or more arbitration modules; The multiple request processing modules correspond one-to-one to the non-overlapping address ranges divided by the crossbar switch; the crossbar switch is used to send the requests corresponding to different address ranges to the corresponding request processing modules respectively; The multiple request processing modules are used to collect the requests for their corresponding address ranges, and after dividing the collected requests according to the address information to be accessed by the requests, pack the requests corresponding to the same address information into a unique request and output it; The one or more arbitration modules correspond one-to-one to the one or more cache blocks, each of the arbitration modules corresponds to multiple of the request processing modules, and is used to send the unique request output by one of the corresponding multiple request processing modules to the corresponding cache block; The one or more cache blocks are used to respond according to the received unique request.
5. The request access system according to claim 4, characterized in that, The request processing module is used to store the address information to be accessed by the request and the request information corresponding to the address information; the request processing module can store multiple pieces of address information, and for each piece of address information, the request processing module can store multiple pieces of request information.
6. The request access system according to claim 5, characterized in that, Each of the cache blocks is used to return response data according to the storage order of the request information in the request processing module after receiving the unique request.
7. The request access system according to claim 5, characterized in that, The request processing module is further used to number the positions where the request information is stored, and each unique request includes all the request information corresponding to a single piece of address information and its position number; each of the cache blocks is used to re-sort the response according to the position number of the request information in the unique request after receiving the unique request.
8. The request access system according to claim 5, characterized in that, The front-end processing module is used for: when there are other duplicate requests for the request after duplicate removal processing, adding flag information and duplicate identity information to the request information of the request after duplicate removal processing; wherein, the flag information is used to indicate whether the request is unique, and the duplicate identity information is the identity information of the request module from which the excluded duplicate request comes.
9. The request access system according to claim 8, characterized in that, Each of the cache blocks is further used for: according to the corresponding request module identity information and duplicate identity information included in the request information in the uniquified request, returning response data to the corresponding request module and the duplicate request modules.
10. A request access method, characterized in that, Including: Dividing the address space accessed by requests into non-overlapping address ranges; Multiple request modules sending access requests; Receiving requests and performing duplicate removal processing on the requests from the multiple request modules within one cycle; Outputting the requests corresponding to different address ranges respectively according to the address to be accessed by the request after duplicate removal processing; Responding to the requests corresponding to different address ranges respectively.
11. The request access method according to claim 10, characterized in that, Performing duplicate removal processing on the requests from the multiple request modules within one cycle includes: when there are no other requests accessing the same address as the current request, passing the current request to the crossbar switch; when there are other requests accessing the same address as the current request, selecting one request from all the requests accessing the same address as the valid request, excluding other duplicate requests, and passing the valid request to the crossbar switch.
12. The request access method according to claim 11, characterized in that, Performing duplicate removal processing on the requests from the multiple request modules within one cycle further includes: when there are other duplicate requests for the request after duplicate removal processing, adding flag information and duplicate identity information to the request information of the request after duplicate removal processing; wherein, the flag information is used to indicate whether the request is unique, and the duplicate identity information is the identity information of the request module from which the excluded duplicate request comes.
13. The request access method according to claim 12, characterized in that, Responding to the requests corresponding to different address ranges respectively includes: according to the corresponding request module identity information and duplicate identity information included in the request information of the request, returning response data to the corresponding request module and the duplicate request modules.
14. The request access method according to claim 10, characterized in that, Outputting the requests corresponding to different address ranges respectively according to the address to be accessed by the request after duplicate removal processing includes: outputting the requests corresponding to different address ranges respectively through a multi-write first-in-first-out memory according to the address to be accessed by the request after duplicate removal processing.
15. The request access method according to any one of claims 10 to 14, characterized in that, Responding to the requests corresponding to different address ranges respectively includes: Collecting the requests in the address range, and dividing the collected requests according to the address information accessed by the requests; Packing the requests corresponding to the same address information into a uniquified request; Responding according to the received uniquified request.
16. The request access method according to claim 15, characterized in that, Responding according to the received uniquified request includes: returning response data according to the order of collecting the requests corresponding to the same address information.
17. The request access method according to claim 15, characterized in that, Responding to the requests corresponding to different address ranges respectively further includes: storing the address information to be accessed by the request and the request information corresponding to the address information; numbering the positions where the request information is stored; and making each uniquified request include all the request information corresponding to a single address information and its position number; Responding according to the received uniquified request includes: re-sorting the response according to the position number of the request information in the uniquified request.
Citation Information
Cited By
Write data merging system, write data merging method, and chip
CN122547704A