Memory merging device, chip product, computer equipment and data reading method
By introducing a request arbitration module and an instruction control module into the memory merging device, merging requests are generated and data is read in batches, solving the problems of insufficient flexibility and deadlock, and improving the flexibility and efficiency of data reading.
Patent Information
- Application Number
- CN202511440128.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-10
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2045-10-10
AI Technical Summary
Existing memory merging devices lack flexibility during data reading and are prone to deadlock issues, especially when merging cache requests exceeds the number of cache modules.
By introducing a request arbitration module and an instruction control module, a merge request is generated and data is read from the data storage module in batches, avoiding deadlock and improving the flexibility of data reading.
This effectively avoids deadlock issues while improving the flexibility and efficiency of data retrieval.
Smart Images

Figure CN120909989A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the chip technical field, and in particular to a memory merging device, a chip product, a computer device and a data reading method. BACKGROUND
[0002] In modern GPU (Graphics Processing Unit, graphics processor) / AI (Artificial Intelligence, artificial intelligence) parallel computing chip design, a memory merging device is an essential component.
[0003] At present, the memory merging device has insufficient flexibility when reading data. SUMMARY
[0004] Embodiments of the present application provide a memory merging device, a chip product, a computer device and a data reading method. The technical scheme provided by the embodiments of the present application includes the following aspects.
[0005] According to an aspect of the embodiments of the present application, a memory merging device is provided, which includes a request arbitration module, an instruction control module and a data storage module. The request arbitration module is configured to generate Q merging requests corresponding to a data reading instruction sent by a first requestor based on the data reading instruction, each of the merging requests being used to read data in a same row belonging to the data storage module, and Q being a positive integer. The request arbitration module is further configured to send control information to the instruction control module in a case where the Q merging requests meet a condition. The instruction control module is configured to read data corresponding to the Q merging requests from the data storage module in multiple batches according to the control information.
[0006] According to an aspect of the embodiments of the present application, a chip product is provided, which includes the memory merging device as described above.
[0007] According to an aspect of the embodiments of the present application, a computer device is provided, which includes the memory merging device as described above.
[0008] According to an aspect of the embodiments of the present application, a data reading method applied to a memory merging device is provided, the memory merging device including a request arbitration module, an instruction control module and a data storage module. The method includes: The request arbitration module generates Q merged requests corresponding to the data read instruction based on the data read instruction sent by the first requestor, each of the merged requests being used to read data belonging to the same row of the data storage module, and Q being a positive integer; The request arbitration module sends control information to the instruction control module in the case that the Q merged requests meet a condition. The instruction control module reads data corresponding to the Q merged requests from the data storage module in multiple batches according to the control information.
[0009] The technical scheme provided by the embodiments of the present application can bring the following beneficial effects: In the case that the Q merged requests corresponding to the data read instruction meet a condition, the data corresponding to the Q merged requests is divided into multiple batches for reading. This way of reading the data corresponding to the Q merged requests from the data storage module in multiple batches instead of reading all the data corresponding to the Q merged requests at one time helps to improve the flexibility of data reading. BRIEF DESCRIPTION OF DRAWINGS
[0010] Figure 1 is a structural block diagram of a memory merging device provided in a possible implementation manner of the present application; Figure 2 is a structural block diagram of a memory merging device provided in another possible implementation manner of the present application; Figure 3 is a schematic diagram of a request identifier corresponding to a merged request provided in a possible implementation manner of the present application; Figure 4 is a schematic diagram of instruction meta information provided in a possible implementation manner of the present application; Figure 5 is a schematic diagram of instruction dispatch information provided in a possible implementation manner of the present application; Figure 6 is a schematic diagram of information in an information buffer in a storage block provided in a possible implementation manner of the present application; Figure 7 is a flowchart of a data reading method applied to a memory merging device provided in a possible implementation manner of the present application. DETAILED DESCRIPTION
[0011] To make the purpose, technical scheme and advantages of the present application more clear, the embodiments of the present application will be further described in detail below with reference to the drawings.
[0012] The memory coalescing device has a coalescing function and a cache function. Upstream of the memory coalescing device is generally a plurality of requestors, and the requestors generally have a parallel execution unit. A task of M threads executed by the parallel execution unit at a time is referred to as an instruction. It should be noted that M is the maximum capability of an instruction, and an instruction can also have only a few threads valid. The present application only considers read operations of the task type of accessing memory, which need to return read data to the requestor downstream. The memory addresses requested by the M threads are different, and the lengths are the same.
[0013] The request sent by the requestor is referred to as an initial thread request, and the request after the coalescing function is referred to as a coalesced cache request. The coalescing granularity of the coalescing function module is generally equal to the cache line size of the cache, for example, assuming that the cache line size is R bytes.
[0014] An initial thread request (that is, an instruction) corresponds to Q coalesced cache requests according to the coalescing of the M memory addresses. The value of Q corresponding to each instruction is different. The data read by each thread can fall within two cache lines (that is, cross cache lines), for example, the data memory address is 0x1078, and the length is 32 bytes, and R is 128 bytes. Therefore, a thread corresponds to at most 2 coalesced cache requests, and M threads correspond to at most 2M coalesced cache requests.
[0015] For the read request sent by the requestor, when the requestor requires to return data, the memory coalescing device needs to collect the data of the M threads and then return the data to the requestor. Apparently, the current memory coalescing device has insufficient flexibility when reading data.
[0016] In addition, the current memory coalescing device also has a deadlock problem caused by insufficient resources: if the addresses of the Q coalesced cache requests of the M threads of an instruction after coalescing fall into the same bank and the same set of the coalesced cache requests, and the number of the coalesced cache requests is greater than the way number U of the cache function module, the requestor cannot receive all the coalesced cache requests, and the return direction cannot return all the read data to the requestor, thereby causing a deadlock.
[0017] To solve the deadlock problem, a conventional solution is to ensure that the number Q of the merge buffer requests corresponding to any instruction is less than or equal to the number U of the lanes of the cache function module, so that the deadlock can be avoided. Alternatively, some hardware resources are added to solve the deadlock problem.
[0018] Since the number of lanes of the cache is limited by the hardware logic timing, it generally does not exceed 16 or 32. Therefore, the first solution is not always applicable. The second solution requires the addition of dedicated hardware resources, and can only be used when Q is greater than U, resulting in a low utilization rate of hardware resources.
[0019] Therefore, the present application provides a more simple and practical memory merging device, which can not only improve the flexibility of data reading, but also solve the deadlock problem.
[0020] Please refer to Figure 1 which shows a structure block diagram of a memory merging device provided in a possible implementation manner of the present application. The memory merging device 10 includes a request arbitration module 11, an instruction control module 12, and a data storage module 14.
[0021] The request arbitration module 11 is configured to generate Q merge requests corresponding to a data reading instruction based on the data reading instruction sent by a first requestor, each of the merge requests being used to read data belonging to a same row of the data storage module 14, and Q being a positive integer.
[0022] The first requestor can be any requestor. The upstream of the memory merging device 10 can include one or more requestors, and the first requestor can be any one of the one or more requestors. Each requestor generally has a parallel execution unit, and M thread tasks executed by the parallel execution unit at a time can be referred to as an instruction. When the instruction sent by the requestor is used to read data from the data storage module 14 of the memory merging device 10, one instruction sent by each requestor is used to read data corresponding to at most M single-pen requests from the data storage module 14. M is a positive integer. That is, the M threads and the M single-pen requests are in one-to-one correspondence. The value of M can be pre-set, and depends on the number of threads of the parallel execution unit in the requestor, for example, M = 4 or 6 or 8, and the present application embodiment does not limit this. It should be noted that M here is the maximum capacity of an instruction, and all M threads of an instruction can be valid, or only some of the M threads of an instruction can be valid.
[0023] Each single request has a corresponding data storage address and data length, where the data storage address is used to indicate the data requested to be read by the single request, a storage address in the data storage module 14; and the data length is used to indicate the length or amount of data requested to be read by the single request. In general cases, the data storage addresses corresponding to the M single requests are different, and the data lengths are the same.
[0024] In some embodiments, the request arbitration module 11 is configured to receive a data read instruction sent by the first requestor, where the data read instruction is used to read data corresponding to X single requests from the data storage module 14, X being a positive integer; and generate Q merged requests corresponding to the data read instruction according to the data storage addresses corresponding to the X single requests. It should be understood that the X single requests are the single requests actually triggered by the data read instruction, and therefore, X is a positive integer less than or equal to M.
[0025] After receiving the data read instruction, the request arbitration module 11 generates Q merged requests corresponding to the data read instruction according to the data storage addresses corresponding to the X single requests and a merging granularity. In general, the merging granularity is one row of the data storage module 14. The request arbitration module 11 merges multiple single requests that need to read data in the same row of the data storage module 14 to generate one merged request. Therefore, each merged request is used to read data in the same row of the data storage module 14.
[0026] In addition, for one single request, the data requested to be read can be located in one row of the data storage module 14, or can be located in two rows of the data storage module 14, such as in two adjacent rows of the data storage module 14, i.e., there is a cross-row case. Therefore, the data requested to be read by one single request can be split into two merged requests, and the data in the two rows is read by the two merged requests, respectively.
[0027] By generating Q merged requests corresponding to the data read instruction according to the data storage addresses corresponding to the X single requests of the data read instruction, the request merging function is realized, multiple single requests accessing data in the same row are merged, which helps to simplify the data read process and improve the data read efficiency.
[0028] The request arbitration module 11 is further configured to send control information to the instruction control module 12 in a case where the Q merged requests meet a condition. Optionally, the control information is used to instruct the instruction control module 12 to read data corresponding to the Q merged requests from the data storage module 14 in multiple batches.
[0029] Optionally, the condition is related to at least one of the following: the number of merged requests, and the location of data to be read by the merged requests in the data storage module 14.
[0030] Exemplarily, the condition includes that the number of the merge requests is greater than a preset threshold. The preset threshold can be set based on hardware performance of the memory merging device or an empirical value, which is not limited in the present application.
[0031] When the number of the merge requests is large, reading the data corresponding to the Q merge requests from the data storage module 14 in multiple batches instead of reading all the data corresponding to the Q merge requests at one time helps to improve the flexibility of data reading.
[0032] In some embodiments, the condition can be a condition preset for determining whether the data reading instruction has a risk of deadlock. The risk of deadlock refers to that the data corresponding to the Q merge requests cannot be read at one time.
[0033] The data storage module 14 can include one or more banks, and each bank can include one or more sets. Optionally, the condition includes that the number of the merge requests falling in the same set of the same bank of the data storage module in the Q merge requests is greater than the number of ways U of the data storage module, U being a positive integer. When the number of the merge requests falling in the same set of the same bank of the data storage module 14 in the Q merge requests is greater than the number of ways U of the data storage module 14, that is, when the Q merge requests satisfy the condition, it indicates that the data reading instruction has a risk of deadlock.
[0034] When the data reading instruction has a risk of deadlock, the data storage module 14 cannot receive all the merge requests in the request direction, and the full amount of data corresponding to the data reading instruction cannot be returned to the first requestor in the return direction. To solve this problem, in the embodiments of the present application, when the request arbitration module 11 monitors that the data reading instruction has a risk of deadlock, the control information is sent to the instruction control module 12, and the instruction control module 12 reads the data corresponding to the Q merge requests from the data storage module 14 in multiple batches according to the control information.
[0035] The request arbitration module 11 can divide the Q merge requests corresponding to the data reading instruction into multiple batches, and the number of the merge requests falling in the same set of the same bank of the data storage module 14 included in each batch is not greater than (i.e. less than or equal to) the number of ways U, so that the occurrence of the deadlock problem can be avoided.
[0036] In a possible implementation, the request arbitration module 11 sends control information to the instruction control module 12, and the control information is used to instruct the instruction control module 12 to read the data corresponding to the Q merge requests from the data storage module 14 in multiple batches. Alternatively, the request arbitration module 11 sends one piece of control information to the instruction control module 12, and the one piece of control information is used to instruct the instruction control module 12 to read the data corresponding to the Q merge requests from the data storage module 14 in multiple batches.
[0037] In another possible implementation, the request arbitration module 11 can send multiple pieces of control information to the instruction control module 12, and each piece of control information is used to read the data corresponding to one batch of merge requests. Moreover, the multiple pieces of control information are sent one by one, and after sending one piece of control information, the data corresponding to one batch of merge requests corresponding to the piece of control information is read first, and after the reading is completed, the next piece of control information is sent, and so on, until the data corresponding to the Q merge requests corresponding to the data reading instruction is completely read.
[0038] In actual application, whether the request arbitration module 11 sends one piece of control information or multiple pieces of control information to the instruction control module 12 can be preconfigured or determined by the capability of the request arbitration module 11, and the present application does not make any limitation in this regard.
[0039] The instruction control module 12 is configured to read the data corresponding to the Q merge requests from the data storage module 14 in multiple batches according to the control information. Alternatively, the instruction control module 12 is further configured to store the read data corresponding to the merge requests into the data collection module 13.
[0040] In the case where the request arbitration module 11 sends one piece of control information to the instruction control module 12, the instruction control module 12 reads the data corresponding to the Q merge requests from the data storage module 14 in multiple batches after receiving the piece of control information, and stores the data into the data collection module 13.
[0041] In the case where the request arbitration module 11 sends multiple pieces of control information to the instruction control module 12, the instruction control module 12 reads the data corresponding to one batch of merge requests corresponding to each piece of control information from the data storage module 14 after receiving the piece of control information, and stores the data into the data collection module 13. Then, after receiving the next piece of control information, the instruction control module 12 reads the data corresponding to one batch of merge requests corresponding to the piece of control information from the data storage module 14, and stores the data into the data collection module 13. The process is repeated until the data corresponding to the Q merge requests corresponding to the data reading instruction is completely read.
[0042] Alternatively, as Figure 1As shown, the memory merging apparatus 10 further comprises a data collection module 13. The data collection module 13 is configured to provide the full-amount data corresponding to the data read instruction to the first requester after collecting the full-amount data corresponding to the data read instruction, wherein the full-amount data corresponding to the data read instruction comprises data corresponding to the Q merging requests.
[0043] In addition, the Q merging requests can be sent to a data storage module 14 by the request arbitration module 11, and the data storage module 14 stores the data corresponding to each merging request into the data collection module 13.
[0044] In this way, the full-amount data corresponding to the data read instruction is collected and then fed back to the first requester, thereby ensuring the completeness of data reading.
[0045] The technical scheme provided by the embodiment of the present application divides the data corresponding to the Q merging requests into multiple batches for reading in the case that the Q merging requests corresponding to the data read instruction meet the condition. This way of reading the data corresponding to the Q merging requests from the data storage module in multiple batches instead of reading all the data corresponding to the Q merging requests at one time helps to improve the flexibility of data reading.
[0046] In addition, in the case that the number of merging requests falling in the same group of the same storage block of the data storage module in the Q merging requests is greater than the number of paths U of the data storage module, that is, in the case that the data read instruction has a risk of deadlock, the number of merging requests falling in the same group of the same storage block of the data storage module included in each batch is not greater than (i.e. less than or equal to) the number of paths U in the way of reading data in multiple batches, so that the occurrence of the deadlock problem can be avoided.
[0047] For reference Figure 2 which shows a structural block diagram of a memory merging apparatus provided in another possible implementation manner of the present application. The memory merging apparatus 10 comprises a request arbitration module 11, an instruction control module 12, a data collection module 13 and a data storage module 14.
[0048] The functions of the request arbitration module 11, the instruction control module 12, the data collection module 13 and the data storage module 14 can be referred to the description in the above embodiment, which will not be repeated here.
[0049] As Figure 2As shown, it is assumed that there are N requestors in the system, N being a positive integer. For the convenience of description, it is assumed that each instruction of each requestor is a read operation task of M threads. An initial thread request or instruction contains M single-pen requests corresponding to data storage addresses respectively, but some threads can be invalid. In some embodiments, there is thread active information on the interface to identify which threads are valid. The above interface can be understood as a physical or logical communication channel between the requestor and the memory consolidation device 10, and the validity of each thread is dynamically identified by the thread active information.
[0050] In the memory consolidation device 10, in the request direction, the data read instruction sent by any requestor of the N requestors will first enter the request arbitration module 11. The request coalesce unit in the request arbitration module 11 is used to implement the coalesce function, which can be one request coalesce unit for each requestor, or one request coalesce unit shared by N requestors. If it is one request coalesce unit for each requestor, N request coalesce units are before the arbitration unit; if it is one request coalesce unit shared by N requestors, the one request coalesce unit is after the arbitration unit. Among them, the role of the arbitration unit is to coordinate the access rights of multiple requestors to the downstream resources. For example, when multiple requestors simultaneously initiate data read instructions, the arbitration unit can determine which requestor's data read instruction can be processed preferentially through a priority strategy (such as fixed priority, polling, etc.). In addition, if each requestor has a request coalesce unit, the single-pen request is first coalesced into a coalesced request by the request coalesce unit, and then enters the arbitration unit, which helps to reduce the arbitration load (because the number of coalesced requests Q is generally less than the original number of threads M) and improve the arbitration efficiency. If N requestors share one request coalesce unit, the arbitration unit first selects the requestor, and then the request of the requestor enters the shared request coalesce unit for coalescing, which helps to save hardware resources (only one request coalesce unit is needed), but the arbitration unit needs to process more fine-grained requests.
[0051] At the same time, there is one or more hash algorithm units in the request arbitration module 11, and the main function of the hash algorithm unit is to obtain storage block identifier (bank_id), group identifier (set_id) and other information (determine which storage block and which group of the data storage module 14 the coalesced request needs to access through the hash algorithm) according to the storage address corresponding to the input coalesced request (to reduce the access times of the hash algorithm unit, generally the data storage address corresponding to the coalesced request), to reduce the storage block / group conflict of the data storage module 14. In addition, in the embodiments of the present application, it is assumed that the data storage module 14 includes a number of storage blocks B, B being a positive integer.
[0052] As Figure 3As shown, a coalesced request (denoted as coalesced_cache_request) sent by the request arbitration module 11 to the data storage module 14 corresponds to a unique request identifier (denoted as cache_req_id). The cache_req_id mainly includes: a requestor identifier (denoted as requestor_id), an instruction identifier (denoted as instruction_id), a bank identifier (denoted as bank_id), and a sub-request identifier (denoted as sub_req_id). The requestor identifier is used to indicate the first requestor, i.e., the identification information of the first requestor. The requestor identifier is used to distinguish from which requestor (N requestors in total) the data read instruction corresponding to the coalesced request comes. The instruction identifier is used to indicate the data read instruction, i.e., the identification information of the data read instruction. The instruction identifier is used to distinguish which instruction is issued by the requestor corresponding to the coalesced request. The bank identifier is used to indicate the storage bank to be accessed by the coalesced request, indicating which storage bank of the data storage module 14 the coalesced request is sent to. The sub-request identifier is the identification information of the coalesced request, used to distinguish the respective coalesced requests in the data read instruction for accessing the data storage module 14. The sub-request identifier represents a range of 0~2M-1.
[0053] In some embodiments, the request arbitration module 11 is configured to send control information to the instruction control module 12 in a case where the Q coalesced requests meet the condition, or in a case where it is monitored that the data read instruction has a risk of deadlock. Optionally, the control information includes: instruction meta information (denoted as instruction_meta_info) corresponding to the data read instruction and instruction dispatch information (denoted as instruction_dispath_info) corresponding to the data read instruction.
[0054] The instruction meta information corresponding to the data read instruction is used to indicate the location of the data corresponding to at least one single-pen request of the data read instruction in the data storage module 14.
[0055] The instruction dispatch information corresponding to the data read instruction is used to indicate that the Q coalesced requests meet the above condition.
[0056] Optionally, in a case where the above condition is a condition for determining whether the data read instruction has a risk of deadlock, the instruction dispatch information corresponding to the data read instruction is used to indicate whether the data read instruction has a risk of deadlock.
[0057] In one possible implementation, when the arbitration module 11 sends a control message to the instruction control module 12, the control message includes instruction metadata indicating the location of the data corresponding to X individual requests of the data read instruction in the data storage module 14. Here, the X individual requests are all the individual requests corresponding to the data read instruction.
[0058] In another possible implementation, when the arbitration module 11 sends multiple control messages to the instruction control module 12, each control message includes instruction metadata indicating the location of the data corresponding to a portion of a single request for a data read instruction in the data storage module 14. This portion of a single request can be a part of all single requests corresponding to a data read instruction, such as a single request corresponding to a batch merging request.
[0059] In the above manner, when the arbitration module 11 determines that the conditions for Q merging requests are met, it sends the instruction element information and instruction dispatch information corresponding to the data reading instruction to the instruction control module 12. On the one hand, this enables the instruction control module 12 to know that the conditions for the current Q merging requests are met. On the other hand, it enables the instruction control module 12 to know the storage location of the data to be read by the data reading instruction when the conditions for the Q merging requests are met, thereby ensuring the accuracy and completeness of subsequent data collection.
[0060] like Figure 4 As shown, the instruction metadata corresponding to the data read instruction includes: instruction length (denoted as instruction_length) and M thread information (denoted as thread_info).
[0061] The instruction length indicates the data length corresponding to a single request. Specifically, it represents the data length shared by M single requests corresponding to a data read instruction.
[0062] There are M thread information entries, each indicating the position of the data corresponding to a single request in a row of the data storage module 14. M is a positive integer and represents the maximum number of single requests.
[0063] It should be understood that M is the maximum number of single requests. The number of single requests indicated by the instruction metadata included in the control information can be less than M or equal to M. If the number of single requests indicated by the instruction metadata included in the control information is less than M, some of the thread information in the above M thread information can be represented by default values to indicate a single request that is not actually to be read.
[0064] By including the instruction length and the M thread information in the instruction element information, the location of the data corresponding to each single request of the data read instruction in the data storage module 14 can be indicated clearly and accurately.
[0065] Optionally, the instruction element information corresponding to the data read instruction further includes an instruction end flag (denoted as instruction_end_flag). The instruction end flag is used to indicate whether the current instruction element information is the last instruction element information of the data read instruction. In the case where the data corresponding to the same single request is located in two rows of the data storage module 14, the data read instruction corresponds to two instruction element information; otherwise, the data read instruction corresponds to one instruction element information.
[0066] Each data read instruction corresponds to one or two instruction element information. Taking the data read instruction sent by the first requestor as an example, in the M single requests corresponding to the data read instruction, in the case where the data corresponding to the same single request is located in two rows of the data storage module 14, the data read instruction corresponds to two instruction element information; otherwise (i.e. in the case where the data corresponding to the same single request is not located in two rows of the data storage module 14), the data read instruction corresponds to one instruction element information.
[0067] If the data read instruction corresponds to one instruction element information, the instruction end flag in the instruction element information corresponding to the data read instruction is a first value. If the data read instruction corresponds to two instruction element information, the instruction end flag in the first instruction element information corresponding to the data read instruction is a second value, and the instruction end flag in the second instruction element information corresponding to the data read instruction is the first value. The first value and the second value are different. Exemplarily, the first value is 1, and the second value is 0.
[0068] Since the data corresponding to one single request can be stored across rows, each data read instruction corresponds to one or two instruction element information, and the number of instruction element information can be indicated clearly and accurately by the instruction end flag, so as to avoid omission of the instruction element information.
[0069] Optionally, as shown in Figure 4 Each thread information includes a thread active flag (denoted as thread_active_flag), a request basic identification (denoted as cache_req_basic_id), a start byte offset (denoted as start_byte_offset), and a valid length (denoted as valid_length).
[0070] The thread valid flag indicates whether the thread is valid in the instruction element information. For example, in the initial thread request, one of the M threads is invalid. For another example, in the data read instruction, the data requested by one thread crosses the row, and in the second instruction element information corresponding to the data read instruction, the thread valid flags corresponding to the threads that do not cross the row are invalid.
[0071] In some embodiments, for the i-th single request in the at least one single request, if the data corresponding to the i-th single request is located in the k-th row and the k+1-th row of the data storage module 14, i is a positive integer, and k is a positive integer, then in the first instruction element information corresponding to the data read instruction, the i-th thread information is used to indicate the position of the data corresponding to the i-th single request in the k-th row. Specifically, the start byte offset and the valid length in the i-th thread information are used to indicate the position of the data corresponding to the i-th single request in the k-th row. In the second instruction element information corresponding to the data read instruction, the i-th thread information is used to indicate the position of the data corresponding to the i-th single request in the k+1-th row. Specifically, the start byte offset and the valid length in the i-th thread information are used to indicate the position of the data corresponding to the i-th single request in the k+1-th row. The i-th thread information is the thread information corresponding to the i-th single request.
[0072] In the above manner, when the data corresponding to one single request crosses the row, the positions of the data corresponding to the single request in the two rows can be clearly and accurately indicated.
[0073] Each instruction also corresponds to one instruction dispatch information. As shown in Figure 5 The instruction dispatch information corresponding to the data read instruction includes an indication flag, a requestor identifier (requestor_id), an instruction identifier (instruction_id), and a block request count (denoted as bank_req_cnt).
[0074] The indication flag is used to indicate that the Q merged requests meet the condition. Optionally, the indication flag is 1 bit. For example, when the indication flag takes the first value, it indicates that the Q merged requests meet the condition; when the indication flag takes the second value, it indicates that the Q merged requests do not meet the condition; and the first value and the second value are different. Exemplarily, the first value is 1, and the second value is 0.
[0075] The request arbitration module 11 is further configured to set the value of the indication flag to a first value when it is determined that the Q merged requests meet the condition, and set the value of the indication flag to a second value when it is determined that the Q merged requests do not meet the condition.
[0076] Optionally, when the condition is a condition for determining whether the data read instruction is at risk of deadlock, the indication flag can also be referred to as an instruction deadlock flag (denoted as instruction_deadlock_flag), which is used to indicate that the data read instruction is at risk of deadlock. Optionally, the instruction deadlock flag is 1 bit. For example, when the value of the instruction deadlock flag is the first value, it indicates that the data read instruction is at risk of deadlock; when the value of the instruction deadlock flag is the second value, it indicates that the data read instruction is not at risk of deadlock; and the first value and the second value are different. For example, the first value is 1 and the second value is 0.
[0077] The request arbitration module 11 is further configured to set the value of the instruction deadlock flag to the first value when it is monitored that the data read instruction is at risk of deadlock, and set the value of the instruction deadlock flag to the second value when it is monitored that the data read instruction is not at risk of deadlock.
[0078] By setting the value of the indication flag by the request arbitration module 11, it can be clearly and intuitively reflected whether the Q merged requests meet the condition, such as whether there is a risk of deadlock.
[0079] The requestor identifier is used to indicate the first requestor. The instruction identifier is used to indicate the data read instruction. The requestor identifier and the instruction identifier are a unique identifier of one instruction. Since the instruction meta information and the instruction dispatch information are one-to-one correspondence, the requestor identifier and the instruction identifier do not need to be stored in the instruction meta information.
[0080] The B block request counts are used to indicate the number of merged requests received by the B storage blocks included in the data storage module 14, and B is a positive integer. For each of the B storage blocks, there is a corresponding block request count, which is used to identify the number of merged requests of the data read instruction falling on the storage block. The M single-pen requests can correspond to at most 2M merged requests, and the 2M merged requests can fall on the same storage block, so the maximum representation range of the block request count needs to be able to represent 2M. For one instruction, the sum of the B block request counts is equal to Q; the B block request counts are used for subsequent logical judgment of whether the full amount of data corresponding to the instruction is already ready.
[0081] By including the indication flag, the requester identifier, the instruction identifier and the B block request count in the instruction dispatch information, it can be clearly and intuitively reflected based on the indication flag whether the Q merged requests meet the condition, such as whether there is a risk of deadlock, it can be effectively distinguished based on the requester identifier and the instruction identifier between the data read instructions of different requesters and the different data read instructions of the same requester, and it can be judged based on the B block request count whether the full amount of data corresponding to the current instruction has been completely ready, thereby fully ensuring the accuracy and integrity of the data read.
[0082] In some embodiments, in a case where the data read instruction corresponds to one piece of instruction element information, the data read instruction corresponds to one piece of instruction dispatch information. In a case where the data read instruction corresponds to two pieces of instruction element information, the data read instruction corresponds to two pieces of instruction dispatch information, wherein the two pieces of instruction dispatch information include one piece of valid instruction dispatch information and one piece of invalid instruction dispatch information.
[0083] Optionally, the instruction control module 12 includes a first storage unit and a second storage unit, the first storage unit is used to store the instruction element information, and the second storage unit is used to store the instruction dispatch information. Optionally, the first storage unit is built using SRAM (Static Random-Access Memory), because the number of bits of one layer of information is large. Optionally, the second storage unit is built using register resources, which is convenient for the operation of other module logics.
[0084] In the above manner, the depths of the storage containers of the instruction element information and the instruction dispatch information are consistent, i.e., the depths of the above first storage unit and the second storage unit are consistent. This can facilitate the storage and reading of the instruction element information and the instruction dispatch information, and also helps to reduce the amount of information in the instruction element information, such as not needing to include the requester identifier and the instruction identifier in the instruction element information.
[0085] In some embodiments, the instruction dispatch information corresponding to the data read instruction further includes a valid flag (valid_flag) used to indicate whether the current instruction dispatch information is valid.
[0086] Since one instruction corresponds to at most two pieces of instruction meta information, that is, two layers in the first storage unit, and one instruction corresponds to only one piece of valid instruction dispatch information, that is, one layer in the second storage unit, in order to make the instruction meta information and the instruction dispatch information correspond to each other (the information in a layer in the first storage unit and the information in the corresponding layer in the second storage unit belong to the information of the same instruction), a 1-bit valid flag is used in the instruction dispatch information to indicate whether the instruction dispatch information in the current layer is valid; if one instruction corresponds to two layers of instruction meta information, then in the second storage unit, one layer of instruction dispatch information is valid information and one layer of instruction dispatch information is invalid information. Among them, the valid flag in the valid instruction dispatch information indicates valid, and the valid flag in the invalid instruction dispatch information indicates invalid.
[0087] By designing the valid flag in the instruction dispatch information, the validity of the instruction dispatch information can be effectively distinguished, the depths of the storage containers of the instruction meta information and the instruction dispatch information are consistent, and thus the storage and reading of the above two pieces of information are simplified.
[0088] Next, the instruction meta information and the instruction dispatch information are introduced and described in combination with two examples.
[0089] Example 1: For example, M=4, R=32 byte, the bank number B=4 of the cache, and the way number U=4. The M threads of one instruction are only threads 0, 1 and 2, and thread 3 is invalid; the M data storage addresses (memory addr) are 0x1000, 0x1004, 0x1008 and an invalid address in sequence; and the request length shared by the M threads is equal to 4 byte. It can be found that the three valid memory addrs+request length all fall in the same 32 byte cache line and do not cross the cache line, so the instruction corresponds to one piece of instruction_meta_info, and instruction_end_flag=1; instruction_length=4.
[0090] The four thread_info are in sequence as follows: thread_active_flag=1, cache_req_basic_id is assumed to be 0x10, start_byte_offset=0, and valid_length=4; thread_active_flag = 1, cache_req_basic_id is the same as threadO, 0x10, start_byte_offset = 0x4, valid_length = 4; thread_active_flag = 1, cache_req_basic_id is the same as threadO, 0x10, start_byte_offset = 0x8, valid_length = 4; thread_active_flag = 0, cache_req_basic_id, start_byte_offset, valid_length are all default values.
[0091] This instruction corresponds to an instruction_meta_info, so instruction_dispath_info only occupies one layer, and valid_flag is valid. The number of coalesced_cache_request after merging Q = 1, Q < U, so instruction_deadlock_flag = 0. The sum of B bank_req_cnt is equal to 1.
[0092] Example 2: M=4, R=32 byte, cache's bank number B=4, way number U=4. M threads of 1 instruction are all active, M data memory addresses are 0x101C, 0x201D, 0x301E and 0x4000 in turn, and the request length shared by M threads is equal to 8 byte. We can find that 3 of 4 memory addresses + request length cross the cache line; 0x101C+8 byte occupies cache line 0x1000 and 0x1020, and the effective length corresponding to the two cache lines is 4 byte in turn; 0x201D+8 byte occupies cache line 0x2000 and 0x2020, and the effective length corresponding to the two cache lines is 3 byte and 5 byte in turn; 0x301E+8 byte occupies cache line 0x3000 and 0x3020, and the effective length corresponding to the two cache lines is 2 byte and 6 byte in turn; 0x4000+8 byte occupies 1 cache line 0x4000. Therefore, the instruction corresponds to two pieces of instruction_meta_info.
[0093] The instruction_end_flag of the 1st piece of instruction_meta_info is 0; and the instruction_length is 8. The 4 thread_info are as follows in turn: thread_active_flag = 1, cache_req_basic_id is assumed to be 0x10, start_byte_offset = 0x1C, and valid_length = 4; thread_active_flag = 1, cache_req_basic_id is assumed to be 0x20, start_byte_offset = 0x1D, and valid_length = 3; thread_active_flag = 1, cache_req_basic_id is assumed to be 0x30, start_byte_offset = 0x1E, and valid_length = 2; thread_active_flag = 1, cache_req_basic_id is assumed to be 0x40, start_byte_offset = 0x0, valid_length = 8.
[0094] The instruction_end_flag of the 2nd instruction_meta_info is 1; instruction_length = 8. The 4 thread_info are in turn: thread_active_flag = 1, cache_req_basic_id is assumed to be 0x11, start_byte_offset = 0x0, valid_length = 4; thread_active_flag = 1, cache_req_basic_id is assumed to be 0x21, start_byte_offset = 0x0, valid_length = 5; thread_active_flag = 1, cache_req_basic_id is assumed to be 0x31, start_byte_offset = 0x0, valid_length = 6; thread_active_flag = 0, cache_req_basic_id, start_byte_offset, valid_length are all default values.
[0095] This instruction corresponds to two instruction_meta_info, so instruction_dispath_info occupies 2 layers, and valid_flag is valid and invalid respectively. The number of coalesced_cache_request after merging Q = 7. If the number of Q coalesced_cache_request falls in the same bank and the same set is greater than U, then instruction_deadlock_flag = 1; otherwise instruction_deadlock_flag = 0. The sum of B bank_req_cnt is equal to 7.
[0096] The instruction meta information and the instruction dispatch information are stored in the first storage unit and the second storage unit of the instruction control module 12. The first storage unit is a 1 read 1 write SRAM, that is, in the same cycle, one layer can be read and one layer can be written.
[0097] In some embodiments, the request arbitration module 11 is further configured to, for any one of the Q merged requests, if the merged request is sent to the first storage block in the data storage module 14, add 1 to the value of the block request count corresponding to the first storage block. The instruction control module 12 is further configured to, after receiving a data ready signal corresponding to the merged request sent by the first storage block, subtract 1 from the value of the block request count corresponding to the first storage block, the data ready signal being used to indicate that the data corresponding to the merged request is in a state that can be read. For a detailed description of the process, please refer to the following. By using the above-mentioned manner, the values of the block request counts corresponding to each storage block in the data storage module 14 are maintained, which can ensure that the data corresponding to the Q merged requests of the data read instruction is read completely, and omission is avoided.
[0098] In some embodiments, the request arbitration module 11 is further configured to, if it is determined that the Q merged requests meet the condition, enter a grant-locked state, the grant-locked state being a state of only processing the data read instruction; and exit the grant-locked state after the data collection module 13 collects the full amount of data corresponding to the data read instruction.
[0099] In the request arbitration module 11, if it is determined that the Q merged requests meet the condition, such as a deadlock risk of the data read instruction, the grant of the arbitration unit needs to be kept to the first requestor, so that the Q merged requests corresponding to the data read instruction of the first requestor are all sent to the downstream storage block, and the data corresponding to the Q merged requests is sent to the data collection module 13, and then the grant is released. By using the above-mentioned manner, it can be ensured that the full amount of data of a single data read instruction is read completely. For specific processing details, please refer to the description hereinafter.
[0100] As Figure 6As shown, for each storage block, the information buffer (denoted as isched buffer) in the storage block stores the merge request and corresponding auxiliary information sent by the upstream, and the information in each layer of the information buffer mainly includes 1-bit valid information (denoted as valid), request identification (denoted as cache_req_id), set identification (set id) and way identification (way id) information, and 1-bit data ready information (denoted as data_ready). If the sector function is provided in the data storage module 14, a sector demand bit (denoted as sector_need_bit) is also needed to indicate the sector information required by the merge request. The request identification is the {requestor_id, instruction_id, bank_id, sub_req_id} information described above. One merge request must be allocated one cacheline, and the set id and way id are the corresponding cacheline information. The data ready information indicates whether the data of the corresponding cacheline is currently valid / ready, and if the cache is hit, the data ready information is 1 (indicating valid / ready), and if the cache is not hit, the data ready information cannot be set to 1 until the data is returned to the storage block.
[0101] Multiple merge requests can correspond to the same cacheline, and vice versa, that is, one cacheline can correspond to multiple layers in the information buffer. After the data of one cacheline is ready, the data ready information of multiple layers in the corresponding information buffer will be set to 1.
[0102] After the 1-layer data ready information in the information buffer is set to 1, the data ready signal generation logic is triggered, and one piece of data ready signal (denoted as data_ready_signal) information is sent to the instruction control module 12, and the information at least includes {requestor_id, instruction_id} information.
[0103] The instruction control module 12 can receive data ready signal information from B storage blocks at the same time, match the instruction dispatch information stored in the valid layer of the second storage unit according to the sent {requestor identification (denoted as requestor_id), instruction identification (denoted as instruction_id)} information, and perform a corresponding minus 1 operation on the B block request count in the matched instruction dispatch information. If it is found that the B block request counts are all reduced to 0, it means that all the data of all the threads of the data read instruction is ready and is stored in the downstream storage blocks. At this time, the data collection completion detection & read data control logic in the instruction control module 12 will trigger two actions at the same time, one action is to send a data read signal (denoted as read_cache_data_signal) to the read data control module of the B storage blocks, and the data read signal at least includes {requestor_id, instruction_id} information; the other action is to read the instruction element information corresponding to the first layer or the second layer from the first storage unit, and send the read instruction element information to the data collection module 13. After the two actions are completed, the instruction element information and the instruction dispatch information corresponding to the data read instruction in the first storage unit and the second storage unit of the instruction control module 12 can be released. If the data of multiple instructions are ready and meet the conditions, the data collection completion detection & read data control logic will select one instruction to trigger the above two actions.
[0104] After the read data control module of the storage block receives the data read signal, it searches for all the merged requests corresponding to {requestor_id, instruction_id} in the information buffer to obtain cacheline information (set id / way id information), and then initiates a read cacheline operation to the data control module. After the cacheline data is taken out from the Data SRAM through the data pipeline (DataPipe), it is sent to the data collection module 13.
[0105] The data collection module 13 receives cacheline data from B storage blocks, and a data packing buffer can be generally set in the module. The cacheline data is packed together in a certain format according to the instruction element information, and then returned to the corresponding requestor.
[0106] The value of Q corresponding to one instruction exceeds the pipeline stage number of bankid / setid that can be monitored by the request arbitration module 11 to merge the cache request, so the request arbitration module 11 needs to monitor whether the instruction deadlock flag (taking the instruction deadlock flag as an example below) needs to be set to 1 (set to 1 represents the risk of deadlock) for many times. When the request arbitration module 11 sends the first merge request of one instruction downstream, it needs to open a layer of space in the second storage unit, set the valid flag of the instruction dispatch information to 1 (indicating that the information of the current layer is valid), assign {requestor_id, instruction_id}, and set the instruction deadlock flag to the initial value 0. When the request arbitration module 11 sends any merge request downstream, it needs to update the value of the B block request count in the instruction dispatch information of the corresponding instruction in the second storage unit (i.e., the add 1 operation) synchronously.
[0107] In some embodiments, the request arbitration module 11 is configured to send, after entering the authorization lock state, a first number of merge requests corresponding to a data read instruction to the corresponding storage block in the data storage module 14, the first number being less than or equal to the number of lanes U of the data storage module 14. The instruction control module 12 is configured to send, in a case where it is determined that the data corresponding to the first number of merge requests are all in a readable state, a data read signal to the corresponding storage block in the data storage module 14, the data read signal being configured to request to read the data corresponding to the merge requests. The data collection module 13 is configured to collect the data corresponding to the first number of merge requests sent by the storage block in the data storage module 14.
[0108] In one batch, by reading the data corresponding to the first number of merge requests, it is possible to avoid the number of merge requests processed at a time being greater than the number of lanes U of the data storage module 14, thereby avoiding the occurrence of deadlock problems.
[0109] In addition, the instruction control module 12 can determine, based on the data ready signal sent by the storage block in the data storage module 14, which data corresponding to the merge requests are in a readable state and which data corresponding to the merge requests are not in a readable state.
[0110] In the above manner, for the data corresponding to each batch of merge requests, in a case where it is determined that all of them are in a readable state, the data collection module 13 is notified to collect the data corresponding to the batch of merge requests, ensuring the completeness of the collection of the data corresponding to each batch of merge requests.
[0111] In some embodiments, the data collection module 13 is further configured to send, after collecting the data corresponding to the first number of merge requests, indication information to the request arbitration module 11, the indication information being used to indicate that the data corresponding to the first number of merge requests has been collected. The request arbitration module 11 is further configured to, in a case where there is an unprocessed merge request in the Q merge requests corresponding to the data read instruction, execute again the step of sending the first number of merge requests corresponding to the data read instruction to the corresponding storage block in the data storage module 14. The request arbitration module 11 is further configured to, in a case where all the Q merge requests corresponding to the data read instruction have been processed, exit the authorization lock state.
[0112] Since the Q merge requests corresponding to the data read instruction are generated by the request arbitration module 11, the request arbitration module 11 knows the number of merge requests, and the data collection module 13 can make the request arbitration module 11 determine the number of merge requests for which the corresponding data has been collected by sending the indication information to the request arbitration module 11. The request arbitration module 11 can accurately determine whether all the Q merge requests have been processed by comparing the number of merge requests for which the corresponding data has been collected with the total number Q of merge requests corresponding to the data read instruction.
[0113] In the above manner, for the scenario of reading data in batches for the full amount of data corresponding to a single data read instruction, it can be ensured that the full amount of data corresponding to the single data read instruction is completely read.
[0114] If the request arbitration module 11 monitors that the data read instruction is in the above-mentioned deadlock risk situation, the instruction deadlock flag is set to 1, and the arbitration logic in the request arbitration module 11 is blocked (that is, enters the authorization lock state); and the instruction meta information of the data read instruction cached in the request arbitration module 11 needs to be immediately written into the first storage unit, and the instruction meta information at this time may be incomplete, and the information of some threads is not stored in the instruction meta information. In order to simplify the processing flow, the instruction control module 12 waits for all instructions before the data read instruction to be completed (the data is returned to the corresponding requester), and then the data collection completion detection & read data control logic selects the data read instruction (of course, the premise is that the B block request count is reduced to 0, and the data ready condition is met), at this time, the data collection completion detection & read data control logic in the instruction control module 12 triggers two actions at the same time, sends a data read signal to the storage block, and then the storage block returns the cacheline data to the data collection module 13. When the data collection module 13 monitors that all the data of the data read instruction has been received, it will notify the request arbitration module 11 of the request direction to unlock the arbitration logic; but this batch of data cannot be returned to the first requester at this time, because the subsequent data has not been collected. After the arbitration logic is unlocked, the next batch of merged requests corresponding to the data read instruction is sent to the storage block, and then the above-mentioned process is repeated, and the first storage unit and the second storage unit in the instruction control module 12 re-open space to store the information of the data read instruction once. Finally, this batch of cacheline data is also moved to the data collection module 13. In order to simplify the processing, if the request arbitration module 11 monitors that the instruction deadlock flag of the data read instruction has been assigned to 1, that is, there is a deadlock risk situation, when the merged requests corresponding to the data read instruction are sent to the downstream storage block, the arbitration logic is temporarily blocked again, and after the multiple batches of cacheline data of the data read instruction are moved to the data collection module 13, the arbitration logic in the request arbitration module 11 is unlocked. Through the above method, the deadlock risk problem described above can be easily solved; and without additional hardware storage resources, whether the instruction deadlock flag is 1 or not, the request and response (data read) path is the same, and the utilization rate of hardware resources is high.
[0115] In addition, in the method, the B storage blocks are independent, and thus the instruction control module 12 needs to be placed outside each storage block. The interaction delay between the instruction control module 12 and the B storage blocks is relatively large. In some embodiments, the B storage blocks can be integrated into one module, and the parallel capability of the cache is implemented by using a manner that there are B hit detection / data control sub-modules in the module. At this time, the function of the instruction control module 12 can also be implemented in the module, so that the interaction delay between the instruction control module 12 and the B storage blocks can be reduced.
[0116] The following is an embodiment of the method of the present application. For details not disclosed in the embodiment of the method of the present application, please refer to the above embodiments.
[0117] Please refer to Figure 7 which shows a flowchart of a data reading method applied to a memory consolidation device provided in a possible implementation manner of the present application. The components of the memory consolidation device can be referred to the introduction and description in the above embodiments. As shown in Figure 7 , the method can include at least one of the following steps 710-730.
[0118] In step 710, the request arbitration module generates Q consolidation requests corresponding to the data reading instruction based on the data reading instruction sent by the first requestor, and each consolidation request is used to read data belonging to the same row of the data storage module, and Q is a positive integer.
[0119] In step 720, the request arbitration module sends control information to the instruction control module in the case that the Q consolidation requests meet a condition. Optionally, the control information is used to instruct the instruction control module to read the data corresponding to the Q consolidation requests from the data storage module in multiple batches.
[0120] In step 730, the instruction control module reads the data corresponding to the Q consolidation requests from the data storage module in multiple batches according to the control information.
[0121] In some embodiments, the condition includes that the number of consolidation requests falling in the same group of the same storage block of the data storage module in the Q consolidation requests is greater than the number U of paths of the data storage module, and U is a positive integer. The number of consolidation requests falling in the same group of the same storage block of the data storage module included in each batch in the multiple batches is less than or equal to the number U of paths.
[0122] In some embodiments, the control information includes instruction element information corresponding to the data reading instruction and instruction dispatch information corresponding to the data reading instruction. The instruction element information corresponding to the data reading instruction is used to indicate the position of the data corresponding to at least one single request of the data reading instruction in the data storage module. The instruction dispatch information corresponding to the data reading instruction is used to indicate that the Q consolidation requests meet the condition.
[0123] In some embodiments, the instruction meta-information corresponding to the data reading instruction comprises: an instruction length and M thread information.
[0124] The instruction length is used to indicate the data length corresponding to a single request.
[0125] The M thread information is used to indicate the position of the data corresponding to a single request in a row of the data storage module, M is a positive integer, and M is the maximum number of single requests.
[0126] In some embodiments, the instruction meta-information corresponding to the data reading instruction further comprises: an instruction end flag, used to indicate whether the current instruction meta-information is the last instruction meta-information of the data reading instruction. Wherein, in the case that the data corresponding to the same single request is located in two rows of the data storage module, the data reading instruction corresponds to two instruction meta-informations; otherwise, the data reading instruction corresponds to one instruction meta-information.
[0127] In some embodiments, for the i-th single request in at least one single request, if the data corresponding to the i-th single request is located in the k-th row and the k+1-th row of the data storage module, i is a positive integer, and k is a positive integer, then: in the first instruction meta-information corresponding to the data reading instruction, the i-th thread information is used to indicate the position of the data corresponding to the i-th single request in the k-th row; in the second instruction meta-information corresponding to the data reading instruction, the i-th thread information is used to indicate the position of the data corresponding to the i-th single request in the k+1-th row; wherein, the i-th thread information is the thread information corresponding to the i-th single request.
[0128] In some embodiments, the instruction dispatch information corresponding to the data reading instruction comprises: an indication flag, a requester identifier, an instruction identifier and B block request counts.
[0129] The indication flag is used to indicate that the Q merged requests meet the condition.
[0130] The requester identifier is used to indicate the first requester.
[0131] The instruction identifier is used to indicate the data reading instruction.
[0132] The B block request counts are used to indicate the number of merged requests received by the B storage blocks included in the data storage module, B is a positive integer.
[0133] In some embodiments, the method further comprises: setting, by the request arbitration module, the value of the indication flag to a first value if the Q merged requests meet the condition; setting, by the request arbitration module, the value of the indication flag to a second value if the Q merged requests do not meet the condition; wherein the first value and the second value are different.
[0134] In some embodiments, the method further comprises: for any one of the Q merged requests, setting, by the request arbitration module, the value of the block request count corresponding to the first storage block to 1 if the request arbitration module sends the merged request to the first storage block; setting, by the instruction control module, the value of the block request count corresponding to the first storage block to 0 after receiving the data ready signal corresponding to the merged request sent by the first storage block, the data ready signal being used to indicate that the data corresponding to the merged request is in a readable state.
[0135] In some embodiments, the data read instruction corresponds to one piece of instruction dispatch information if the data read instruction corresponds to one piece of instruction element information; the data read instruction corresponds to two pieces of instruction dispatch information if the data read instruction corresponds to two pieces of instruction element information, wherein the two pieces of instruction dispatch information include one piece of valid instruction dispatch information and one piece of invalid instruction dispatch information.
[0136] In some embodiments, the instruction dispatch information corresponding to the data read instruction further comprises: a valid flag, used to indicate whether the current instruction dispatch information is valid.
[0137] In some embodiments, the method further comprises: storing, by the instruction control module, the data corresponding to the Q merged requests in the data collection module; providing, by the data collection module, the full amount of data to the first requestor after collecting the full amount of data corresponding to the data read instruction, the full amount of data including the data corresponding to the Q merged requests.
[0138] In some embodiments, the method further comprises: entering, by the request arbitration module, an authorization lock state if the Q merged requests meet the condition, the authorization lock state being a state for processing only the data read instruction; exiting, by the request arbitration module, the authorization lock state after the data collection module collects the full amount of data corresponding to the data read instruction.
[0139] In some embodiments, the method further comprises: requesting the arbitration module to send, after entering the authorized lock state, a first number of merge requests corresponding to the data read instruction to the corresponding storage block in the data storage module, the first number being less than or equal to the number U. The instruction control module sends a data read signal to the corresponding storage block in the data storage module in a case where the data corresponding to the first number of merge requests are all in a readable state, the data read signal being used to request reading the data corresponding to the merge requests. The data collection module collects the data corresponding to the first number of merge requests sent by the storage block in the data storage module.
[0140] In some embodiments, the method further comprises: the data collection module sending indication information to the request arbitration module after collecting the data corresponding to the first number of merge requests, the indication information being used to indicate that the data corresponding to the first number of merge requests have been collected. The request arbitration module re-executes the step of sending the first number of merge requests corresponding to the data read instruction to the corresponding storage block in the data storage module in a case where there is an unprocessed merge request in the Q merge requests. The request arbitration module exits the authorized lock state in a case where all the Q merge requests have been processed.
[0141] An exemplary embodiment of the present application further provides a chip product, which comprises the memory merging device described above. Optionally, the chip product can be a GPU chip, an AI chip, a TP (Tensor Processing Unit) chip, an NPU (Neural Processing Unit) chip, etc.
[0142] An exemplary embodiment of the present application further provides a computer device, which comprises the memory merging device described above. Optionally, the computer device can be a personal computer, a workstation, a game console, and some mobile devices (such as a tablet computer, a smart phone, etc.), can also be a vehicle terminal device, a smart home device, a smart television, a smart robot, etc., can also be a server, a server cluster, an artificial intelligence computing cluster, a cloud computing cluster, etc., wherein the artificial intelligence computing cluster can also be referred to as an intelligent computing cluster or a smart computing cluster, and the present application does not limit this.
[0143] It should be understood that "multiple" mentioned herein refers to two or more than two. The "and / or" describes the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B can represent the following three cases: A exists alone, A and B exist together, and B exists alone. The character " / " generally represents that the associated objects before and after it are in an "or" relationship. In addition, the step numbers described herein only exemplarily show a possible execution order between steps, and in some other embodiments, the above steps can also be executed in a non-numbered order, such as two different numbered steps being executed simultaneously, or two different numbered steps being executed in an order opposite to that shown in the figure, and the embodiments of the present application are not limited in this regard.
[0144] The above only describes exemplary embodiments of the present application and is not intended to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. A memory merging device, characterized by, The memory merging device comprises a request arbitration module, an instruction control module and a data storage module; The request arbitration module is configured to generate Q merging requests corresponding to a data read instruction sent by a first requestor based on the data read instruction, each of the merging requests being configured to read data belonging to a same row of the data storage module, and Q being a positive integer; The request arbitration module is further configured to send control information to the instruction control module in a case where the Q merging requests satisfy a condition; The instruction control module is configured to read data corresponding to the Q merging requests from the data storage module in multiple batches according to the control information.
2. The memory consolidation device of claim 1, wherein, The condition comprises a number of merging requests falling in a same group of a same memory block of the data storage module in the Q merging requests being greater than a number U of paths of the data storage module, and U being a positive integer; Each of the multiple batches comprises a number of merging requests falling in the same group of the same memory block of the data storage module, and the number is less than or equal to the number U of paths.
3. The memory consolidation device of claim 1, wherein, The control information comprises instruction element information corresponding to the data read instruction and instruction dispatch information corresponding to the data read instruction; The instruction element information corresponding to the data read instruction is configured to indicate a position of data corresponding to at least one single request of the data read instruction in the data storage module; The instruction dispatch information corresponding to the data read instruction is configured to indicate that the Q merging requests satisfy the condition.
4. The memory consolidation device of claim 3, wherein, The instruction element information corresponding to the data read instruction comprises: an instruction length configured to indicate a data length corresponding to the single request; M thread information pieces, each of which is configured to indicate a position of data corresponding to one single request in one row of the data storage module, M being a positive integer, and M being a maximum number of the single requests.
5. The memory consolidation device of claim 4, wherein, The instruction element information corresponding to the data read instruction further comprises: an instruction end flag configured to indicate whether the current instruction element information is the last instruction element information of the data read instruction; wherein, in a case where data corresponding to the same single request is located in two rows of the data storage module, the data read instruction corresponds to two pieces of the instruction element information; otherwise, the data read instruction corresponds to one piece of the instruction element information.
6. The memory consolidation device of claim 5, wherein, For an i-th single request of the at least one single request, if data corresponding to the i-th single request is located in a k-th row and a k+1-th row of the data storage module, i being a positive integer and k being a positive integer, then: an i-th thread information piece in a first piece of instruction element information corresponding to the data read instruction is configured to indicate a position of data corresponding to the i-th single request in the k-th row; an i-th thread information piece in a second piece of instruction element information corresponding to the data read instruction is configured to indicate a position of data corresponding to the i-th single request in the k+1-th row; wherein, the i-th thread information piece is thread information corresponding to the i-th single request.
7. The memory consolidation device of claim 3, wherein, The instruction dispatch information corresponding to the data read instruction comprises: an indication flag, used to indicate that the Q merge requests meet the condition; a requester identifier, used to indicate the first requester; an instruction identifier, used to indicate the data read instruction; B block request counts, used to indicate the number of merge requests received by B storage blocks included in the data storage module respectively, B being a positive integer.
8. The memory consolidation device of claim 7, wherein, The request arbitration module is further configured to: set a value of the indication flag to a first value in a case where it is determined that the Q merge requests meet the condition; set the value of the indication flag to a second value in a case where it is determined that the Q merge requests do not meet the condition; wherein the first value and the second value are different.
9. The memory merging apparatus according to claim 7, wherein the request arbitration module is further configured to, for any one of the Q merge requests, increase a value of a block request count corresponding to a first storage block in the data storage module by 1 in a case where the merge request is sent to the first storage block; the instruction control module is further configured to, after receiving a data ready signal corresponding to the merge request sent by the first storage block, decrease the value of the block request count corresponding to the first storage block by 1, the data ready signal being used to indicate that data corresponding to the merge request is in a readable state.
10. The memory merging apparatus according to claim 3, wherein in a case where the data read instruction corresponds to one piece of instruction element information, the data read instruction corresponds to one piece of instruction dispatch information; in a case where the data read instruction corresponds to two pieces of instruction element information, the data read instruction corresponds to two pieces of instruction dispatch information, wherein the two pieces of instruction dispatch information include one piece of valid instruction dispatch information and one piece of invalid instruction dispatch information.
11. The memory consolidation device of claim 10, wherein, the instruction dispatch information corresponding to the data read instruction further includes: a validity flag, used to indicate whether the current instruction dispatch information is valid.
12. The memory merge device of any of claims 1 to 11, wherein, The memory merging apparatus further includes a data collection module. the instruction control module is further configured to store data corresponding to the Q merge requests into the data collection module; the data collection module is configured to, after collecting full-amount data corresponding to the data read instruction, provide the full-amount data to the first requester, the full-amount data including the data corresponding to the Q merge requests.
13. The memory consolidation device of claim 12, wherein, The request arbitration module is further configured to: enter an authorization lock state in a case where it is determined that the Q merge requests meet the condition, the authorization lock state being a state of only processing the data read instruction; exit the authorization lock state after the data collection module collects full-amount data corresponding to the data read instruction.
14. The memory merging apparatus according to claim 13, wherein the request arbitration module is configured to, after entering the authorization lock state, send a first number of merge requests corresponding to the data read instruction to corresponding storage blocks in the data storage module, the first number being less than or equal to a number U of paths of the data storage module. The instruction control module is configured to, in a case where it is determined that the data corresponding to the first number of merge requests are all in a readable state, send a data reading signal to the corresponding storage block in the data storage module, the data reading signal being configured to request reading of the data corresponding to the merge requests; The data collection module is configured to collect the data corresponding to the first number of merge requests sent by the storage block in the data storage module.
15. The memory merging apparatus according to claim 14, wherein The data collection module is further configured to, after collecting the data corresponding to the first number of merge requests, send indication information to the request arbitration module, the indication information being configured to indicate that the data corresponding to the first number of merge requests has been collected completely; The request arbitration module is further configured to, in a case where there is an unprocessed merge request in the Q merge requests, execute again the step of sending the data reading instruction corresponding to the first number of merge requests to the corresponding storage block in the data storage module. The request arbitration module is further configured to, in a case where the Q merge requests have all been processed, exit the authorization locking state.
16. The memory merge device of any of claims 1 to 11, wherein, The request arbitration module is configured to: receive the data reading instruction sent by the first requestor, the data reading instruction being configured to read data corresponding to X single requests from the data storage module, X being a positive integer; generate the Q merge requests corresponding to the data reading instruction according to the data storage addresses corresponding to the X single requests.
17. A chip product, characterized by The chip product comprises the memory merging apparatus according to any one of claims 1 to 16.
18. A computer device, comprising: The computer device comprises the memory merging apparatus according to any one of claims 1 to 16.
19. A data reading method applied to a memory consolidation device, characterized by, The memory merging apparatus comprises a request arbitration module, an instruction control module and a data storage module. The method comprises: The request arbitration module generates Q merge requests corresponding to the data reading instruction based on the data reading instruction sent by the first requestor, each of the merge requests being configured to read data in a same row belonging to the data storage module, Q being a positive integer; The request arbitration module sends control information to the instruction control module in a case where the Q merge requests meet a condition; The instruction control module reads the data corresponding to the Q merge requests from the data storage module in multiple batches according to the control information.
Citation Information
Patent Citations
Storage system, method for reading data from storage system and method for writing data to storage system
CN102023809A
Apparatus and method for improved cache utilization and efficiency
CN110276710A
Instruction cache, instruction cache group and request merging method thereof
CN113867801A
Operation instruction processing method and device, computer equipment and storage medium
CN118331897A
Data reading method for chip, and chip, computer device, storage medium and computer program product
WO2025139618A1