Data request processing method and device, electronic equipment, storage medium and program product

By setting a request buffer corresponding to the cache unit in the GPU and merging data requests, the problem of low data request processing efficiency in a multi-threaded environment is solved, and more efficient cache access and memory management are achieved.

CN120029933APending Publication Date: 2025-05-23MOORE THREADS TECH CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510192110.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

How to improve the efficiency of data request processing of GPUs when processing large amounts of data, especially in multi-threaded environments.

Method used

By setting up a plurality of pending request buffers that correspond one by one to the multiple cache units, the pending data requests allocated to the cache unit are buffered, and the data requests in the pending request buffer are merged to form a merge request to reduce cache access.

Benefits of technology

Improves the efficiency of cache access, allows a single cache unit to process more threads, reduces the frequency of access to cache and memory, and reduces the burden of memory access.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029933A_ABST
    Figure CN120029933A_ABST
Patent Text Reader

Abstract

The invention relates to a data request processing method and device, electronic equipment, a storage medium and a program product. The method comprises the following steps: for any to-be-processed request buffer, merging to-be-processed data requests in the to-be-processed request buffer to obtain at least one first-level merged request, and writing the at least one first-level merged request into a first-level merged request buffer corresponding to the to-be-processed request buffer; for any first-level merging request extracted from any first-level merging request buffer, responding to any second-level merging request in a second-level merging request buffer corresponding to the first-level merging request buffer and meeting a preset merging condition with the first-level merging request; combining the first-level combination request with the second-level combination request to obtain a new second-level combination request; and taking out the second-level merging request from the second-level merging request buffer for processing. According to the invention, the cache access efficiency can be improved on the whole.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular to a data request processing method, a data request processing device, an electronic device, a computer-readable storage medium, and a computer program product. Background Art

[0002] In modern computing architecture, GPU (Graphics Processing Unit) is widely used in graphics rendering, scientific computing and deep learning due to its powerful parallel processing capabilities. A key component in GPU design is multi-level cache. Multi-level cache is designed to reduce the latency when the GPU accesses external storage devices, thereby improving overall performance.

[0003] The multi-level cache inside the GPU is a group of high-speed memories that are arranged in a hierarchy to provide fast data access. These caches are usually divided into at least two levels, such as L1, L2, etc., and each level has its specific access speed and capacity. The design of the multi-level cache allows the GPU to quickly read and write data from the cache when processing large amounts of data, rather than exchanging data from slower external memory every time.

[0004] Unlike a CPU (Central Processing Unit), a GPU often processes multiple program blocks at the same time, and each program block is further divided into multiple threads for processing. How to improve the efficiency of data request processing is a technical problem that needs to be solved urgently. Summary of the invention

[0005] The present disclosure provides a technical solution for processing data requests.

[0006] According to one aspect of the present disclosure, a data request processing method is provided, wherein a cache comprises a plurality of cache units, wherein the plurality of cache units correspond one-to-one to a plurality of pending request buffers, wherein the pending request buffer corresponding to any cache unit is used to buffer pending data requests allocated to the cache unit, and the method comprises:

[0007] For any pending request buffer, merge the pending data requests in the pending request buffer to obtain at least one first-level merged request, and write the at least one first-level merged request into the first-level merged request buffer corresponding to the pending request buffer;

[0008] For any first-level merge request taken out from any first-level merge request buffer, in response to any second-level merge request in a second-level merge request buffer corresponding to the first-level merge request buffer and the first-level merge request satisfying a preset merge condition, merge the first-level merge request with the second-level merge request to obtain a new second-level merge request;

[0009] The second-level merge request is taken out from the second-level merge request buffer for processing.

[0010] In a possible implementation, the preset merging condition includes:

[0011] The data requested by the first-level merge request and the second-level merge request are located in N adjacent cache lines, where N is an integer greater than or equal to 2.

[0012] In a possible implementation, the preset merging condition includes:

[0013] The first-level merge request and the second-level merge request are data requests received within M adjacent clock cycles, where M is an integer greater than or equal to 2.

[0014] In a possible implementation, the preset merging condition includes at least one of the following:

[0015] The first-level merge request and the second-level merge request come from the same request end;

[0016] The first-level merge request and the second-level merge request correspond to the same system-level cache policy;

[0017] The first-level merge request and the second-level merge request correspond to the same consistency strategy.

[0018] In a possible implementation, the method further includes:

[0019] For any first-level merge request taken out from any first-level merge request buffer, in response to the fact that there is no second-level merge request that satisfies the preset merge condition with the first-level merge request in the second-level merge request buffer corresponding to the first-level merge request buffer, the first-level merge request is written into the second-level merge request buffer as a new second-level merge request.

[0020] In a possible implementation, taking out the second-level merge request from the second-level merge request buffer for processing includes:

[0021] In response to the second-level merge request most recently written into the second-level merge request buffer carrying instruction end marker information, the second-level merge request is taken out from the second-level merge request buffer for processing, wherein the instruction end marker information is used to indicate the last data request in the request sequence.

[0022] In a possible implementation, merging the pending data requests in the pending request buffer to obtain at least one first-level merged request, and writing the at least one first-level merged request into a first-level merged request buffer corresponding to the pending request buffer includes:

[0023] In response to the number of pending data requests in the pending request buffer being greater than or equal to L, taking out L pending data requests from the pending request buffer, where L is an integer greater than or equal to 2;

[0024] The L pending data requests are merged to obtain K first-level merged requests, and the K first-level merged requests are written into a first-level merged request buffer corresponding to the pending request buffer, where K is an integer greater than 0 and less than or equal to L.

[0025] In a possible implementation, merging the L to-be-processed data requests to obtain K first-level merged requests includes:

[0026] Among the L data requests to be processed, data requests whose requested data are located in the same cache line are merged to obtain K first-level merged requests.

[0027] In a possible implementation, the method further includes:

[0028] Determine, from a plurality of cache units, a target cache unit corresponding to the data request to be processed;

[0029] The to-be-processed data request is written into a target to-be-processed request buffer corresponding to the target cache unit.

[0030] In a possible implementation, determining, from a plurality of cache units, a target cache unit corresponding to the data request to be processed includes:

[0031] Performing a hash operation on the request address of the data request to be processed to obtain a hash operation result corresponding to the data request to be processed;

[0032] According to the hash operation result, a target cache unit corresponding to the to-be-processed data request is determined from a plurality of cache units.

[0033] According to one aspect of the present disclosure, a data request processing device is provided, wherein a cache comprises a plurality of cache units, wherein the plurality of cache units correspond one-to-one to a plurality of pending request buffers, wherein the pending request buffer corresponding to any cache unit is used to buffer pending data requests allocated to the cache unit, and the device comprises:

[0034] A first merging module is configured to merge the pending data requests in any pending request buffer to obtain at least one first-level merged request, and write the at least one first-level merged request into a first-level merged request buffer corresponding to the pending request buffer;

[0035] a second merging module, configured to, for any first-level merge request taken out from any first-level merge request buffer, merge the first-level merge request with the second-level merge request in response to any second-level merge request in a second-level merge request buffer corresponding to the first-level merge request buffer and the first-level merge request satisfying a preset merging condition, to obtain a new second-level merge request;

[0036] A processing module is used to take out the second-level merge request from the second-level merge request buffer for processing.

[0037] In a possible implementation, the preset merging condition includes:

[0038] The data requested by the first-level merge request and the second-level merge request are located in N adjacent cache lines, where N is an integer greater than or equal to 2.

[0039] In a possible implementation, the preset merging condition includes:

[0040] The first-level merge request and the second-level merge request are data requests received within M adjacent clock cycles, where M is an integer greater than or equal to 2.

[0041] In a possible implementation, the preset merging condition includes at least one of the following:

[0042] The first-level merge request and the second-level merge request come from the same request end;

[0043] The first-level merge request and the second-level merge request correspond to the same system-level cache policy;

[0044] The first-level merge request and the second-level merge request correspond to the same consistency strategy.

[0045] In a possible implementation manner, the device further includes:

[0046] A first writing module is configured to write, for any first-level merge request taken out from any first-level merge request buffer, the first-level merge request as a new second-level merge request into a second-level merge request buffer in response to the absence of a second-level merge request that satisfies the preset merge condition with the first-level merge request in a second-level merge request buffer corresponding to the first-level merge request buffer.

[0047] In a possible implementation, the processing module is used to:

[0048] In response to the second-level merge request most recently written into the second-level merge request buffer carrying instruction end marker information, the second-level merge request is taken out from the second-level merge request buffer for processing, wherein the instruction end marker information is used to indicate the last data request in the request sequence.

[0049] In a possible implementation, the first merging module is used to:

[0050] In response to the number of pending data requests in the pending request buffer being greater than or equal to L, taking out L pending data requests from the pending request buffer, where L is an integer greater than or equal to 2;

[0051] The L pending data requests are merged to obtain K first-level merged requests, and the K first-level merged requests are written into a first-level merged request buffer corresponding to the pending request buffer, where K is an integer greater than 0 and less than or equal to L.

[0052] In a possible implementation, the first merging module is used to:

[0053] Among the L data requests to be processed, data requests whose requested data are located in the same cache line are merged to obtain K first-level merged requests.

[0054] In a possible implementation manner, the device further includes:

[0055] A determination module, used to determine a target cache unit corresponding to a data request to be processed from a plurality of cache units;

[0056] The second writing module is used to write the to-be-processed data request into a target to-be-processed request buffer corresponding to the target cache unit.

[0057] In a possible implementation manner, the determining module is used to:

[0058] Performing a hash operation on the request address of the data request to be processed to obtain a hash operation result corresponding to the data request to be processed;

[0059] According to the hash operation result, a target cache unit corresponding to the to-be-processed data request is determined from a plurality of cache units.

[0060] According to one aspect of the present disclosure, an electronic device is provided, comprising: one or more processors; a memory for storing executable instructions; wherein the one or more processors are configured to call the executable instructions stored in the memory to execute the above method.

[0061] According to one aspect of the present disclosure, a computer-readable storage medium is provided, on which computer program instructions are stored, and the computer program instructions implement the above method when executed by a processor.

[0062] According to one aspect of the present disclosure, a computer program product is provided, including a computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in an electronic device, a processor in the electronic device executes the above method.

[0063] In an embodiment of the present disclosure, a plurality of pending request buffers corresponding one to one to a plurality of cache units are set to buffer pending data requests assigned to the cache units. For any pending request buffer, the pending data requests in the pending request buffer are merged to obtain at least one first-level merge request, and the at least one first-level merge request is written into a first-level merge request buffer corresponding to the pending request buffer. For any first-level merge request taken out from any first-level merge request buffer, in response to any second-level merge request in a second-level merge request buffer corresponding to the first-level merge request buffer satisfying a preset merge condition with the first-level merge request, the first-level merge request is merged with the second-level merge request to obtain a new second-level merge request, and the second-level merge request is taken out from the second-level merge request buffer for processing. Thus, by buffering the pending data requests in the request buffer, the multi-threaded architecture in the related art is improved, which is beneficial for a single cache unit to process more threads, thereby improving the efficiency of cache access as a whole. Moreover, by setting a first-level merge request buffer and a second-level merge request buffer, the parallel multi-data request merging scheme in the related art is changed to a serial multi-data request merging scheme, which is conducive to the scalable processing of more threads by a single cache unit and has a more flexible configuration. In addition, the frequency of cache and memory access can be reduced and the access efficiency of cache and memory can be improved by merging data requests. Furthermore, by adopting a second-level merge request buffer, it is possible to fully merge data requests, reduce the number of data requests sent to the cache unit, and reduce the number of accesses to the cache unit, thereby reducing the burden of memory access.

[0064] In the related art, a single cache unit can only process 8 threads in parallel at most, and cannot achieve high timing and efficiency. The disclosed embodiment can be expanded to 16, 32, 64, 128 threads, which can bring higher timing benefits.

[0065] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure.

[0066] Further features and aspects of the present disclosure will become apparent from the following detailed description of exemplary embodiments with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] The drawings herein are incorporated into the specification and constitute a part of the specification. These drawings illustrate embodiments consistent with the present disclosure and are used to illustrate the technical solutions of the present disclosure together with the specification.

[0068] Figure 1 A flow chart of a data request processing method provided by an embodiment of the present disclosure is shown.

[0069] Figure 2 A schematic diagram showing the hardware architecture of the data request processing method provided by an embodiment of the present disclosure.

[0070] Figure 3 A schematic diagram showing an instruction expansion module.

[0071] Figure 4 A schematic diagram showing an address balancing allocation module and a first-level merge request buffer in a data request processing method provided in an embodiment of the present disclosure.

[0072] Figure 5 Another schematic diagram showing the hardware architecture of the data request processing method provided by an embodiment of the present disclosure.

[0073] Figure 6 A schematic diagram showing a second-level merge request buffer in the data request processing method provided in an embodiment of the present disclosure.

[0074] Figure 7 A schematic diagram showing a hit history information table in a data request processing method provided in an embodiment of the present disclosure.

[0075] Figure 8 A block diagram of a data request processing device provided by an embodiment of the present disclosure is shown.

[0076] Fig. 9 A block diagram of an electronic device 1900 provided by an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0077] Various exemplary embodiments, features and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference numerals in the accompanying drawings represent elements with the same or similar functions. Although various aspects of the embodiments are shown in the accompanying drawings, the drawings are not necessarily drawn to scale unless otherwise specified.

[0078] The word “exemplary” is used exclusively herein to mean “serving as an example, example, or illustration.” Any embodiment described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments.

[0079] The term "and / or" herein is only a description of the association relationship of the associated objects, indicating that there may be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the term "at least one" herein represents any combination of at least two of any one or more of a plurality of. For example, including at least one of A, B, and C can represent including any one or more elements selected from the set consisting of A, B, and C.

[0080] In addition, in order to better illustrate the present disclosure, numerous specific details are given in the following specific embodiments. It should be understood by those skilled in the art that the present disclosure can also be implemented without certain specific details. In some examples, methods, means, components and circuits well known to those skilled in the art are not described in detail in order to highlight the subject matter of the present disclosure.

[0081] The basic operation instructions of GPU (Graphics Processing Unit) memory multithreading include read operation (load) instructions, write operation (store) instructions, atomic operation (atomic) instructions, etc. In order to improve the efficiency of GPU instruction operation and reduce the burden of downstream memory access, instructions are usually merged and streamlined before cache operations.

[0082] In the related art, multiple data requests are usually merged by horizontal merging (i.e. parallel merging). This solution is simple to implement, but for hardware, the hardware scalability of this solution is poor, it cannot support more threads, the timing resource consumption is large, and the hardware does not support higher frequencies. Usually, the instruction merging in the related art supports the merging of data requests of 4 or 8 threads at most, but does not support the merging of data requests of more threads. When processing multiple data requests, it cannot achieve higher timing and frequency.

[0083] In order to solve technical problems similar to those described above, the present disclosure provides a data request processing method, by setting multiple pending request buffers corresponding to multiple cache units one by one, for buffering pending data requests allocated to the cache units, for any pending request buffer, merging the pending data requests in the pending request buffer to obtain at least one first-level merge request, and writing the at least one first-level merge request into the first-level merge request buffer corresponding to the pending request buffer, for any first-level merge request taken out from any first-level merge request buffer, in response to any second-level merge request in the second-level merge request buffer corresponding to the first-level merge request buffer satisfying a preset merge condition with the first-level merge request, merging the first-level merge request with the second-level merge request to obtain a new second-level merge request, and taking the second-level merge request from the second-level merge request buffer for processing, thereby buffering the pending data requests by the request buffer, improving the multi-threaded architecture in the related art, which is beneficial for a single cache unit to process more threads, thereby improving the efficiency of cache access as a whole. Moreover, by setting a first-level merge request buffer and a second-level merge request buffer, the parallel multi-data request merging scheme in the related art is changed to a serial multi-data request merging scheme, which is conducive to the scalable processing of more threads by a single cache unit and has a more flexible configuration. In addition, the frequency of cache and memory access can be reduced and the access efficiency of cache and memory can be improved by merging data requests. Furthermore, by adopting a second-level merge request buffer, it is possible to fully merge data requests, reduce the number of data requests sent to the cache unit, and reduce the number of accesses to the cache unit, thereby reducing the burden of memory access.

[0084] In the related art, a single cache unit can only process 8 threads in parallel at most, and cannot achieve high timing and efficiency. The disclosed embodiment can be expanded to 16, 32, 64, 128 threads, which can bring higher timing benefits.

[0085] The data request processing method provided by the embodiment of the present disclosure is described in detail below with reference to the accompanying drawings.

[0086] Figure 1A flow chart of a data request processing method provided in an embodiment of the present disclosure is shown. In one possible implementation, the execution subject of the data request processing method may be a data request processing device. For example, the data request processing method may be executed by a terminal device or a server or other electronic device. The terminal device may be a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device or a wearable device, etc. In some possible implementations, the data request processing method may be implemented by a processor calling computer-readable instructions stored in a memory. In an embodiment of the present disclosure, the cache includes a plurality of cache units, and the plurality of cache units correspond one-to-one to a plurality of pending request buffers, wherein the pending request buffer corresponding to any cache unit is used to buffer pending data requests assigned to the cache unit. As Figure 1 As shown, the data request processing method includes steps S11 to S13.

[0087] In step S11, for any pending request buffer, the pending data requests in the pending request buffer are merged to obtain at least one first-level merged request, and the at least one first-level merged request is written into the first-level merged request buffer corresponding to the pending request buffer.

[0088] In step S12, for any first-level merge request taken out from any first-level merge request buffer, in response to any second-level merge request in the second-level merge request buffer corresponding to the first-level merge request buffer and the first-level merge request satisfying a preset merge condition, the first-level merge request and the second-level merge request are merged to obtain a new second-level merge request.

[0089] In step S13, the second-level merge request is taken out from the second-level merge request buffer for processing.

[0090] In the embodiment of the present disclosure, the cache includes a plurality of cache units, and the cache unit may represent a unit obtained by dividing the cache. For example, the cache unit may refer to a cache bank. Of course, the cache unit may also be defined in other ways, which are not limited here.

[0091] In a possible implementation, the method further includes: determining a target cache unit corresponding to the data request to be processed from a plurality of cache units; and writing the data request to be processed into a target pending request buffer corresponding to the target cache unit.

[0092] In this implementation, in response to a data request to be processed, one of the multiple cache units may be determined as a target cache unit corresponding to the data request to be processed, wherein the target cache unit may represent a cache unit corresponding to the data request to be processed.

[0093] In one possible implementation, before determining the target cache unit corresponding to the data request to be processed from multiple cache units, the method also includes: in response to an original data request from any thread or any program block, splitting the original data request into at least one data request to be processed according to a preset splitting granularity.

[0094] The thread may be a GPU thread or a DMA (Direct Memory Access) thread, etc., which is not limited here.

[0095] In this implementation, the preset splitting granularity may be a cache line size, 1 / 2 of the cache line size, or an even multiple of the cache line size, etc., which is not limited here.

[0096] As an example of this implementation, in response to an original data request from any thread or any program block, the instruction expansion module may split the original data request into at least one data request to be processed according to a preset splitting granularity.

[0097] In this implementation, in response to an original data request from any thread or any program block, the starting address and burst length of the data requested by the original data request can be obtained, and the original data request can be split into at least one data request to be processed according to the starting address, the burst length, the address boundary of the cache line, and the preset split granularity. By splitting the original data request according to the address boundary of the cache line, the address of the data requested by the split data request to be processed can be aligned to the cache line.

[0098] As an example of this implementation, the preset splitting granularity is the cache line size. In this example, each of the intermediate data requests to be processed obtained by splitting the original data request can correspond to a cache line one-to-one, wherein the intermediate data request to be processed can represent the data requests to be processed other than the first data request to be processed and the last data request to be processed among the multiple data requests to be processed obtained by splitting the original data request. Among them, each of the intermediate data requests to be processed corresponds to a cache line one-to-one, which can mean that each of the intermediate data requests to be processed can correspond to each complete cache line one-to-one. The first data request to be processed obtained by splitting the original data request may correspond to a complete cache line, or to a portion of a cache line, or to a portion of a cache line and a complete cache line. The last data request to be processed obtained by splitting the original data request may correspond to a complete cache line, or to a portion of a cache line, or to a complete cache line and a portion of a cache line.

[0099] In this implementation, by responding to an original data request from any thread or any program block, the original data request is split into at least one data request to be processed according to a preset splitting granularity, thereby improving the efficiency of subsequent data request processing.

[0100] As an example of this implementation, the method further includes: determining the preset splitting granularity according to a bit width of a specified data interface.

[0101] In this example, the designated data interface may be a data interface for returning information for the original data request. That is, the preset splitting granularity may be determined based on the data interface for returning information for the original data request. Of course, depending on the actual application scenario, the designated data interface may also be other data interfaces, which are not limited here.

[0102] In this example, the preset split granularity is determined according to the bit width of the specified data interface, thereby being able to determine a suitable split granularity.

[0103] As another example of this implementation, the preset splitting granularity may be a default value.

[0104] As an example of this implementation, after splitting the original data request into at least one data request to be processed, the method further includes: for any data request to be processed, in response to the data requested by the data request to be processed being located in more than two cache lines, marking the data request to be processed.

[0105] In this example, for any pending data request, an address boundary check can be performed on the data requested by the pending data request to determine whether the data requested by the pending data request spans cache lines (i.e., to determine whether the data requested by the pending data request is located in at least two cache lines).

[0106] In one example, for any pending data request, the instruction expansion module may perform an address boundary check on the data requested by the pending data request to determine whether the data requested by the pending data request spans cache lines.

[0107] In this example, the first data request to be processed and the last data request to be processed of the at least one data request to be processed may cross cache lines. In one example, when the first data request to be processed and / or the last data request to be processed cross cache lines, the instruction expansion module may output a signal to indicate so as to mark the first data request to be processed and / or the last data request to be processed.

[0108] In this example, by marking any one of the at least one pending data request, in response to the data requested by the pending data request being located in more than two cache lines, the pending data request is enabled to access the complete data in the cache.

[0109] In another possible implementation, the original data requests from each thread and each program block may be respectively determined as data requests to be processed. In this implementation, the original data requests may not be split.

[0110] As an example of this implementation, when the original data request is determined as a pending data request, the data request in the pending request buffer can be directly sent to the cache unit for cache reading and writing processing without merging processing.

[0111] In one possible implementation, determining a target cache unit corresponding to a data request to be processed from a plurality of cache units includes: performing a hash operation on a request address of the data request to be processed to obtain a hash operation result corresponding to the data request to be processed; and determining, based on the hash operation result, a target cache unit corresponding to the data request to be processed from a plurality of cache units.

[0112] As an example of this implementation method, the address balancing allocation module can perform a hash operation on the request address of the data request to be processed to obtain a hash operation result corresponding to the data request to be processed, and determine the target cache unit corresponding to the data request to be processed from multiple cache units based on the hash operation result.

[0113] Since the cache unit can only receive a single data request in sequence each time to perform subsequent cache read and write operations, a hash operation is performed on the request address of the data request to be processed to obtain a hash operation result corresponding to the data request to be processed, and based on the hash operation result, the target cache unit corresponding to the data request to be processed is determined from multiple cache units, thereby alleviating the interaction pressure between different data requests and cache units.

[0114] In another possible implementation, determining the target cache unit corresponding to the data request to be processed from the plurality of cache units includes: randomly determining the target cache unit corresponding to the data request to be processed from the plurality of cache units.

[0115] In the embodiment of the present disclosure, the cache unit corresponds to the pending request buffer one-to-one, that is, the multiple cache units correspond to the multiple pending request buffers one-to-one. The target pending request buffer may represent the pending request buffer corresponding to the target cache unit. In the embodiment of the present disclosure, after determining the target cache unit corresponding to the pending data request, the pending data request may be written into the target pending request buffer corresponding to the target cache unit.

[0116] Since different data requests may be assigned to the same cache unit, by setting a pending request buffer, data requests from multiple threads can be buffered. The depth of the pending request buffer can be set according to the actual application scenario requirements. For example, the depth of the pending request buffer can be 16 or 32, etc., which is not limited here. The depth of the pending request buffer is 16, which means that the pending request buffer can buffer 16 pending data requests; the depth of the pending request buffer is 32, which means that the pending request buffer can buffer 32 pending data requests; and so on.

[0117] In a possible implementation, the pending request buffer may be set in the address balanced allocation module.

[0118] In the disclosed embodiment, the type of data request to be merged may be a read (load) request, an atomic operation (atomic) request, a write (store) request, etc., which is not limited here. By merging data requests, the number of requests sent to the cache unit can be reduced.

[0119] In the embodiment of the present disclosure, for any pending request buffer, the pending data requests in the pending request buffer are merged to obtain at least one first-level merged request, and the at least one first-level merged request is written into the first-level merged request buffer corresponding to the pending request buffer. For example, the data requests in the target pending request buffer are merged to obtain a first-level merged request, and the first-level merged request is written into the target first-level merged request buffer corresponding to the target pending request buffer.

[0120] In the embodiment of the present disclosure, the first-level merge request buffer corresponding to any pending request buffer can be used to buffer the first-level merge request obtained by merging the pending data requests in the pending request buffer. The depth of the first-level merge request buffer can be flexibly set according to the actual application scenario requirements and is not limited here. For example, the depth of the first-level merge request buffer can be 3, 4, 8, 10, etc.

[0121] In the embodiment of the present disclosure, the first-level merge request may represent a data request obtained by merging data requests in the request buffer to be processed.

[0122] In a possible implementation, the data requests may be merged by an instruction merging unit, wherein the instruction merging unit may include a first-level instruction merging subunit and a second-level instruction merging subunit.

[0123] The first-level instruction merging subunit may merge the data requests in the pending request buffer to obtain a first-level merged request, and write the first-level merged request into the first-level merged request buffer corresponding to the pending request buffer. For example, the first-level instruction merging subunit may merge the data requests in the target pending request buffer to obtain a first-level merged request, and write the first-level merged request into the target first-level merged request buffer corresponding to the target pending request buffer.

[0124] In one example, the number of pending request buffers, first-level instruction merging subunits, and second-level instruction merging subunits may be the same. For example, four pending request buffers, four first-level instruction merging subunits, and four second-level instruction merging subunits may be provided.

[0125] In one possible implementation, the merging of the pending data requests in the pending request buffer to obtain at least one first-level merged request, and writing the at least one first-level merged request into a first-level merged request buffer corresponding to the pending request buffer, includes: in response to the number of pending data requests in the pending request buffer being greater than or equal to L, taking out L pending data requests from the pending request buffer, where L is an integer greater than or equal to 2; merging the L pending data requests to obtain K first-level merged requests, and writing the K first-level merged requests into a first-level merged request buffer corresponding to the pending request buffer, where K is an integer greater than 0 and less than or equal to L.

[0126] For example, L can be 4, D can be 16, and D represents the depth of the pending request buffer. If there are 16 pending data requests in a pending request buffer, namely pending data requests 0-15, then pending data requests 0-3 can be merged to obtain at least one first-level merged request, pending data requests 4-7 can be merged to obtain at least one first-level merged request, pending data requests 8-11 can be merged to obtain at least one first-level merged request, and pending data requests 12-15 can be merged to obtain at least one first-level merged request. Moreover, the first-level merged request including the valid request address in each first-level merged request can be written into the first-level merged request buffer. That is, if any first-level merged request obtained by the merge does not include a valid request address, the first-level merged request may not be written into the first-level merged request buffer.

[0127] For another example, L may be 8, and D may be 16, where D represents the depth of the pending request buffer. If there are 16 pending data requests in a pending request buffer, namely pending data requests 0-15, then pending data requests 0-7 may be merged to obtain at least one first-level merged request, and pending data requests 8-15 may be merged to obtain at least one first-level merged request. Furthermore, the first-level merged request including the valid request address in each first-level merged request may be written into the first-level merged request buffer.

[0128] In this implementation, in response to the number of data requests in the pending request buffer being greater than or equal to L, L pending data requests are taken out from the pending request buffer, where L is an integer greater than or equal to 2, and the L pending data requests are merged to obtain K first-level merged requests, and the K first-level merged requests are written into the first-level merged request buffer corresponding to the pending request buffer, thereby improving the efficiency of data request merging.

[0129] In a possible implementation, taking out the L pending data requests from the pending request buffer includes: taking out the first L pending data requests written into the pending request buffer.

[0130] In this implementation, in response to the number of data requests in the pending request buffer being greater than or equal to L, the first L pending data requests written are taken out from the pending request buffer, where L is an integer greater than or equal to 2, and the L pending data requests are merged to obtain K first-level merged requests, and the K first-level merged requests are written into the first-level merged request buffer corresponding to the pending request buffer, thereby enabling orderly processing of data requests and improving system stability.

[0131] In one possible implementation, the merging of the L pending data requests to obtain K first-level merged requests includes: merging data requests in which the requested data are located in the same cache line among the L pending data requests to obtain K first-level merged requests.

[0132] In this implementation, the request addresses of L pending data requests can be compared pairwise according to the address of the cache line size, and the data requests whose requested data are located in the same cache line can be merged to obtain K first-level merged requests, and the K first-level merged requests can be written into the first-level merged request buffer in sequence.

[0133] In this implementation, by merging the data requests whose requested data are located in the same cache line among the L data requests to be processed, K first-level merged requests are obtained, thereby improving the efficiency of subsequent cache reading and writing.

[0134] In a possible implementation, a mapping relationship between identification information of a data request to be processed and identification information of a first-level merged request may be recorded in a request mapping table. As an example of this implementation, the request mapping table may be a buffer built based on SRAM (Static Random Access Memory).

[0135] In the disclosed embodiment, for any first-level merge request buffer, one first-level merge request may be taken out from the first-level merge request buffer for subsequent processing each time. For example, the first-level merge request written first may be taken out from the first-level merge request buffer for processing each time. By taking out the first-level merge request written first from the first-level merge request buffer for processing, orderly processing of the first-level merge requests may be achieved, thereby improving the stability of the system.

[0136] In a possible implementation, one first-level merge request may be taken out from each first-level merge request buffer in each clock cycle for subsequent processing.

[0137] In an embodiment of the present disclosure, when a first-level merge request is taken out from a first-level merge request buffer, the corresponding second-level merge request buffer may be checked to find out whether there is a second-level merge request that can be merged with the current first-level merge request. Among them, it may be determined whether two requests can be merged according to a preset merge condition. For any first-level merge request taken out from any first-level merge request buffer, in response to any second-level merge request in a second-level merge request buffer corresponding to the first-level merge request buffer satisfying a preset merge condition with the first-level merge request, the first-level merge request may be merged with the second-level merge request to obtain a new second-level merge request.

[0138] In the disclosed embodiment, the second-level merge request buffer may be used to buffer the second-level merge request, wherein the second-level merge request may be a data request obtained by merging two or more first-level merge requests, or may be a single first-level merge request.

[0139] In a possible implementation, the preset merge condition includes: the data requested by the first-level merge request and the second-level merge request are located in N adjacent cache lines, where N is an integer greater than or equal to 2.

[0140] In this implementation, the first-level instruction merging subunit may only support merging of data requests whose requested data are located in the same cache line, while the second-level instruction merging subunit may support merging of data requests whose requested data are located in N adjacent cache lines. That is, for any first-level merge request, the data requested by each data request obtained by merging the first-level merge request is located in the same cache line; for any second-level merge request, the data requested by each data request obtained by merging the second-level merge request may be located in different cache lines.

[0141] For example, the size of the cache line is 32 bytes, the first-level instruction merging subunit may only support the merging of 32-byte data requests, while the second-level instruction merging subunit may support the merging of 64-byte or 128-byte data requests.

[0142] In this implementation, by setting a preset merge condition, the first-level merge request and the second-level merge request are allowed to be merged when the requested data is located in N adjacent cache lines, thereby reducing the number of accesses to the memory or cache, thereby reducing memory access latency.

[0143] In a possible implementation, the preset merging condition includes: the first-level merge request and the second-level merge request are data requests received within M adjacent clock cycles, where M is an integer greater than or equal to 2.

[0144] In this implementation, compared with the first-level instruction merging subunit, the second-level instruction merging subunit can merge data requests within more time cycles. For example, the first-level instruction merging subunit only supports merging data requests within a single time cycle, while the second-level instruction merging subunit can support merging data requests within M adjacent clock cycles.

[0145] In this implementation, the preset merging condition is a time-based merging strategy, that is, the first-level merge request and the second-level merge request are data requests received within M adjacent clock cycles, where M is an integer greater than or equal to 2. Within M clock cycles, multiple requests are merged into a larger request, reducing the number of requests to the bus and cache controller, thereby reducing their burden.

[0146] In a possible implementation, the preset merge condition includes at least one of the following: the first-level merge request and the second-level merge request come from the same request end; the first-level merge request and the second-level merge request correspond to the same system-level cache policy (SLC Cache Policy, System Level Cache CachePolicy); the first-level merge request and the second-level merge request correspond to the same consistency policy.

[0147] As an example of this implementation, the preset merge condition may include: the first-level merge request and the second-level merge request come from the same request end (master).

[0148] In this example, for any first-level merge request and any second-level merge request, if the first-level merge request and the second-level merge request come from different request ends, it can be determined that the first-level merge request and the second-level merge request do not meet the preset merge condition, and thus the first-level merge request and the second-level merge request will not be merged.

[0149] As an example of this implementation, the preset merge condition may include: the first-level merge request and the second-level merge request correspond to the same system-level cache policy, wherein the system-level cache policy may include policies such as how data is cached, when to perform cache replacement, and cache consistency maintenance.

[0150] In this example, for any first-level merge request and any second-level merge request, if the first-level merge request and the second-level merge request correspond to different system-level cache strategies, it can be determined that the first-level merge request and the second-level merge request do not meet the preset merge condition, and thus the first-level merge request and the second-level merge request will not be merged.

[0151] As an example of this implementation, the preset merge condition may include: the first-level merge request and the second-level merge request correspond to the same consistency strategy.

[0152] In this example, for any first-level merge request and any second-level merge request, if the first-level merge request and the second-level merge request correspond to different consistency strategies, it can be determined that the first-level merge request and the second-level merge request do not meet the preset merge condition, and thus the first-level merge request and the second-level merge request will not be merged.

[0153] In one possible implementation, the method also includes: for any first-level merge request taken out from any first-level merge request buffer, in response to the fact that there is no second-level merge request that satisfies the preset merge condition with the first-level merge request in the second-level merge request buffer corresponding to the first-level merge request buffer, writing the first-level merge request as a new second-level merge request into the second-level merge request buffer.

[0154] In this implementation, when a first-level merge request is taken out from the first-level merge request buffer, if a second-level merge request that meets the preset merge conditions is not found in the corresponding second-level merge request buffer, then the current first-level merge request will be directly written into the second-level merge request buffer as a new second-level merge request.

[0155] This implementation can dynamically decide whether to merge the current first-level merge request according to the existing second-level merge request in the second-level merge request buffer. By adopting this implementation, cache resources can be used more effectively, unnecessary memory access can be reduced, and the processing efficiency of the entire system can be improved.

[0156] In a possible implementation, second-level merge request buffers may be respectively set for different request ends and different types of data requests. In this implementation, for any second-level merge request buffer, each second-level merge request in the second-level merge request buffer comes from the same request end and belongs to the same type of data request.

[0157] In one example, the request end includes a MUSA processor engine (MUSAProcessor Engine, MPE) and a programmable data sequencer (Programmable Data Sequencer, PDS). A second-level merge request buffer B11 may be set for a read request from the MUSA processor engine; a second-level merge request buffer B12 may be set for an atomic operation request from the MUSA processor engine; a second-level merge request buffer B13 may be set for a write request from the MUSA processor engine; a second-level merge request buffer B21 may be set for a read request from the programmable data sequencer; a second-level merge request buffer B22 may be set for a write request from the programmable data sequencer; and so on.

[0158] In one example, the depth of the second-level merge request buffer B11 may be 16, the depth of the second-level merge request buffer B21 may be 4, and the depth of the second-level merge request buffer B12 may be 1. Of course, the depth of each second-level merge request buffer can be flexibly set according to the actual application scenario requirements, and is not limited here.

[0159] Among them, the depth of any second-level merge request buffer can represent the number of storage units (entries) in the second-level merge request buffer, that is, the depth of any second-level merge request buffer can represent the maximum number of second-level merge requests that can be stored in the second-level merge request buffer.

[0160] In an example, the second-level merge request buffer B12 can store requests for atomic operation from the MUSA processor engine that require data to be returned.

[0161] In a possible implementation, each storage unit in the second-level merge request buffer may include a merge flag bit. The merge flag bit of any storage unit may be used to record whether the data requested by the second-level merge request in the storage unit has been merged.

[0162] In a possible implementation, any storage unit in the second-level merge request buffer may be used to store tag information (tag) of the second-level merge request.

[0163] As an example of this implementation method, for the second-level merge request buffer corresponding to the read request, the tag information stored in any storage unit in the second-level merge request buffer may include at least part of the context identification information (ctxt_pasid), the address of the requested data (vaddr), the burst length (burst), the identification information of the second-level merge request (req_id), etc.

[0164] As an example of this implementation method, for the second-level merge request buffer corresponding to the atomic operation request, the tag information stored in any storage unit in the second-level merge request buffer may include at least part of the context identification information (ctxt_pasid), the address of the requested data (vaddr), the target data (data), the mask (mask), the identification information of the second-level merge request (req_id), etc.

[0165] In a possible implementation, each storage unit in the second-level merge request buffer may include a most recent merge time flag (age). The most recent merge time flag of any storage unit may be used to record the time when the second-level merge request in the storage unit was most recently merged. For example, the most recent merge time flag of any storage unit may be used to record the time period that has passed since the second-level merge request in the storage unit was last updated.

[0166] In an example, each storage unit in the second-level merge request buffer B11 may include a most recent merge time flag, and the second-level merge request buffer B21 and the second-level merge request buffer B12 may not include a most recent merge time flag.

[0167] In an embodiment of the present disclosure, a second-level merge request can be taken out from a second-level merge request buffer and sent to a cache unit (e.g., a cache block) for cache read and write processing. For any second-level merge request in the second-level merge request buffer, the second-level merge request can be cleared from the second-level merge request buffer in response to the completion of the execution of the second-level merge request. The completion of the execution of the second-level merge request may indicate that all related memory operation requests of the second-level merge request have been processed, and the corresponding data has been written back or read. The relevant information of the second-level merge request is cleared from the second-level merge request buffer to release resources.

[0168] In one possible implementation, taking out the second-level merge request from the second-level merge request buffer for processing includes: in response to the second-level merge request most recently written in the second-level merge request buffer carrying instruction end marker information, taking out the second-level merge request from the second-level merge request buffer for processing, wherein the instruction end marker information is used to indicate the last data request in the request sequence.

[0169] In this implementation, if any data request carries instruction end marker information, it can indicate that the data request is the last data request in the request sequence. For example, the instruction end marker information can be "end_ofinstruction".

[0170] In one example, the requesting end includes a MUSA processor engine, and the instruction end marker information corresponding to the MUSA processor engine can include a Cache Flash Invalid (CFI) barrier instruction.

[0171] In this implementation, for any second-level merge request buffer, if the latest written second-level merge request in the second-level merge request buffer carries instruction end marker information, all the written second-level merge requests can be kicked out from the second-level merge request buffer for cache read / write processing.

[0172] In a possible implementation, for any second-level merge request buffer, in response to the second-level merge request buffer being full, second-level merge requests can be taken out from the second-level merge request buffer for cache read / write processing according to a preset replacement policy.

[0173] For example, the preset replacement policy can be: replacing the second-level merge request in the second-level merge request buffer that has not been merged for the longest time.

[0174] Another example, the preset replacement policy can be: replacing the second-level merge request in the second-level merge request buffer whose write duration is greater than or equal to a preset duration and the number of merges is less than or equal to a preset number of times.

[0175] Of course, the replacement policy can be flexibly set according to the actual application scenario requirements and is not limited here.

[0176] In a possible implementation, the first-level instruction merge subunit and the second-level instruction merge subunit may not merge the instructions for maintaining cache coherence. For the instructions for maintaining cache coherence, the first-level instruction merge subunit and the second-level instruction merge subunit can receive and forward (issue) the instruction to subsequent processing stages, such as cache or memory. In addition, the first-level instruction merge subunit and the second-level instruction merge subunit can also support the ordering of different data requests and instructions. Inside a single instruction, different threads can execute in an out-of-order manner. Out-of-order execution allows multiple threads to operate simultaneously without having to wait for them to complete in order, thus improving performance.

[0177] The data request processing method provided in the embodiments of the present disclosure can be applied to technical fields such as GPU, multi-instruction, multi-threading, AI (Artificial Intelligence), cache coherence, etc., which are not limited here. In addition, the data request processing method provided in the embodiments of the present disclosure can be applied to application scenarios such as GPU / DMA multi-threaded parallel reading and writing to improve the efficiency of GPU / DMA multi-threaded parallel reading and writing, which are not limited here.

[0178] The data request processing method provided by the embodiment of the present disclosure is described below through a specific application scenario. Figure 2 A schematic diagram showing the hardware architecture of the data request processing method provided by the embodiment of the present disclosure. Figure 2 It can be seen that compared with the related art, the hardware architecture provided by the embodiment of the present disclosure allows a single cache unit (such as a cache block) to process more threads. For example, the number of threads that can be processed in parallel by a single cache unit may be 8, 16, 32, 64, 128, etc. Figure 2 It shows that a single cache unit can process 8 threads in parallel (see Figure 2 Threads 0 to 7, Threads J to J+7) and 16 threads (see Figure 2 An example of threads 0 to 15, thread J to thread J+15).

[0179] Figure 3 FIG. 4 is a schematic diagram showing the instruction expansion module. Figure 3 As shown, the instruction expansion module may include an address boundary check submodule, an address cross-cache line check submodule, an address calculation submodule and a data alignment submodule. Among them, the address boundary check submodule can determine whether the request address of the original data request is an address boundary OOB (Out Of Boundary, beyond the boundary) according to the address boundary, starting address, output interface bit width and burst length of the cache line. That is, the address boundary detection submodule can be used to determine whether the request address of the original data request exceeds the address boundary of the cache line. The address cross-cache line check module can be used to determine whether the request address of the data request to be processed crosses the cache line, and the data request to be processed can be marked by requesting a cross-cache line when the request address of the data request to be processed crosses the cache line. The address calculation module can be used to expand the request address of the original data request, for example, splitting the original data request into multiple data requests to be processed. The data alignment module can be used to align the written data to the cache line for the write instruction.

[0180] like Figure 2As shown, for any one of at least one pending data request, a hash operation can be performed on the request address of the pending data request through the address balancing allocation module to obtain a hash operation result corresponding to the pending data request, and based on the hash operation result, a target cache unit corresponding to the pending data request is determined from multiple cache units (such as cache blocks).

[0181] exist Figure 2 In the embodiment, the instruction merging unit may include a first-level instruction merging subunit and a second-level instruction merging subunit. The first-level instruction merging subunit may be used to merge the pending data requests in any pending request buffer to obtain at least one first-level merge request, and write the at least one first-level merge request into the first-level merge request buffer corresponding to the pending request buffer. The second-level instruction merging subunit may be used to merge the first-level merge request with the second-level merge request for any first-level merge request taken out from any first-level merge request buffer, in response to any second-level merge request in the second-level merge request buffer corresponding to the first-level merge request buffer meeting a preset merge condition with the first-level merge request, to obtain a new second-level merge request. The second-level merge request may be taken out from the second-level merge request buffer and sent to a cache unit (e.g., a cache block) for cache read and write processing.

[0182] Figure 4 A schematic diagram showing an address balancing allocation module and a first-level merge request buffer in a data request processing method provided by an embodiment of the present disclosure. Figure 4 In the example shown, 16 data requests are processed in parallel. Figure 4 As shown, the address balancing allocation module may include a hash operation submodule and a pending request buffer. The hash operation submodule may be used to perform a hash operation on the request address of the pending data request to obtain a hash operation result corresponding to the pending data request, and according to the hash operation result, determine the target cache unit corresponding to the pending data request from multiple cache units (cache blocks), thereby determining the target pending request buffer and the target first-level merge request buffer corresponding to the pending data request. The pending request buffer may be used to buffer data requests. The pending data request may be taken out from the pending request buffer for merging to obtain a first-level merge request, and the first-level merge request may be written into the first-level merge request buffer corresponding to the pending request buffer. The first-level merge request written first may be taken out from the first-level merge request buffer for subsequent processing.

[0183] Figure 5Another schematic diagram showing the hardware architecture of the data request processing method provided by the embodiment of the present disclosure. Figure 5 In the example shown, the request end includes MUSA processor engine 0 (MPE0), MUSA processor engine 1 (MPE1), MUSA processor engine 2 (MPE2), MUSA processor engine 3 (MPE3) and a programmable data sequencer (PDS). Among them, the data request issued by the programmable data sequencer can be distributed through the PDS request distribution module. The merging unit engine (CU engine, Coalesce Unit engine) can include an instruction expansion module, an address balancing allocation module and an instruction merging unit. Figure 5 In the example shown, there are merging unit engines 0, merging unit engines 1, merging unit engines 2, and merging unit engines 3. The instruction merging unit may include a first-level instruction merging subunit and a second-level instruction merging subunit. The merging unit engine may communicate with the cache block via a crossbar switch (xbar). Figure 5 In the example shown, the cache blocks may include cache block 0, cache block 1, cache block 2, and cache block 3 of the L1 cache. In addition, the write request output by the merge unit engine may enter the write operation processing flow. A demultiplexer (DEMUX) may be used to send write requests to multiple targets, such as different cache lines or cache units. When the result of the write operation needs to be sent to a common store (CS), the demultiplexer may be used to distribute these write requests to the correct target address or cache line so that the data is written to the correct location.

[0184] Figure 6 A schematic diagram showing a second-level merge request buffer in the data request processing method provided in an embodiment of the present disclosure. Figure 6 The second-level merge request buffer B11 corresponding to the read request of the MUSA processor engine, the second-level merge request buffer B21 corresponding to the read request of the programmable data sequencer, and the second-level merge request buffer B12 corresponding to the atomic operation request of the MUSA processor engine are shown. Figure 6In the example shown, the depth of the second-level merge request buffer B11 is 16, the depth of the second-level merge request buffer B21 is 4, and the depth of the second-level merge request buffer B12 is 1. Each storage unit in the second-level merge request buffer B11, the second-level merge request buffer B21, and the second-level merge request buffer B12 includes a merge flag bit. In addition, each storage unit in the second-level merge request buffer B11, the second-level merge request buffer B21, and the second-level merge request buffer B12 can be used to store tag information (tag) of the second-level merge request. Among them, the tag information stored in each storage unit in the second-level merge request buffer B11 and the second-level merge request buffer B21 may include context identification information (ctxt_pasid), the address of the requested data (vaddr), the burst length (burst), and the identification information (req_id) of the second-level merge request. The tag information stored in each storage unit in the second-level merge request buffer B12 may include context identification information (ctxt_pasid), the address of the requested data (vaddr), the target data (data), the mask (mask), and the identification information of the second-level merge request (req_id). In addition, the second-level merge request buffer B11 may also include the most recent merge time flag (age). Figure 6 As shown, the second-level merge requests output by multiple second-level merge request buffers can be processed by a multiplexer (MUX).

[0185] Figure 7 A schematic diagram of a hit history information table in a data request processing method provided by an embodiment of the present disclosure is shown. The hit history information table can be used to record the mapping relationship between the first-level merge request and the second-level merge request. Figure 7 In the example shown, the depth of the hit history information table is 64. Of course, the depth of the hit history information table can be flexibly set according to the actual application scenario requirements, and is not limited here.

[0186] exist Figure 7In the example shown, the first column (flag) of the hit history information table is the instruction flag bit, which can be used to distinguish whether the second-level merge request is a load instruction or a non-load instruction. The instruction flag bit can help manage the execution of the second-level merge request and the loading of data more effectively. The second column (tag = req_id[7:0]) can be used to record the lowest 8 bits of the identification information of the second-level merge request. Here, 7:0 is a bit width representation, indicating from bit 7 to bit 0. Of course, it may not be 8 bits according to the configuration of the actual application scenario. The third column (hit req_id) can be used to record the identification information of the first-level merge request corresponding to the second-level merge request. The fourth column (return) can be used to record whether the data has been returned.

[0187] For any first-level merge request taken out from any first-level merge request buffer, if there is no second-level merge request in the second-level merge request buffer corresponding to the first-level merge request that meets the preset merge condition, the first-level merge request can be written into the second-level merge request buffer as a new second-level merge request, and the lowest 8 bits of the identification information of the new second-level merge request can be recorded in the second column of the hit history information table.

[0188] For any first-level merge request taken out from any first-level merge request buffer, if any second-level merge request in the second-level merge request buffer corresponding to the first-level merge request meets the preset merge condition, the first-level merge request and the second-level merge request can be merged to obtain a new second-level merge request, and the identification information of the first-level merge request can be recorded in the third column of the hit history information table.

[0189] In Figure 7 the second-level merge request can be sent to the L1 cache first. If the second-level merge request misses in the L1 cache, it can be sent to a larger L2 cache or other higher-level caches. RSTL (Register Slice Timingbuffer) is a component used to synchronize different clock domains or handle timing issues in processor design. The mix logic unit can be used to process the identification information of the second-level merge request and the returned data. The merge unit (CU) can perform data return processing based on the returned data, the mapping relationship between the second-level merge request and the first-level merge request, and the mapping relationship between the first-level merge request and the original data request (i.e., the "data request to be processed" in the above text).

[0190] This application scenario improves the multi-threaded architecture in the related technology, and changes the parallel multi-threaded streamlining and merging scheme in the related technology to a serial multi-threaded streamlining and merging scheme, which is conducive to a single cache unit processing more threads and has a more flexible configuration.

[0191] In addition, this application scenario can improve the efficiency of a single cache unit. In related technologies, a single cache unit can only process 4 or 8 threads in parallel at most, and cannot achieve high timing and efficiency. This application scenario can be expanded to 16, 32, 64, and 128 threads, and serially merged, which can bring higher timing benefits.

[0192] It can be understood that the above-mentioned various method embodiments mentioned in the present disclosure can be combined with each other to form a combined embodiment without violating the principle logic. Due to space limitations, the present disclosure will not repeat them. It can be understood by those skilled in the art that in the above-mentioned method of the specific implementation method, the specific execution order of each step should be determined according to its function and possible internal logic.

[0193] In addition, the present disclosure also provides a data request processing device, an electronic device, a computer-readable storage medium, and a computer program product, all of which can be used to implement any data request processing method provided by the present disclosure. The corresponding technical solutions and technical effects can be found in the corresponding records in the method part and will not be repeated here.

[0194] Figure 8 The block diagram of the data request processing device provided by the embodiment of the present disclosure is shown. The cache includes a plurality of cache units, and the plurality of cache units correspond to a plurality of pending request buffers one by one, wherein the pending request buffer corresponding to any cache unit is used to buffer the pending data request assigned to the cache unit. Figure 8 As shown, the data request processing device includes:

[0195] A first merging module 81 is configured to merge the pending data requests in any pending request buffer to obtain at least one first-level merged request, and write the at least one first-level merged request into a first-level merged request buffer corresponding to the pending request buffer;

[0196] A second merging module 82 is configured to, for any first-level merge request taken out from any first-level merge request buffer, merge the first-level merge request with the second-level merge request in a second-level merge request buffer corresponding to the first-level merge request buffer to obtain a new second-level merge request in response to any second-level merge request in a second-level merge request buffer corresponding to the first-level merge request buffer and the first-level merge request satisfying a preset merge condition;

[0197] The processing module 83 is used to take out the second-level merge request from the second-level merge request buffer for processing.

[0198] In a possible implementation, the preset merging condition includes:

[0199] The data requested by the first-level merge request and the second-level merge request are located in N adjacent cache lines, where N is an integer greater than or equal to 2.

[0200] In a possible implementation, the preset merging condition includes:

[0201] The first-level merge request and the second-level merge request are data requests received within M adjacent clock cycles, where M is an integer greater than or equal to 2.

[0202] In a possible implementation, the preset merging condition includes at least one of the following:

[0203] The first-level merge request and the second-level merge request come from the same request end;

[0204] The first-level merge request and the second-level merge request correspond to the same system-level cache policy;

[0205] The first-level merge request and the second-level merge request correspond to the same consistency strategy.

[0206] In a possible implementation manner, the device further includes:

[0207] A first writing module is configured to write, for any first-level merge request taken out from any first-level merge request buffer, the first-level merge request as a new second-level merge request into a second-level merge request buffer in response to the absence of a second-level merge request that satisfies the preset merge condition with the first-level merge request in a second-level merge request buffer corresponding to the first-level merge request buffer.

[0208] In a possible implementation, the processing module 83 is used to:

[0209] In response to the second-level merge request most recently written into the second-level merge request buffer carrying instruction end marker information, the second-level merge request is taken out from the second-level merge request buffer for processing, wherein the instruction end marker information is used to indicate the last data request in the request sequence.

[0210] In a possible implementation, the first merging module 81 is used to:

[0211] In response to the number of pending data requests in the pending request buffer being greater than or equal to L, taking out L pending data requests from the pending request buffer, where L is an integer greater than or equal to 2;

[0212] The L pending data requests are merged to obtain K first-level merged requests, and the K first-level merged requests are written into a first-level merged request buffer corresponding to the pending request buffer, where K is an integer greater than 0 and less than or equal to L.

[0213] In a possible implementation, the first merging module 81 is used to:

[0214] Among the L data requests to be processed, data requests whose requested data are located in the same cache line are merged to obtain K first-level merged requests.

[0215] In a possible implementation manner, the device further includes:

[0216] A determination module, used to determine a target cache unit corresponding to a data request to be processed from a plurality of cache units;

[0217] The second writing module is used to write the to-be-processed data request into a target to-be-processed request buffer corresponding to the target cache unit.

[0218] In a possible implementation manner, the determining module is used to:

[0219] Performing a hash operation on the request address of the data request to be processed to obtain a hash operation result corresponding to the data request to be processed;

[0220] According to the hash operation result, a target cache unit corresponding to the to-be-processed data request is determined from a plurality of cache units.

[0221] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the method described in the above method embodiments. Its specific implementation and technical effects can refer to the description of the above method embodiments. For the sake of brevity, they will not be repeated here.

[0222] The present disclosure also provides a computer-readable storage medium having computer program instructions stored thereon, wherein the computer program instructions implement the above method when executed by a processor. The computer-readable storage medium may be a non-volatile computer-readable storage medium, or may be a volatile computer-readable storage medium.

[0223] The embodiment of the present disclosure further provides a computer program, including a computer-readable code. When the computer-readable code is executed in an electronic device, a processor in the electronic device executes the above method.

[0224] The embodiments of the present disclosure also provide a computer program product, including a computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in an electronic device, a processor in the electronic device executes the above method.

[0225] An embodiment of the present disclosure also provides an electronic device, comprising: one or more processors; and a memory for storing executable instructions; wherein the one or more processors are configured to call the executable instructions stored in the memory to execute the above method.

[0226] The electronic device may be provided as a terminal, a server, or a device in other forms.

[0227] Fig. 9 1 is a block diagram of an electronic device 1900 provided in an embodiment of the present disclosure. For example, the electronic device 1900 may be provided as a server or a terminal device. Fig. 9 , the electronic device 1900 includes a processing component 1922, which further includes one or more processors, and a memory resource represented by a memory 1932 for storing instructions executable by the processing component 1922, such as an application. The application stored in the memory 1932 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to perform the above method.

[0228] The electronic device 1900 may also include a power supply component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and an input / output interface 1958 (I / O interface). The electronic device 1900 may operate based on an operating system stored in the memory 1932, such as a Microsoft Server operating system (Windows Server 2003). TM ), a graphical user interface operating system launched by Apple (MacOS X TM ), a multi-user, multi-process computer operating system (Unix TM ), a free and open source Unix-like operating system (Linux TM ), an open source Unix-like operating system (FreeBSD TM ) or similar.

[0229] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 1932 including computer program instructions, which can be executed by the processing component 1922 of the electronic device 1900 to perform the above method.

[0230] The present disclosure may be a system, a method and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.

[0231] A computer-readable storage medium may be a tangible device that can hold and store instructions used by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples of computer-readable storage media (a non-exhaustive list) include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disk read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punch card or a raised structure in a groove on which instructions are stored, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium is not to be interpreted as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse through a fiber optic cable), or an electrical signal transmitted through a wire.

[0232] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in the computer-readable storage medium in each computing / processing device.

[0233] The computer program instructions for performing the operation of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages, such as Smalltalk, C++, etc., and conventional procedural programming languages, such as "C" language or similar programming languages. Computer-readable program instructions may be executed completely on a user's computer, partially on a user's computer, as an independent software package, partially on a user's computer, partially on a remote computer, or completely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., using an Internet service provider to connect via the Internet). In some embodiments, an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may be customized by utilizing the state information of the computer-readable program instructions, and the electronic circuit may execute the computer-readable program instructions, thereby realizing various aspects of the present disclosure.

[0234] Various aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present disclosure. It should be understood that each box in the flowchart and / or block diagram and the combination of each box in the flowchart and / or block diagram can be implemented by computer-readable program instructions.

[0235] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processor of the computer or other programmable data processing device, a device that implements the functions / actions specified in one or more boxes in the flowchart and / or block diagram is generated. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause the computer, programmable data processing device, and / or other equipment to work in a specific manner, so that the computer-readable medium storing the instructions includes a manufactured product, which includes instructions for implementing various aspects of the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0236] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operating steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.

[0237] The flow chart and block diagram in the accompanying drawings show the possible architecture, function and operation of the system, method and computer program product according to multiple embodiments of the present disclosure. In this regard, each square box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and a part of the module, program segment or instruction includes one or more executable instructions for realizing the specified logical function. In some alternative implementations, the function marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous square boxes can actually be executed substantially in parallel, and they can sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs the specified function or action, or can be implemented with a combination of special hardware and computer instructions.

[0238] The computer program product may be implemented in hardware, software or a combination thereof. In one optional embodiment, the computer program product is embodied as a computer storage medium, and in another optional embodiment, the computer program product is embodied as a software product, such as a software development kit (SDK) and the like.

[0239] The above description of various embodiments tends to emphasize the differences between the various embodiments. The same or similar aspects can be referenced to each other, and for the sake of brevity, they will not be repeated herein.

[0240] If the technical solution of the embodiments of the present disclosure involves personal information, the product using the technical solution of the embodiments of the present disclosure has clearly informed the personal information processing rules and obtained the individual's voluntary consent before processing the personal information. If the technical solution of the embodiments of the present disclosure involves sensitive personal information, the product using the technical solution of the embodiments of the present disclosure has obtained the individual's separate consent before processing the sensitive personal information, and at the same time meets the requirement of "explicit consent". For example, at personal information collection devices such as cameras, clear and prominent signs are set to inform that the personal information collection scope has been entered and personal information will be collected. If the individual voluntarily enters the collection scope, it is deemed that he or she agrees to collect his or her personal information; or on the device for processing personal information, when the personal information processing rules are notified by obvious signs / information, the individual's authorization is obtained through pop-up information or by asking the individual to upload his or her personal information; the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the type of personal information processed.

[0241] The embodiments of the present disclosure have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and changes will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical applications, or improvements to the technology in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.

Claims

1. A data request processing method, characterized in that: The cache includes a plurality of cache units, and the plurality of cache units correspond one-to-one to a plurality of pending request buffers, wherein the pending request buffer corresponding to any cache unit is used to buffer pending data requests allocated to the cache unit, and the method includes: For any pending request buffer, merge the pending data requests in the pending request buffer to obtain at least one first-level merged request, and write the at least one first-level merged request into the first-level merged request buffer corresponding to the pending request buffer; For any first-level merge request taken out from any first-level merge request buffer, in response to any second-level merge request in a second-level merge request buffer corresponding to the first-level merge request buffer and the first-level merge request satisfying a preset merge condition, merge the first-level merge request with the second-level merge request to obtain a new second-level merge request; The second-level merge request is taken out from the second-level merge request buffer for processing.

2. The method according to claim 1, characterized in that The preset merging conditions include: The data requested by the first-level merge request and the second-level merge request are located in N adjacent cache lines, where N is an integer greater than or equal to 2.

3. The method according to claim 1, characterized in that The preset merging conditions include: The first-level merge request and the second-level merge request are data requests received within M adjacent clock cycles, where M is an integer greater than or equal to 2.

4. The method according to any one of claims 1 to 3, characterized in that The preset merging condition includes at least one of the following: The first-level merge request and the second-level merge request come from the same request end; The first-level merge request and the second-level merge request correspond to the same system-level cache policy; The first-level merge request and the second-level merge request correspond to the same consistency strategy.

5. The method according to any one of claims 1 to 3, characterized in that: The method further comprises: For any first-level merge request taken out from any first-level merge request buffer, in response to the fact that there is no second-level merge request that satisfies the preset merge condition with the first-level merge request in the second-level merge request buffer corresponding to the first-level merge request buffer, the first-level merge request is written into the second-level merge request buffer as a new second-level merge request.

6. The method according to any one of claims 1 to 3, characterized in that: The taking out the second-level merge request from the second-level merge request buffer for processing includes: In response to the second-level merge request most recently written into the second-level merge request buffer carrying instruction end marker information, the second-level merge request is taken out from the second-level merge request buffer for processing, wherein the instruction end marker information is used to indicate the last data request in the request sequence.

7. The method according to any one of claims 1 to 3, characterized in that The merging of the pending data requests in the pending request buffer to obtain at least one first-level merged request, and writing the at least one first-level merged request into a first-level merged request buffer corresponding to the pending request buffer, comprises: In response to the number of pending data requests in the pending request buffer being greater than or equal to L, taking out L pending data requests from the pending request buffer, where L is an integer greater than or equal to 2; The L pending data requests are merged to obtain K first-level merged requests, and the K first-level merged requests are written into a first-level merged request buffer corresponding to the pending request buffer, where K is an integer greater than 0 and less than or equal to L.

8. The method according to claim 7, characterized in that The step of merging the L to-be-processed data requests to obtain K first-level merged requests includes: Among the L data requests to be processed, data requests whose requested data are located in the same cache line are merged to obtain K first-level merged requests.

9. The method according to any one of claims 1 to 3, characterized in that: The method further comprises: Determine, from a plurality of cache units, a target cache unit corresponding to the data request to be processed; The to-be-processed data request is written into a target to-be-processed request buffer corresponding to the target cache unit.

10. The method according to claim 9, characterized in that The step of determining, from among the plurality of cache units, a target cache unit corresponding to the data request to be processed comprises: Performing a hash operation on the request address of the data request to be processed to obtain a hash operation result corresponding to the data request to be processed; According to the hash operation result, a target cache unit corresponding to the to-be-processed data request is determined from a plurality of cache units.

11. A data request processing device, characterized in that: The cache includes a plurality of cache units, and the plurality of cache units correspond one-to-one to a plurality of pending request buffers, wherein the pending request buffer corresponding to any cache unit is used to buffer pending data requests allocated to the cache unit, and the device includes: A first merging module is configured to merge the pending data requests in any pending request buffer to obtain at least one first-level merged request, and write the at least one first-level merged request into a first-level merged request buffer corresponding to the pending request buffer; a second merging module, configured to, for any first-level merge request taken out from any first-level merge request buffer, merge the first-level merge request with the second-level merge request in response to any second-level merge request in a second-level merge request buffer corresponding to the first-level merge request buffer and the first-level merge request satisfying a preset merging condition, to obtain a new second-level merge request; A processing module is used to take out the second-level merge request from the second-level merge request buffer for processing.

12. An electronic device, characterized in that: include: one or more processors; a memory for storing executable instructions; The one or more processors are configured to call the executable instructions stored in the memory to execute the method according to any one of claims 1 to 10.

13. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 10 is implemented.

14. A computer program product comprising computer readable code, or a non-computer program product carrying computer readable code A volatile computer-readable storage medium, characterized in that When the computer readable code is executed in an electronic device, The processor in the electronic device executes the method according to any one of claims 1 to 10.

Citation Information

Cited By

  • Request processing method, secondary merging device and request processing system

    CN120803973A

  • Data access method, cache device, chip product and computer equipment

    CN120929222A

  • Data access method, cache device, chip product and computer equipment

    CN120929222B