Test method and device, equipment, storage medium and program product

By generating test cases based on preset merging relationships, the problem of low iteration efficiency of test cases in the prior art is solved, and accurate testing and efficient iteration of the merging function of the memory merging module are realized.

CN120045397APending Publication Date: 2025-05-27MOORE THREADS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510192553.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

When testing memory merging modules, the test cases need to be redeveloped after configuration changes, and the iteration efficiency is low.

Method used

By obtaining test cases and expected merge results, each request and expected merge results are determined based on the merge relationship between a preset number of requests, and input these requests into the memory merge module to obtain the actual merge results for verification.

Benefits of technology

It realizes the generation of test cases based on preset merging relationships, tests typical merging scenarios, accurately tests the merging function of the memory merging module, and improves the iterative efficiency of the test.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045397A_ABST
    Figure CN120045397A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a test method and device, equipment, a storage medium and a program product, and the method comprises the steps: obtaining a test case and an expected merging result of a memory merging module; each request in the test case and the expected merging result are determined based on a merging relationship among a preset number of requests; the merging relation is used for representing whether adjacent requests in the preset number of requests are merged or not; inputting each request in the test case into the memory merging module to obtain an actual merging result; and verifying the actual merging result based on the expected merging result to obtain a test result of the memory merging module. Therefore, the test case is generated based on the preset merging relation, and the typical merging scene is tested, so that the request merging behavior of the test case can be predicted in advance, the merging function of the memory merging module can be accurately tested, and the iteration efficiency of the test is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of chip testing, and in particular, to a testing method, device, equipment, storage medium, and program product. Background Art

[0002] In the design of parallel acceleration chips such as graphics processing units (GPUs) and neural processing units (NPUs), the access to the data path (operations such as loading and storing memory access operations issued by GPU cores and NPU cores to access the memory subsystem) is generally processed in parallel. To improve the efficiency and throughput of data access, there is generally a memory merge module. When multiple threads access adjacent memory locations, multiple access requests can be merged, reducing the number of requests, which can greatly improve the bandwidth utilization of the system.

[0003] For the testing of the memory merge module, the general testing method is to develop corresponding test cases according to the current configuration. Once a certain configuration changes, the test cases need to be redeveloped, and the iteration efficiency of this testing method is relatively low. Summary of the Invention

[0004] In view of this, at least one testing method, device, equipment, storage medium, and program product are provided in the embodiments of this application.

[0005] The technical solution of the embodiments of this application is implemented as follows:

[0006] On the one hand, the embodiments of this application provide a testing method, and the method includes: obtaining a test case of a memory merge module and an expected merge result; each request in the test case and the expected merge result are determined based on the merge relationship between a preset number of requests; the merge relationship is used to represent whether adjacent requests among the preset number of requests are merged; inputting each request in the test case into the memory merge module to obtain an actual merge result; and verifying the actual merge result based on the expected merge result to obtain the test result of the memory merge module.

[0007] On the other hand, an embodiment of the present application provides a testing device, which includes: an acquisition module for acquiring test cases and expected merging results of a memory merging module; each request in the test cases and the expected merging results are determined based on the merging relationship between a preset number of requests; the merging relationship is used to represent whether adjacent requests among the preset number of requests are merged; an input module for inputting each request in the test cases into the memory merging module to obtain an actual merging result; a verification module for verifying the actual merging result based on the expected merging result to obtain the test result of the memory merging module.

[0008] In yet another aspect, an embodiment of the present application provides a computer device, including a memory and a processor. The memory stores a computer program that can run on the processor, and when the processor executes the program, it implements some or all of the steps in the above method.

[0009] In still another aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements some or all of the steps in the above method.

[0010] In still another aspect, an embodiment of the present application provides a computer program product, including a computer program or instruction. When the computer program or instruction is executed by a processor, it implements some or all of the steps in the above method.

[0011] In the embodiments of the present application, each request and the expected merging result in the test cases are determined based on the merging relationship between a preset number of requests, and the actual merging result is verified based on the expected merging result to obtain the test result of the memory merging module. In this way, test cases are generated based on the preset merging relationship to test typical merging scenarios, the request merging behavior of the test cases can be predicted in advance, the merging function of the memory merging module can be accurately tested, and the iteration efficiency of the test can be improved.

[0012] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the technical solution of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification. These drawings illustrate embodiments consistent with the present application and, together with the specification, are used to explain the technical solution of the present application.

[0014] Figure 1 Schematic diagram of the implementation process of a testing method provided by an embodiment of the present application Figure 1 ;

[0015] Figure 2Schematic implementation process of a test method provided by an embodiment of the present application Figure 2 ;

[0016] Figure 3 Schematic implementation process of a test method provided by an embodiment of the present application Figure 3 ;

[0017] Figure 4 Schematic implementation process of a test method provided by an embodiment of the present application Figure 4 ;

[0018] Figure 5 Schematic implementation process of a test method provided by an embodiment of the present application Figure 5 ;

[0019] Figure 6 Schematic diagram of the composition structure of a test device provided by an embodiment of the present application;

[0020] Figure 7 Schematic diagram of the hardware entity of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0021] In order to make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions of the present application will be further elaborated in detail below with reference to the accompanying drawings and embodiments. The described embodiments should not be construed as limiting the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present application.

[0022] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict. The terms "first / second / third" involved are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when allowed, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0023] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing this application and are not intended to limit this application.

[0024] An embodiment of this application provides a testing method, which can be executed by a processor of a computer device. Herein, the computer device may refer to a device with data processing capabilities such as a server, a laptop computer, a tablet computer, a desktop computer, a smart TV, a set-top box, a mobile device (such as a mobile phone, a portable video player, a personal digital assistant, a dedicated messaging device, a portable gaming device), etc.

[0025] Figure 1 Schematic diagram of the implementation process of a testing method provided by an embodiment of this application Figure 1 , such as Figure 1 shown, this method includes the following steps S101 to step S103:

[0026] Step S101, obtain test cases of the memory merging module and expected merging results; each request in the test cases and the expected merging results are determined based on the merging relationship between a preset number of requests; the merging relationship is used to represent whether adjacent requests among the preset number of requests are merged.

[0027] Among them, the memory merging module is a software module in the chip, which is used to merge requests for a thread to access memory. Exemplarily, multiple threads of a GPU access the same address within 32 bytes. Assuming there are 4 threads, the addresses of the 4 access requests are 0x10000, 0x10004, 0x10008, 0x1000c respectively, and the length of each request is 4 bytes. Then these 4 requests can be merged into the same request, with the address being 0x10000 and the length being 16 bytes.

[0028] Among them, the expected merging result is the merging result obtained by the memory merging module normally merging the requests in the test cases. The expected merging result is a predicted value, and based on the test cases and the merging method of the memory merging module, the expected merging result corresponding to the test cases can be determined in advance. Exemplarily, the expected merging result is the number of request mergings output by the memory merging module.

[0029] Among them, the memory merging module can parallel process requests of N threads in the same cycle. That is to say, at most N requests of a memory merging module are merged together in the same cycle.

[0030] Among them, when generating test cases and expected merging results, it is necessary to determine the merging relationship between requests in advance. Each request in the test cases and the expected merging results are determined based on the merging relationship between a preset number of requests. The merging relationship represents whether adjacent requests among the preset number of requests are merged. The size of the preset number can be the maximum number of requests N that the memory merging module parallel processes, or other values less than the maximum number of requests N.

[0031] In some embodiments, the merging relationship between the preset number of requests is used to represent that the preset number of requests are divided into P sets based on the merging factor K, and the requests within one set correspond to one merging result output by the memory merging module.

[0032] Among them, the merging relationship includes at least one of the following: the first merging relationship, the second merging relationship. Among them, the first merging relationship can be understood as that among the preset number of requests, every K requests are merged; the second merging relationship can be understood as that among the preset number of requests, only K requests are merged; K is a positive integer and less than or equal to the preset number. The first merging relationship ensures that as many requests as possible are merged together, and the second merging relationship ensures that as few requests as possible are merged together. When K is equal to the preset number, for K requests, the results after merging through the first merging relationship and the second merging relationship are the same, both merging the K requests together.

[0033] To understand the two situations of merging every K requests among the above-mentioned preset number of requests and only K requests among the preset number of requests, the following will illustrate this with the description of the division results: The first merging relationship is used to represent the existence of P - 1 first sets and 1 second set; the second merging relationship is used to represent the existence of 1 first set and N - K third sets.

[0034] Among them, each first set includes K requests, the second set includes O requests, 1 ≤ O < K, each third set includes 1 request, and N is the preset number.

[0035] Exemplarily, the preset number is 8, and the above-mentioned merging factor K can take positive integer values between 1 and 8. When the merging factor K is 2, for the first merging relationship, the 8 requests can be divided into 4 sets based on the merging factor K. Among them, there are 3 first sets, including: the set where the 1st request and the 2nd request are merged together, the set where the 3rd request and the 4th request are merged together, and the set where the 5th request and the 6th request are merged together; there is 1 second set, including the set where the 7th request and the 8th request are merged together. Similarly, assuming the preset number is 7 and other parameters remain unchanged, the first sets remain the same, and the second set is one, including the set where the 7th request exists. It can be seen that the number of requests in the first sets is all the merging factor K, and the number of requests in the second set is less than or equal to the merging factor K.

[0036] In the case where the merging factor K is 2, for the second merging relationship, among the 8 requests, only 2 requests are merged. That is to say, the first set is 1, including the set formed by merging the 1st request and the 2nd request; the third set is 6, and each third set includes 1 request, which are the 3rd request to the 8th request respectively. It can be seen that since only K requests can be merged, there is only 1 first set, and the remaining requests cannot be merged. Each 1 request can be used as a third set. Therefore, the number of the third sets is N - K.

[0037] In the embodiments of the present application, through the first merging relationship, among a preset number of requests, every K requests are merged; through the second merging relationship, among a preset number of requests, only K requests are merged; K is a positive integer and less than or equal to the preset number. In this way, by constructing the first merging relationship and the second merging relationship, two typical merging scenarios can be covered, and test cases can be generated targeted to verify the accuracy of the memory merging module.

[0038] In some embodiments, configuration parameters of the memory merging module are preset. The configuration parameters include the number W of memory merging modules, the request allocation amount M, the request processing threshold N, the request length L, and the merging factor K. In actual application, the number of memory merging modules can be one or at least two. When the number of memory merging modules is at least two, at least two memory merging modules jointly process multiple memory access requests and merge the requests. The request allocation amount M represents the number of requests allocated to a single memory merging module. The request processing threshold N represents the maximum number of requests that a memory merging module processes in parallel at a time. The merging factor K represents the number of preset requests to be merged among N requests. The request length L is the length of the memory address accessed by a single request. In the embodiments of the present application, the request lengths of the addresses of all requests are the same.

[0039] Among them, the total number of requests in the test case is W * M. A memory merging module processes N requests in parallel at a time. M is a positive integer multiple of N. A memory merging module processes M requests and needs to execute M / N parallel processing processes. For every N requests, the memory merging module selects every K requests to be merged or only K requests to be merged based on the merging relationship.

[0040] Among them, each request in the test case includes a destination request address and a request length. Based on the target request address and the request length, the address accessed by the request can be determined. Exemplarily, the target request address of a request is 0x10000 and the request length is 4 bytes. The memory address accessed by this request is from 0x10000 to 0x10003.

[0041] In some embodiments, the number of test cases is one. After determining the merging relationship, the number W of memory merging modules, the request allocation amount M, the request processing threshold N, and the merging factor K, a test case corresponding to the configuration parameters is generated based on the determined configuration parameters.

[0042] In some embodiments, the number of test cases is at least two. In the case where the merging relationship is determined, each merging factor corresponds to one test case. Exemplarily, the request processing threshold N is 8, and the merging factor K takes values from 1 to 8, and 8 test cases need to be generated.

[0043] In some embodiments, a test case includes the target request address corresponding to the initial request identifier of each request. Obtaining the test case of the memory merging module includes: in the case where the merging relationship is the first merging relationship, dividing the N requests processed in parallel by the memory merging module each time into two categories: the first category of requests and the second category of requests, then determining the request category corresponding to the initial request identifier, and determining the target request address based on the mapping method corresponding to the request category. The first category of requests refers to the requests that can be completely merged based on the merging factor, and the second category of requests refers to the requests that cannot be completely merged based on the merging factor. The mapping method is a pre-set method for determining the target request address from the initial request identifier, and the mapping method corresponds to the request category. The mapping method needs to ensure that for the N requests processed in parallel each time, as many requests as possible are merged based on the merging factor. Exemplarily, the request processing threshold N is 9, and the merging factor K is 2. For 9 requests, every two requests are merged, and the last request cannot be merged. Therefore, the first 8 requests can be completely merged, and the last request cannot be completely merged, that is, the first 8 requests are the first category of requests, and the 9th request is the second category of requests.

[0044] In some embodiments, obtaining the test case of the memory merging module includes: in the case where the merging relationship is the second merging relationship, dividing the N requests processed in parallel by the memory merging module each time into two categories: the third category of requests and the fourth category of requests, then determining the request category corresponding to the initial request identifier, and determining the target request address based on the mapping method corresponding to the request category. The third category of requests refers to the requests that are merged, and the fourth category of requests refers to the requests that are not merged. The mapping method needs to ensure that for the N requests processed in parallel each time, as few requests as possible are merged based on the merging factor. Exemplarily, the request processing threshold N is 9, and the merging factor K is 2. For 9 requests, only the first 2 requests can be merged, and the remaining 7 requests cannot be merged. Therefore, the first 2 requests are the third category of requests, and the last 7 requests are the fourth category of requests.

[0045] In some embodiments, it is desirable that the merged result is the number of merged requests output by the memory merging module. Based on the number of memory request modules W, the request allocation amount M, the request processing threshold N, the merging relationship, and the merging factor K, the number of merged requests output by the memory merging module can be determined. Exemplarily, when the merging relationship is the first merging relationship, the number of merged requests output by the memory merging module is W*(M / N)*ceil(N / K), where ceil() is the ceiling function. When the merging relationship is the second merging relationship, the number of merged requests output by the memory merging module is W*(M / N)*(N-K+1).

[0046] In some embodiments, it is desirable that the merged result is the address constraint condition of the merged requests output by the memory merging module. The output merged requests need to satisfy the address constraint condition. Exemplarily, the address constraint condition is that the target request address <W*M*L, that is, the target request address of the requests output by the memory merging module should be less than W*M*L.

[0047] Step S102: Input each request in the test case into the memory merging module to obtain the actual merged result.

[0048] Among them, after obtaining the test case, the target request address and request length of each request in the test case can be obtained. Based on the target request address and request length of each request, each request in the test case is input into the memory merging module. The memory merging module merges the requests based on the target request address and request length of each request and outputs the merged requests. Specifically, the memory merging module first determines the merging boundary corresponding to each request, and only the requests falling within the same merging boundary can be merged. Within the same merging boundary, the memory merging module further determines whether the memory addresses accessed by the requests are continuous, and merges the requests with continuous addresses into one request. Each request in the test case can be sent by a thread.

[0049] In some embodiments, the actual merged result is the number of requests actually output by the memory merging module. After inputting each request in the test case into the memory merging module, the memory merging module executes the request merging process based on the access address corresponding to each request, merges the requests that can be merged together, and finally outputs several merged requests. The number of the merged requests is used as the actual merged result.

[0050] In some embodiments, the actual merging result is the target request address of the request output by the memory merging module. After the memory merging module executes the request merging process, it outputs the target request address of each merged request. When the memory merging module performs normal merging, the target request address of the output request needs to meet certain constraint conditions. Therefore, the target request address of the request output by the memory merging module is used as the actual merging result.

[0051] Step S103: Verify the actual merging result based on the expected merging result to obtain the test result of the memory merging module.

[0052] Among them, in order to test whether the request merging function of the memory merging module is normal, it is necessary to verify the actual merging result output by the memory merging module based on the expected merging result to obtain the test result of the memory merging module.

[0053] In some embodiments, the expected merging result is the number of requests expected to be output by the memory merging module, and the actual merging result is the number of requests actually output by the memory merging module. Compare the expected number of requests with the actual number of requests. When the expected number of requests is the same as the actual number of requests, it is determined that the test result of the memory merging module is passed; when the expected number of requests is different from the actual number of requests, it is determined that the test result of the memory merging module is not passed.

[0054] In some embodiments, the expected merging result is the constraint condition of the target request address of the request output by the memory merging module, and the actual merging result is the target request address of the request output by the memory merging module. Compare the target request address of each output request with the constraint condition to obtain the comparison result of each output request. The comparison result includes that the target request address meets the constraint condition and the target request address does not meet the constraint condition. When the comparison results of all output requests are that the target request address meets the constraint condition, it is determined that the test result of the memory merging module is passed; when the comparison result of at least one output request is that the target request address does not meet the constraint condition, it is determined that the test result of the memory merging module is not passed.

[0055] In the embodiments of the present application, each request and the expected merging result in the test case are determined based on the merging relationship between a preset number of requests, and the actual merging result is verified based on the expected merging result to obtain the test result of the memory merging module. In this way, test cases are generated based on the preset merging relationship to test typical merging scenarios, the request merging behavior of the test cases can be predicted in advance, the merging function of the memory merging module can be accurately tested, and the iteration efficiency of the test can be improved.

[0056] Figure 2 It is a schematic implementation process of a test method provided by the embodiments of the present applicationFigure 2 , which can be executed by a processor of a computer device. Based on Figure 1 , the test case includes a target request address corresponding to the initial request identifier of each request; Figure 1 in "obtaining test cases for the memory merging module" can be updated to S201 to S203, and will be described in combination with Figure 2 the steps shown.

[0057] Step S201, based on the merging relationship, configuration parameters, and the initial request identifier, determine a first allocation amount and a second allocation amount.

[0058] Among them, the test case includes at least two requests, and each request includes an initial request identifier and a destination request address. The initial request identifier represents the sequence number of the request, and the destination request address represents the starting address of the request to access the memory. Exemplarily, the initial request identifier of the first request is 0, the initial request identifier of the second request is 1, the target request address of the first request is 0x0, and the target request address of the second request is 0x4.

[0059] Among them, the configuration parameters are related parameters for initialization, and are used for the initialization process of the memory merging module to perform request merging. Exemplarily, the configuration parameters include the number of memory merging modules, request allocation amounts, request processing thresholds, request lengths, and merging factors. The number of requests in the test case can be determined by the number of memory merging modules and the request allocation amounts, the number of requests allocated to a single memory merging module can be determined by the request allocation amounts, the number of requests processed in parallel by the memory merging module at a single time can be determined by the request processing threshold, the address length corresponding to a single request can be determined by the request length, and the number of requests to be merged can be determined by the merging factor.

[0060] In some embodiments, when determining the target request address of each request in the test case, it is necessary to first map the initial request identifier to a memory request identifier, and determine the corresponding target request address based on the memory request identifier. The initial request identifier can be regarded as the sequence number of each request before merging, and the memory request identifier can be regarded as the sequence number after the initial request identifier is mapped based on the mapping relationship. The principle of the mapping relationship is that for N requests processed in parallel by the memory merging module at a single time, the memory request identifiers of the requests that can be merged together are continuous, and the memory request identifiers of the requests that cannot be merged together are not continuous. In this way, the target request address is determined based on the memory request identifier, so that the access addresses of the requests merged together are continuous, and the access addresses of the requests that cannot be merged together are not continuous.

[0061] Exemplarily, for the first merging relationship, when the requested allocation amount is 1024, the request processing threshold is 8, and the merging factor is 2, the memory request identifiers corresponding to the initial request identifiers of each request are shown in Table 1. As can be seen from Table 1, for the 8 requests processed in parallel by the memory merging module, every two requests are merged, and the memory request identifiers of the requests to be merged are consecutive.

[0062] Table 1

[0063] Initial request identifier Memory request identifier 0 0 1 1 2 256 3 257 4 512 5 513 6 768 7 769 8 2 9 3 10 258 11 259 12 514 13 515 14 770 15 771 … …

[0064] In some embodiments, in order to be able to determine the corresponding memory request identifier based on the initial request identifier, it is necessary to determine the number of memory request identifiers occupied by other requests before the memory request identifier corresponding to this initial request identifier. Exemplarily, for the initial request identifier 12, the finally determined corresponding memory request identifier is 514, indicating that the memory request identifiers from 0 to 513 are occupied by the remaining requests. Therefore, when determining the memory request identifier corresponding to the initial request identifier 12, it is necessary to calculate the number of memory request identifiers occupied by other requests.

[0065] Among them, for each initial request identifier, the number of memory request identifiers occupied by other requests is determined by two parts: the first allocation amount and the second allocation amount. The requests corresponding to each request processing threshold are divided into several merging items, and each merging item includes several requests merged together. Exemplarily, for the first merging relationship, when the request processing threshold is 8 and the merging factor is 2, every 2 requests are merged. Therefore, the 8 requests are divided into 4 merging items, and each merging item includes 2 requests. The first allocation amount represents the number of memory request identifiers occupied by other requests before the current merging item where the initial request identifier is located; the second allocation amount represents the number of memory request identifiers occupied by other requests before the current initial request identifier within the current merging item where the initial request identifier is located. Exemplarily, as shown in Table 1, for the initial request identifier 12, the first allocation amount is 514 and the second allocation amount is 0; for the initial request identifier 13, the first allocation amount is 514 and the second allocation amount is 1.

[0066] In some embodiments, the configuration parameters include a request allocation quantity M, a request processing threshold N, and a merging factor K. In the case where the merging relationship is the first merging relationship, divide the request processing threshold by the merging factor, and then multiply by the merging factor to obtain the starting unmerged request identifier t_remain_start_id_in_warp within the sub-request group; divide the initial request identifier by the request allocation quantity to obtain the request group identifier block_id; take the remainder of the initial request identifier with respect to the request allocation quantity to obtain the in-block request identifier thread_id_in_block; based on the starting unmerged request identifier within the sub-request group, the request group identifier, the in-block request identifier, the request allocation quantity, and the request processing threshold, determine the first allocation quantity.

[0067] Exemplarily, the first allocation quantity is (thread_id_in_block / / N)*(N - t_remain_start_id_in_warp)+t_remain_start_id_in_warp*(M / / N)+block_id*M, or (thread_id_in_block / / N)*K+((thread_id_in_block % N) / / K)*(M / / N)*K+block_id*M. Wherein, " / / " is the integer division operation, and "%" is the remainder operation.

[0068] In some embodiments, the configuration parameters include a request allocation quantity M, a request processing threshold N, and a merging factor K. In the case where the merging relationship is the second merging relationship, take the remainder of the in-block request identifier with respect to the request processing threshold to obtain the in-sub-request-group request identifier thread_id_in_N; divide the in-block request identifier by the request processing threshold to obtain the sub-request group identifier warp_id, and based on the request group identifier block_id, the in-block request identifier thread_id_in_block, the in-sub-request-group request identifier thread_id_in_N, the sub-request group identifier warp_id, the request allocation quantity M, and the request processing threshold N, determine the first allocation quantity.

[0069] Exemplarily, the first allocation quantity is warp_id*K+block_id*M or M / / N*thread_id_in_N+warp_id+block_id*M.

[0070] Wherein, the second allocation quantity is used to represent the number of memory request identifiers occupied by other requests before the current initial request identifier within the current merging item. A merging item is the smallest merging unit, and there is at least one request in a merging item.

[0071] In some embodiments, when the merging relationship is the first merging relationship and the merging attribute is that the requests corresponding to the initial request identifier are not merged, the second allocation amount is determined based on the request identifier in the block thread_id_in_block, the starting unmerged request identifier in the sub-request group t_remain_start_id_in_warp, and the request processing threshold N. In some embodiments, when the merging relationship is the first merging relationship and the merging attribute is that the requests corresponding to the initial request identifier are merged, the second allocation amount is determined based on the request identifier in the block thread_id_in_block, the request processing threshold N, and the merging factor K.

[0072] In some embodiments, when the merging relationship is the second merging relationship and the merging attribute is that the requests corresponding to the initial request identifier are merged, the request identifier thread_id_in_N in the sub-request group is used as the second allocation amount. In some embodiments, when the merging relationship is the second merging relationship and the merging attribute is that the requests corresponding to the initial request identifier are not merged, the second allocation amount is 0.

[0073] Step S202: Determine the memory request identifier based on the first allocation amount and the second allocation amount.

[0074] Among them, the first allocation amount and the second allocation amount are added to obtain the memory request identifier corresponding to the initial request identifier.

[0075] Step S203: Determine the target request address corresponding to the initial request identifier based on the merging relationship, the configuration parameter, the initial request identifier, the first allocation amount, and the memory request identifier.

[0076] Among them, the memory request identifiers of the requests merged together are consecutive, and the memory request identifiers of the requests not merged together are not consecutive. Since the request length of each request in the embodiments of the present application is fixed, the target request address can be determined based on the memory request identifier and the request length.

[0077] In some embodiments, after obtaining the memory request identifier of each request, multiply the memory request identifier by the request length to obtain the target request address, and determine the address for the request to access the memory based on the target request address and the request length. Exemplarily, for the first merging relationship, with a request allocation amount of 1024, a request processing threshold of 8, a merging factor of 2, and a request length of 4, the target request addresses corresponding to the initial request identifiers of each request are shown in Table 2. For the initial request identifier 0, the target request address is 0x0, the request length is 4, and the corresponding address for the request to access the memory is from 0x0 to 0x3; for the initial request identifier 1, the target request address is 0x4, the request length is 4, and the corresponding address for the request to access the memory is from 0x4 to 0x7. It can be seen that at this target request address, the memory addresses accessed by the requests to be merged are also consecutive.

[0078] Table 2

[0079] Initial request identifier Memory request identifier Target request address 0 0 0x0 1 1 0x4 2 256 0x400 3 257 0x404

[0080] In some embodiments, the addresses accessed by the requests to be merged must be within the same merging boundary, and each merging boundary corresponds to an address range. Exemplarily, the merging boundary corresponds to the address range 0x0000 to 0x0020, and only the requests with access addresses within the address range 0x0000 to 0x0020 can be merged together. The starting address of the merging boundary is the boundary starting address, the ending address is the boundary ending address, and the length of the merging boundary is the boundary granularity. Each request has a corresponding merging boundary. To ensure that the access address of each request is within the corresponding merging boundary, it is necessary to determine the boundary ending address corresponding to the initial request identifier, and adjust the product of the memory request address and the request length based on the boundary ending address to obtain the final target request address.

[0081] Among them, the boundary ending address is determined based on the first allocation amount, the boundary granularity, and the request length. Exemplarily, multiply the first allocation amount by the request length to obtain the boundary starting address, divide the boundary starting address by the boundary granularity, multiply the result by the boundary granularity, and then add the boundary granularity to obtain the boundary ending address.

[0082] In the embodiments of the present application, the first allocation amount is determined based on the merging relationship, the configuration parameters, and the initial request identifier, the memory request identifier is determined in combination with the first allocation amount, and then the target request address is determined in combination with the memory request identifier. In this way, by first determining the first allocation amount and the memory request identifier to determine the target request address, the target request address of each request can be effectively determined.

[0083] Figure 3 is a schematic implementation process of a testing method provided by the embodiments of the present application Figure 3 , and this method can be executed by the processor of a computer device. Based on Figure 2, the configuration parameters include: a request processing threshold, which is used to represent the maximum number of requests processed by the memory merging module at a single time; a request allocation quantity, which is used to represent the number of requests allocated to the memory merging module; a merging factor, which is used to represent the maximum number of requests in a set. Figure 2 S201 in Figure 3 can be updated to S301 to S303, and will be described in conjunction with

[0084] Step S301: Based on the initial request identifier and the request allocation quantity, determine the in-block request identifier; the in-block request identifier is used to represent the bit order of the request corresponding to the initial request identifier in the corresponding request group; wherein, each request group corresponds to a memory merging module, and the request group includes all requests allocated to the memory merging module.

[0085] Among them, the request allocation quantity is the number of requests that a single memory merging module needs to process. Exemplarily, the request allocation quantity is 1024, indicating that the memory merging module needs to perform merging processing on 1024 requests.

[0086] Among them, each memory merging module corresponds to a request group. The request group includes all requests allocated to the memory merging module, and the number of requests in the request group is the request allocation quantity. Exemplarily, the request allocation quantity is 1024, the initial request identifier is from 0 to 2047, and the number of memory merging modules is 2. Therefore, there are 2 request groups in total. The first request group includes requests corresponding to the initial request identifier from 0 to 1023, and the second request group includes requests corresponding to the initial request identifier from 1024 to 2047.

[0087] Among them, the in-block request identifier is the bit order of the request corresponding to the initial request identifier in the corresponding request group. Perform a modulo operation on the initial request identifier with the request allocation quantity, and use the result as the in-block request identifier. Exemplarily, the first request group includes requests corresponding to the initial request identifier from 0 to 1023, the second request group includes requests corresponding to the initial request identifier from 1024 to 2047, the request allocation quantity is 1024. For the request with the initial request identifier of 1020, perform a modulo operation on 1020 with 1024, and the obtained in-block request identifier is 1020, indicating that this request is in the 1021st position in the first request group; for the request with the initial request identifier of 1030, perform a modulo operation on 1030 with 1024, and the obtained in-block request identifier is 6, indicating that this request is in the 7th position in the second request group.

[0088] Step S302: Based on the merging relationship, the merging factor, the request processing threshold, and the in-block request identifier, determine the merging attribute corresponding to the initial request identifier; the merging attribute is used to determine whether the request corresponding to the initial request identifier is merged.

[0089] Among them, the merging factor represents the number of requests to be merged in a merging relationship. In the first merging relationship, among the requests corresponding to the request processing threshold, every merging factor requests are merged; in the second merging relationship, among the requests corresponding to the request processing threshold, only the merging factor requests are merged. The request processing threshold is the number of requests that the memory merging module processes in parallel at a time.

[0090] From the perspective of the merging result, under a merging relationship, a preset number of requests can be divided into a first set and a second set, or divided into a first set and a third set. The requests within each set will be merged. Based on the foregoing embodiments, the maximum number of sets within the above sets is the merging factor, and at the same time, it is also the number of requests within the first set.

[0091] To more clearly illustrate the role of the above merging factor, the merging factor will be described below by taking the partitioning processes corresponding to the first merging relationship and the second merging relationship respectively. For the first merging relationship, among the preset number of requests, the merging factor requests are sequentially taken out as the first set until the number of remaining requests is less than or equal to the merging factor, and the remaining requests are taken as the second set; for the second merging relationship, among the preset number of requests, the merging factor requests are taken out as 1 first set, and for the remaining all requests, according to the rule that one request forms one third set, the third set is obtained. Thus, it can be seen that the merging factor is used to represent the maximum number of requests within a set.

[0092] Among them, the merging attribute is used to determine whether the request corresponding to the initial request identifier is merged. Based on the merging attribute, each request can be divided into two categories. For the first merging relationship, the requests that are completely merged based on the merging factor are determined as the first category of requests, and the requests that cannot be completely merged based on the merging factor are determined as the second category of requests. Exemplarily, the request processing threshold is 11 and the merging factor is 3. Every 3 requests are merged, and the last 2 requests are merged together. Therefore, the first 9 requests are determined as the first category of requests, and the last 2 requests are determined as the second category of requests. For the second merging relationship, the requests that are merged based on the merging factor are determined as the third category of requests, and the requests that cannot be merged based on the merging factor are determined as the fourth category of requests. Exemplarily, the request processing threshold is 11 and the merging factor is 3. Only 3 requests are merged. Therefore, the first 3 requests are determined as the third category of requests, and the last 8 requests are determined as the fourth category of requests.

[0093] In some embodiments, when the merging relationship is the first merging relationship, the starting unmerged request identifier t_remain_start_id_in_warp within the sub-request group is determined based on the request processing threshold N and the merging factor K, and the merging attribute is determined based on the request processing threshold N, the starting unmerged request identifier t_remain_start_id_in_warp within the sub-request group, and the request identifier thread_id_in_block within the block.

[0094] Among them, the starting unmerged request identifier within the sub-request group is used to represent the bit order of the first second-type request within the sub-request group. Each request group includes several sub-request groups, and each sub-request group includes the number of requests equal to the request processing threshold. Exemplarily, the request group includes 1024 requests with the initial request identifier ranging from 0 to 1023, and the request processing threshold is 8. Therefore, the request group includes 128 sub-request groups. Exemplarily, the expression of the starting unmerged request identifier t_remain_start_id_in_warp within the sub-request group is (N / / K)*K. When N is 9 and K is 2, t_remain_start_id_in_warp is 8, indicating that the first second-type request within the sub-request group is the 9th request within the sub-request group.

[0095] Among them, when N!= t_remain_start_id_in_warp and (thread_id_in_block % N) >= t_remain_start_id_in_warp, it is determined that the request corresponding to the initial request identifier is a second-type request, that is, the merging attribute is that the request corresponding to the initial request identifier is not merged. In other cases, it is determined that the request corresponding to the initial request identifier is a first-type request, that is, the merging attribute is that the request corresponding to the initial request identifier is merged. N!= t_remain_start_id_in_warp indicates that there is a second-type request in the corresponding sub-request group, and (thread_id_in_block % N) >= t_remain_start_id_in_warp indicates that the current request is a second-type request.

[0096] In some embodiments, when the merging relationship is the second merging relationship, the identifier thread_id_in_N within the sub-request group is determined based on the request processing threshold N and the request identifier thread_id_in_block within the block, and the merging attribute is determined based on the identifier thread_id_in_N within the sub-request group and the merging factor K.

[0097] Among them, the bit order of the request with the initial request identifier in the corresponding sub-request group. Exemplarily, the expression of thread_id_in_N in the sub-request group is thread_id_in_block % N. Exemplarily, thread_id_in_block is 2 and N is 8, and the calculated thread_id_in_N is 2, indicating that the request with the initial request identifier is in the third position in the corresponding sub-request group.

[0098] Among them, when thread_id_in_N < K, it is determined that the request corresponding to the initial request identifier is a third type of request, that is, the merging attribute is to merge the request corresponding to the initial request identifier. In other cases, it is determined that the request corresponding to the initial request identifier is a fourth type of request, that is, the merging attribute is that the request corresponding to the initial request identifier is not merged.

[0099] Step S303: Determine the first allocation amount based on the merging attribute, the merging relationship, the initial request identifier, the in-block request identifier, the request processing threshold, and the request allocation amount.

[0100] In some embodiments, when the merging relationship is the first merging relationship and the merging attribute is that the request corresponding to the initial request identifier is not merged, determine the third allocation amount based on the in-block request identifier thread_id_in_block, the starting unmerged request identifier t_remain_start_id_in_warp in the sub-request group, and the request processing threshold N; determine the fourth allocation amount based on the starting unmerged request identifier t_remain_start_id_in_warp in the sub-request group, the request allocation amount M, and the request processing threshold N; determine the fifth allocation amount based on the initial request identifier init_thread_id and the request allocation amount M. Add the third allocation amount, the fourth allocation amount, and the fifth allocation amount to obtain the first allocation amount. Among them, the third allocation amount represents the number of memory request identifiers occupied by other second type of requests in the remaining sub-request groups before the sub-request group where the current initial request identifier is located in the request group; the fourth allocation amount represents the number of memory request identifiers occupied by all first type of requests in the current request group; the fifth allocation amount represents the number of memory request identifiers occupied in the remaining request groups before the request group where the current initial request identifier is located.

[0101] Exemplarily, the expression of the third allocation amount is (thread_id_in_block / / N) * (N - t_remain_start_id_in_warp), the expression of the fourth allocation amount is t_remain_start_id_in_warp * (M / / N), and the expression of the fifth allocation amount is init_thread_id / / M * M.

[0102] In some embodiments, when the merging relationship is the first merging relationship and the merging attribute is to merge the requests corresponding to the initial request identifier, the sixth allocation amount is determined based on the request identifier in the block thread_id_in_block, the request processing threshold N, and the merging factor K; the seventh allocation amount is determined based on the request identifier in the block thread_id_in_block, the request processing threshold N, the merging factor K, and the request allocation amount M; the sixth allocation amount, the seventh allocation amount, and the fifth allocation amount are added together to obtain the first allocation amount. Among them, the sixth allocation amount represents the number of memory request identifiers occupied by other first-type requests in the bit order of the current merging item in the remaining sub-request groups before the sub-request group where the current initial request identifier is located in the request group; the seventh allocation amount represents the number of memory request identifiers occupied by other first-type requests in the bit order of the remaining merging items before the bit order of the current merging item where the current initial request identifier is located in all sub-request groups in the request group. Each sub-request group includes at least one merging item, and each merging item includes at least one request merged together.

[0103] Exemplarily, the expression of the sixth allocation amount is (thread_id_in_block / / N) * K, and the expression of the seventh allocation amount is ((thread_id_in_block % N) / / K) * (M / / N) * K.

[0104] In some embodiments, when the merging relationship is the second merging relationship and the merging attribute is to merge the requests corresponding to the initial request identifier, the sub-request group identifier warp_id is determined based on the request identifier in the block thread_id_in_block and the request processing threshold N; the eighth allocation amount is determined based on the sub-request group identifier warp_id and the merging factor K; the eighth allocation amount and the fifth allocation amount are added together to obtain the first allocation amount. Among them, the sub-request group identifier represents the sub-request group where the current initial request identifier is located; the eighth allocation amount represents the number of memory request identifiers occupied by the third-type requests in the remaining sub-request groups before the sub-request group where the current initial request identifier is located in the request group.

[0105] Exemplarily, the expression of the sub-request group identifier is thread_id_in_block / / N, and the expression of the eighth allocation amount is warp_id * K.

[0106] In some embodiments, when the merging relationship is the second merging relationship and the merging attribute is that the request corresponding to the initial request identifier is not merged, the request identifier thread_id_in_N within the sub-request group is determined based on the request identifier thread_id_in_block within the block and the request processing threshold N. Based on the request identifier thread_id_in_N within the sub-request group, the request allocation quantity M, and the request processing threshold N, the ninth allocation quantity is determined; the sub-request group identifier is used as the tenth allocation quantity; the ninth allocation quantity, the tenth allocation quantity, and the fifth allocation quantity are added together to obtain the first allocation quantity. Among them, the ninth allocation quantity represents the number of occupied memory request identifiers at other ordinal positions before the ordinal position of the current initial request identifier within the sub-request group among all sub-request groups within the request group. The tenth allocation quantity represents the number of occupied memory request identifiers at the ordinal position of the current initial request identifier within the remaining sub-request groups before the sub-request group where the current initial request identifier is located within the request group.

[0107] Exemplarily, the expression for the request identifier within the sub-request group is thread_id_in_block % N, and the expression for the ninth allocation quantity is M / / N * thread_id_in_N.

[0108] Step S304: Determine the second allocation quantity based on the merging attribute, the merging relationship, the request identifier within the block, the request processing threshold, and the request allocation quantity.

[0109] Among them, the second allocation quantity is used to represent the number of occupied memory request identifiers by other requests before the current initial request identifier within the merging item where the current initial request identifier is located. The merging item is the smallest merging unit, and there is at least one request in the merging item.

[0110] In some embodiments, when the merging relationship is the first merging relationship and the merging attribute is that the request corresponding to the initial request identifier is not merged, the second allocation quantity is determined based on the request identifier thread_id_in_block within the block, the starting unmerged request identifier t_remain_start_id_in_warp within the sub-request group, and the request processing threshold N. Exemplarily, the expression for the second allocation quantity is (thread_id_in_block % N) - t_remain_start_id_in_warp.

[0111] In some embodiments, when the merging relationship is the first merging relationship and the merging attribute is that the request corresponding to the initial request identifier is merged, the second allocation quantity is determined based on the request identifier thread_id_in_block within the block, the request processing threshold N, and the merging factor K. Exemplarily, the expression for the second allocation quantity is (thread_id_in_block % N) % K.

[0112] In some embodiments, when the merging relationship is the second merging relationship and the merging attribute is to merge the requests corresponding to the initial request identifier, the request identifier thread_id_in_N within the sub-request group is used as the second allocation quantity.

[0113] In some embodiments, when the merging relationship is the second merging relationship and the merging attribute is that the requests corresponding to the initial request identifier are not merged, the second allocation quantity is 0.

[0114] In the embodiments of the present application, based on the initial request identifier and the request allocation quantity, the request identifier within the block is determined; based on the merging relationship, the merging factor, the request processing threshold, and the request identifier within the block, the merging attribute corresponding to the initial request identifier is determined; based on the merging attribute, the merging relationship, the initial request identifier, the request identifier within the block, the request processing threshold, and the request allocation quantity, the first allocation quantity is determined, and based on the merging attribute, the merging relationship, the request identifier within the block, the request processing threshold, and the request allocation quantity, the second allocation quantity is determined. In this way, by considering the merging attribute of each request to determine the first allocation quantity and the second allocation quantity, a memory request identifier that more conforms to the actual test requirements can be obtained, and then an accurate target request address can be obtained, so as to facilitate the verification of subsequent test results.

[0115] Figure 4 It is a schematic implementation process of a test method provided by the embodiments of the present application Figure 4 , and this method can be executed by the processor of a computer device. Based on Figure 2 , the configuration parameter further includes: a boundary granularity, which is used to characterize the length between the boundary start address and the boundary end address; a request length, which is used to characterize the length of the address of the request. Figure 2 S203 in Figure 4 can be updated to S401 to S404, and will be described in conjunction with

[0116] Step S401, based on the first allocation quantity and the request length, determine the boundary base address.

[0117] Among them, the addresses accessed by the requests to be merged must be within the same merging boundary, and each merging boundary corresponds to an address range. The start address of the merging boundary is the boundary start address, the end address is the boundary end address, and the length of the merging boundary is the boundary granularity. The boundary granularity is used to characterize the length between the boundary start address and the boundary end address; the request length is used to characterize the length of the address of the request. In the embodiments of the present application, the request lengths of all requests are fixed values.

[0118] In some embodiments, the product of the first allocation quantity and the request length is used as the boundary base address. Based on the boundary base address, the boundary start address can be determined.

[0119] Step S402: Determine the boundary termination address based on the boundary base address and the boundary granularity.

[0120] In some embodiments, the expression of the boundary termination address is (coalesce_base_start_addr / / R)*R+R, where coalesce_base_start_addr is the boundary base address, R is the boundary granularity, and " / / " is the integer division operation. Considering that the memory coalescing module divides the memory addresses into multiple coalescing boundaries based on the boundary granularity, the boundary termination address of each coalescing boundary must be an integer multiple of the boundary granularity. (coalesce_base_start_addr / / R)*R can be regarded as the boundary start address.

[0121] Step S403: Determine the first request address based on the memory request identifier and the request length.

[0122] Among them, the first request address is used to represent the start address of the memory access request.

[0123] In some embodiments, multiply the memory request identifier by the request length to obtain the first request address. Exemplarily, for the initial request identifier 0, the corresponding memory request identifier is 0 and the request length is 4, so the first request address is 0x0; for the initial request identifier 1, the corresponding memory request identifier is 1 and the request length is 4, so the first request address is 0x4; for the initial request identifier 2, the corresponding memory request identifier is 2 and the request length is 4, so the first request address is 0x8. In this way, the memory addresses accessed by the requests to be coalesced are made continuous.

[0124] In some embodiments, when the coalescing relationship is the second coalescing relationship and the coalescing attribute is that the request corresponding to the initial request identifier has not been coalesced, multiply the memory request identifier by the boundary granularity to obtain the first request address, and directly use the first request address as the target request address.

[0125] Step S404: Determine the target request address based on the first request address, the request length, and the boundary termination address.

[0126] In some embodiments, the first request address and the request length are added together to obtain the termination address for accessing the memory, and the termination address is compared with the boundary termination address: in the case where the termination address is less than or equal to the boundary termination address, the first request address is directly used as the target request address; in the case where the termination address is greater than the boundary termination address, the difference obtained by subtracting the request length from the boundary termination address is used as the target request address.

[0127] Among them, if the directly calculated termination address is used as the target request address, it is easy to cause the termination address to exceed the boundary termination address of the merge boundary, resulting in the inability to perform normal merging of this request. Therefore, the termination address is compared with the boundary termination address. In the case where the termination address is greater than the boundary termination address, the difference obtained by subtracting the request length from the boundary termination address is used as the target request address, so that the target request address is within the merge boundary, ensuring that the access addresses of multiple requests merged together are all restricted within the same merge boundary.

[0128] In some embodiments, a thread issues at least two requests, each request corresponding to a loop identifier. Based on the loop identifier, the number of memory request modules, the request allocation amount, the request length, and the boundary granularity, the loop base address within the current loop is determined; the first request address and the request length are added together to obtain the termination address for accessing the memory; the termination address is compared with the boundary termination address to obtain the second request address; the second request address and the loop base address are added together to obtain the target request address.

[0129] In the embodiments of the present application, based on the first allocation amount and the request length, the boundary start address is determined; based on the boundary start address and the boundary granularity, the boundary termination address is determined; based on the memory request identifier and the request length, the first request address is determined; based on the first request address, the request length, and the boundary termination address, the target request address is determined. In this way, by first determining the boundary termination address corresponding to the request and adjusting the target request address based on the boundary termination address, an accurate target request address can be obtained, ensuring that the access addresses of the requests merged together are within the same merge boundary.

[0130] Figure 5 is the implementation process schematic of a test method provided by the embodiments of the present application Figure 5 , which can be executed by the processor of the computer device. Based on Figure 4 , the configuration parameter further includes a request allocation amount, the number of memory request modules, and a loop identifier. The request allocation amount is used to represent the number of requests allocated to the memory merge module. The memory merge module is used to execute the request merging process for at least two loops. The loop identifier is used to represent the current loop sequence corresponding to the request of the initial request identifier; Figure 4 in S404 can be updated to S501 to S503, which will be combined with Figure 5The steps shown will be described.

[0131] Step S501: Based on the loop identifier, the number of memory request modules, the request allocation amount, the request length, and the boundary granularity, determine the loop base address corresponding to the loop identifier.

[0132] In some embodiments, each thread issues at least two requests, each request corresponding to a loop order, and the test case includes all the requests issued by each thread. Among the at least two requests issued by the thread, the initial request identifier of the requests remains unchanged, but the target request addresses are different, and these two requests are distinguished based on the target request addresses. In each loop order, there is a loop base address, and the loop base address is used to determine the target request addresses occupied by the previous loop orders.

[0133] Among them, based on the loop identifier LOOP_id, the number of memory request modules W, the request allocation amount M, the request length L, and the boundary granularity R, determine the loop base address corresponding to the loop identifier. Exemplarily, the expression of the loop base address is LOOP_id * ceil((W * M * L) / R) * R, where ceil is the ceiling function.

[0134] Step S502: Based on the first request address, the request length, and the boundary termination address, determine the second request address.

[0135] Among them, the second request address is the starting address for the request to access memory determined without considering the loop order of the request.

[0136] In some embodiments, compare the sum of the first request address and the request length with the boundary termination address: when the sum of the first request address and the request length is greater than the boundary termination address, use the result of subtracting the request length from the boundary termination address as the second request address; when the sum of the first request address and the request length is less than or equal to the boundary termination address, use the first request address as the second request address.

[0137] Step S503: Based on the second request address and the loop base address, determine the target request address.

[0138] Among them, use the sum of the second request address and the loop base address as the target request address.

[0139] In the embodiments of the present application, based on the loop identifier, the number of memory request modules, the request allocation amount, the request length, and the boundary granularity, the loop base address corresponding to the loop identifier is determined; based on the first request address, the request length, and the boundary termination address, the second request address is determined; based on the second request address and the loop base address, the target request address is determined. In this way, by considering the loop base address, in the scenario where multiple requests are issued by a thread, multiple requests can be distinguished to determine the target request address of each request.

[0140] In some embodiments, the expected merge result includes at least one of the following: the expected number of request merges and the address constraint item; the expected number of request merges is used to represent the number of requests after normal merging of each request in the test case, and the address constraint item is used to represent the range of the target request addresses after normal merging of each request in the test case.

[0141] Among them, when generating the test case, based on the merge relationship, the merge behavior of each request in the test case can be determined in advance. Therefore, when each request in the test case is input into the memory merge module, the number of requests output by the memory merge module under normal circumstances is also determined in advance. Based on this, the expected number of request merges can be used as the expected merge result to test the memory merge module. The expected number of request merges is used to represent the number of requests after normal merging of each request in the test case.

[0142] Among them, after generating the test case, the target request address of each request in the test case is determined. Based on the merge relationship, the target request addresses of the requests output by the memory merge module under normal circumstances need to meet specific address constraints. Therefore, the address constraint item can be used as the expected merge result to test the memory merge module. The address constraint item is used to represent the range of the target request addresses after normal merging of each request in the test case.

[0143] In the embodiments of the present application, the expected merge result includes at least one of the following: the expected number of request merges and the address constraint item. In this way, it is possible to verify the number of requests and the request addresses output by the memory merge module.

[0144] In some embodiments, the configuration parameters include: a request processing threshold, which is used to represent the maximum number of requests processed by the memory merge module at a single time; a request allocation amount, which is used to represent the number of requests allocated to the memory merge module; the configuration parameters further include the number of memory request modules; the test case corresponds to a merge factor, which is used to represent the number of requests merged together; obtaining the expected merge result of the memory merge module includes: based on the number of memory request modules, the request allocation amount, the request processing threshold, the merge factor, and the merge relationship, determining the expected number of request merges.

[0145] Among them, based on the number of memory request modules W, the request allocation amount M, the request processing threshold N, the merging factor K, and the merging relationship, the expected number of request merges is determined. Exemplarily, when the merging relationship is the first merging relationship, the expression for the expected number of request merges is W*(M / N)*ceil(N / K), where ceil() is the ceiling function. When the merging relationship is the second merging relationship, the expression for the expected number of request merges is W*(M / N)*(N-K+1).

[0146] In some embodiments, the configuration parameter further includes a loop threshold. The loop threshold LOOP_CNT represents the maximum loop order corresponding to the request. At least two requests are issued by a single thread, and each request corresponds to a loop identifier. The value range of the loop identifier is from 0 to LOOP_CNT-1. Based on the loop threshold LOOP_CNT, the number of memory request modules W, the request allocation amount M, the request processing threshold N, the merging factor K, and the merging relationship, the expected number of request merges is determined. Exemplarily, when the merging relationship is the first merging relationship, the expression for the expected number of request merges is LOOP_CNT*W*(M / N)*ceil(N / K), where ceil() is the ceiling function. When the merging relationship is the second merging relationship, the expression for the expected number of request merges is LOOP_CNT*W*(M / N)*(N-K+1).

[0147] In the embodiments of the present application, since the method for obtaining the test cases provided in the foregoing embodiments is adopted, the corresponding expected number of request merges can quickly and accurately obtain a fixed expected number of request merges through the embodiments of the present application, which is convenient for testing the memory merging module; in addition, based on the number of memory request modules, the request allocation amount, the request processing threshold, the merging factor, and the merging relationship, the expected number of request merges is determined, and an accurate expected number of request merges can be obtained. Furthermore, the merging function of the memory merging module can be verified from the dimension of the number of merged requests, improving the testing ability.

[0148] In some embodiments, the actual merging result includes the actual number of request merges. Verifying the actual merging result based on the expected merging result to obtain the test result of the memory merging module includes: comparing the actual number of request merges with the expected number of request merges to obtain a comparison result; when the comparison result indicates that the actual number of request merges is the same as the expected number of request merges, the test result is a pass; when the comparison result indicates that the actual number of request merges is different from the expected number of request merges, the test result is a failure.

[0149] Among them, the actual number of requests to be merged is the number of requests output after each request in the test case is input into the memory merging module and the memory merging module performs request merging. The memory merging module determines the memory access addresses of the input requests, and requests within the same merging boundary and with consecutive access addresses can be merged.

[0150] Among them, the actual number of requests to be merged is compared with the expected number of requests to be merged to obtain a comparison result. When the comparison result indicates that the actual number of requests to be merged is the same as the expected number of requests to be merged, the test result is that the test passes; when the comparison result indicates that the actual number of requests to be merged is different from the expected number of requests to be merged, the test result is that the test fails.

[0151] In the embodiments of the present application, the actual number of requests to be merged is verified based on the expected number of requests to be merged to obtain the test result of the memory merging module. In this way, it can effectively verify whether the number of requests output by the memory merging module is correct.

[0152] In some embodiments, the configuration parameters include: the request allocation amount, which is used to represent the number of requests allocated to the memory merging module; the request length, which is used to represent the number of requests merged together; the boundary granularity, which is used to represent the length between the boundary start address and the boundary end address; the configuration parameters further include the number of memory request modules; the test case includes the target request address corresponding to the initial request identifier of each request; the address constraint items include a first constraint item, a second constraint item, and a third constraint item; the first constraint item is used to represent the relationship between the target request address of the output request and the request length, the second constraint item is used to represent the relationship between the target request address of the output request and the boundary granularity, and the third constraint item is used to represent the relationship between the target request address of the output request and the request allocation amount; obtaining the expected merging result of the memory merging module includes: determining the first constraint item based on the target request address and the request length; determining the second constraint item based on the target request address, the request length, and the boundary granularity; determining the third constraint item based on the target request address, the number of memory request modules, the request allocation amount, and the request length.

[0153] Among them, the expected merging result includes address constraint items, and the address constraint items include a first constraint item, a second constraint item, and a third constraint item. Different constraint items represent different constraint conditions corresponding to the target request address of the output request.

[0154] Among them, based on the target request address memory_addr and the request length L, the first constraint item is determined; exemplarily, the first constraint item is (memory_addr % L) == 0, indicating that the target request address is an integer multiple of the request length.

[0155] Among them, based on the target request address memory_addr, the request length L, and the boundary granularity R, the second constraint item is determined; exemplarily, the second constraint item is ((memory_addr + L) % R) == 0, indicating that the sum of the target request address and the request length is an integer multiple of the boundary granularity.

[0156] Among them, based on the target request address memory_addr, the number of memory request modules W, the request allocation amount M, and the request length L, the third constraint item is determined; exemplarily, the third constraint item is memory_addr < W * M * L, indicating that the target request address is less than the product of the number of memory merging modules, the request allocation amount, and the request length.

[0157] In some embodiments, considering the scenario where a single thread issues at least two requests, and each request corresponds to a loop identifier, at this time, the target request address of the input request includes the loop base address. Therefore, the address constraint item also needs to include the loop base address LOOP_base_memory_address. Each loop identifier corresponds to an address constraint item. Among them, the first constraint item is ((memory_addr - LOOP_base_memory_address) % L) == 0, the second constraint item is ((memory_addr - LOOP_base_memory_address + L) % R) == 0, and the third constraint item is memory_addr - LOOP_base_memory_address < W * M * L.

[0158] In the embodiments of the present application, based on the target request address and the request length, the first constraint item is determined; based on the target request address, the request length, and the boundary granularity, the second constraint item is determined; based on the target request address, the number of memory request modules, the request allocation amount, and the request length, the third constraint item is determined. In this way, accurate address constraint conditions for the target request address can be obtained.

[0159] In some embodiments, the actual merging result includes at least one target request address of an output request; verifying the actual merging result based on the expected merging result to obtain a test result of the memory merging module, including: matching each of the target request addresses with the first constraint item, the second constraint item, and the third constraint item respectively to obtain a matching result corresponding to each of the target request addresses; wherein, when the target request address satisfies at least one of the first constraint item and the second constraint item and the target request address satisfies the third constraint item, the matching result is a successful match; in other cases, the matching result is a failed match; when all the matching results are successful matches, the test result is a passed test; when at least one of the matching results is a failed match, the test result is a failed test.

[0160] Among them, each request in the test case is input into the memory merging module, the memory merging module merges the requests, outputs the target request address of the merged request, and matches the output target request address of each request with the address constraint item to obtain the matching result of each output request.

[0161] Among them, when the target request address satisfies at least one of the first constraint item and the second constraint item and the target request address satisfies the third constraint item, the matching result is a successful match; in other cases, the matching result is a failed match.

[0162] Among them, the test result of the memory merging module is determined based on the matching result of each output request. When all the matching results are successful matches, the test result is a passed test; when at least one of the matching results is a failed match, the test result is a failed test.

[0163] In the embodiments of the present application, each target request address is respectively matched with the first constraint item, the second constraint item, and the third constraint item to obtain a matching result corresponding to each target request address; when all the matching results are successful matches, the test result is a passed test; when at least one of the matching results is a failed match, the test result is a failed test. In this way, by verifying the target request address of the output request based on the first constraint item, the second constraint item, and the third constraint item to obtain the test result of the memory merging module, it can effectively verify whether the target request address output by the memory merging module is correct.

[0164] In some embodiments, a large number of random test cases are constructed based on a random merging relationship, and the memory merging module is tested based on the random test cases to cover scenarios not covered by typical test cases. The merging rule of requests within the request processing threshold in the random test cases is random. Therefore, the target request address of each request in the random test cases is random, and the setting of the expected merging result can be relaxed. Exemplarily, the expected number of request merges in the random test cases is set within a range.

[0165] In some embodiments, each test case corresponds to a merging factor, and the value range of the merging factor is a positive integer between 1 and N, where N is the request processing threshold. Different values can also be set for the request length L. Each test case corresponds to a request length, and the value range of the request length is a positive integer between 1 and L_Max, where L_Max is a preset maximum request length. Therefore, the number of test cases is N * L_Max.

[0166] In some embodiments, for a test case, the number of requests sent by a thread is controlled by the loop threshold LOOP_CNT. The number of threads is the product of the number of memory merging modules W and the request allocation amount M. Then, the total number of requests input by all memory merging modules is W * M * LOOP_CNT.

[0167] The following describes the application of the test method provided by the embodiments of the present application in an actual scenario.

[0168] In the design of modern parallel acceleration chips such as GPUs and NPUs, the access to their data paths (operations such as load, store, and atomic operations issued by GPU and NPU cores to access the memory subsystem) is generally processed in parallel. To improve the efficiency and throughput of data access, there is generally a memory merging module (MemoryCoalescing). For example, multiple threads of a GPU access the same address within 32 bytes. Assuming there are 4 threads, and the addresses of the 4 requests to access memory are 0x10000, 0x10004, 0x10008, and 0x1000c respectively, and the length of each request is 4 bytes, then these 4 requests can be merged into the same request with an address of 0x10000 and a length of 16 bytes. The number of requests is reduced, which can greatly improve the bandwidth utilization rate of the system.

[0169] For a task of software, it is generally dispatched to multiple parallel execution units (assuming there are W, corresponding to the number of memory merging modules in the above embodiments), and a parallel execution unit processes a certain number of subtasks (assuming a total of M threads of subtasks, corresponding to the request allocation amount in the above embodiments). Assume that there is exactly one memory merging module in a parallel execution unit, and a memory merging module can parallelly process subtasks of N threads (corresponding to the request processing threshold in the above embodiments) in the same cycle, that is, at most N threads' sub-requests of a memory merging module can be merged together in the same cycle. To process the subtasks of M threads, a parallel execution unit needs M / N times (assuming M is a positive integer multiple of N). Assume that all parallel execution units are used for the software task, that is, the task is divided into W*M threads for parallel execution. For each thread subtask, a thread number (id) needs to be given to it. Assume that the calculation method of the thread id is init_thread_id = block_id*M + thread_id_in_block, where init_thread_id (corresponding to the initial request identifier in the above embodiments) is a global thread number (0 <= init_thread_id < W*M), block_id (corresponding to the request group identifier in the above embodiments) is the serial number of the parallel execution unit where the current thread is located (0 <= block_id < W), and thread_id_in_block (corresponding to the in-block request identifier in the above embodiments) is the thread number within a parallel execution unit (0 <= thread_id_in_block < M). Corresponding to the unified computing device architecture CUDA, W corresponds to the number of blocks block_num, M corresponds to the number of threads thread_num, and N corresponds to the warp.

[0170] For the memory merging module, generally, a granularity R of a coalesce boundary is set (generally, the same configuration is globally used, corresponding to the boundary granularity in the above embodiments). For example, 32 bytes, indicating that requests whose addresses fall within the same 32 bytes can be merged.

[0171] For each memory request, its main parameters include the address (corresponding to the target request address in the above embodiments), the length (corresponding to the request length in the above embodiments), and the mask (only requests of types such as write / atomic operations need to consider); in a parallel acceleration chip, the lengths of multiple threads of parallel tasks (corresponding to the request lengths in the above embodiments) are generally equal and less than or equal to R (L <= R), that is, the sum of the address (corresponding to the target request address in the above embodiments) and the length of each thread crosses the boundary at most once. Hereinafter, this common request length is referred to as L. When performing request merging, generally, this request is first split (unrolled) according to the merging boundary into one or two sub-requests, and each sub-request is aligned with the merging boundary, that is, the sum of the address and the length of each sub-request is within the address range of the same merging boundary. For example, for a load request, address = 0x10010, length = 32 bytes, assuming the coalesce boundary = 32 bytes; then this load request needs to be split into 2 sub-requests. The address of sub-request 0 is 0x10010, length = 16 bytes, and the address of sub-request 1 is 0x10020, length = 16 bytes.

[0172] For a test case, generally, a loop threshold LOOP_CNT is set to control the number of requests issued by a thread. There are a total of W * M threads. Then, for this test case, the total number of input requests of all memory merging modules in the system is W * M * LOOP_CNT.

[0173] For the memory merging module, the general test method is to develop corresponding test cases according to the current W / M / N / R / L / LOOP_CNT configurations. Once one of the W / M / N / R / L / LOOP_CNT configurations changes, the test cases need to be re-developed. Therefore, the embodiments of this application propose a test method adaptable to any W / M / N / R / L / LOOP_CNT parameters to improve the iteration efficiency of testing.

[0174] Without considering the order among N threads and without considering the cross-boundary issues after address merging, the total number F(N) of address merging cases for N threads is similar to the integer splitting problem in number theory (where the number 1 represents that the address of 1 thread cannot be merged with other addresses, 2 represents that the addresses of 2 threads can be merged and cannot be merged with other addresses, and so on). Using the dynamic programming method, the number of different ways to split N into the sum of positive integers can be obtained. When N is small, all possible address merging scenarios can be constructed by traversal. However, if N is large, the value of F(N) is also very large, and it is not very realistic to construct so many test cases in a targeted manner. Therefore, it is necessary to make a choice for the test scenarios. For this purpose, the embodiments of this application provide two groups of typical address merging test cases, which are applicable to any W / M / N / R / L / LOOP_CNT parameters. Beyond the typical scenarios, random test cases are added to cover all combination scenarios.

[0175] The embodiments of this application give two groups of typical address merging test cases, which are respectively called test case pattern 1 and test case pattern 2.

[0176] For the test of the memory merging module, in the first step, the above two typical scenarios are tested; in the second step, any random test cases are constructed to cover all combination scenarios. Among them, in the first step, test cases are constructed in a targeted manner, and the merging behavior of the test cases can be predicted in advance, and a targeted development tool checker is made to verify the correctness of the module; in the second step, a large number of random test cases are constructed to cover the scenarios not covered by the typical scenarios. The merging behavior of the random test cases cannot be predicted in advance, so the checker can be relaxed. According to the above test process, the iterative efficiency of the test can be improved.

[0177] In the embodiments of this application, the rule for constructing the memory address for each loop is the same. Only in each loop, a loop base address of LOOP_id*ceil((W*M*L) / R)*R (hereinafter referred to as LOOP_memory_base_address) needs to be added, where LOOP_id is the serial number of this loop (corresponding to the loop identifier in the above embodiments), numbered from 0, 0 <= id < LOOP_CNT; ceil means rounding up. The construction principles of the following test case patterns are all for one loop.

[0178] Among them, the construction principle of test case pattern 1 (corresponding to the first merging relationship in the above embodiments) is that for each loop, for any test_case_id (corresponding to the merging factor in the above embodiments) (1 <= test_case_id <= N), every test_case_id sub-requests among the N sub-requests are merged, that is, as many sub-requests as possible are merged. Any one of the following test cases corresponds to multiple sub-cases, and the lengths (L) of the requests of different sub-cases are set differently, where 0 < L <= L_Max, and L_Max is a preset value representing the maximum length of the requests that the memory merging module can process. Therefore, this group of test cases has N * L_Max sub-cases.

[0179] Test case pattern 1 is shown in Table 3.

[0180] Table 3

[0181]

[0182]

[0183] To construct the final test case, first, a mapping method of thread IDs needs to be defined to map the initial task thread ID to the thread ID for memory access (init_thread_id --> memory_thread_id). init_thread_id corresponds to the initial request identifier in the above embodiments, and memory_thread_id corresponds to the memory request identifier in the above embodiments. The principle of the mapping relationship is that among the N threads processed in the same cycle, the memory_thread_ids of the threads that can be merged together are consecutive, and the memory_thread_ids of the threads that cannot be merged together are not consecutive. Without considering the merging boundary, the address (memory_addr, corresponding to the target request address in the above embodiments) where a certain thread ID requests to access memory is equal to memory_thread_id multiplied by L; however, for the above N * L_Max sub-cases, such calculation is likely to cause the memory_addr + request_length of a thread to cross the merging boundary, and with the superposition of the merging among multiple sub-requests, the situation is relatively complex and not conducive to the subsequent verification of the correctness of the test case. Therefore, if it is found that the memory_addr + L of a certain thread crosses the merging boundary of this merge, then adjust the memory_addr of this thread so that memory_addr + L is within the merging boundary of this merge; that is, ensure that the addresses + lengths of multiple threads that can be merged together are all limited within the same R range.

[0184] Given W / M / N / R / L / LOOP_CNT / test_case_id, the calculation formula for the number of output requests of a final test address merging module for a test case is LOOP_CNT * W * (M / N) * ceil(N / test_case_id), assuming that M is a positive integer multiple of N, and ceil represents rounding up. The total number is used as a verification point for verifying the correctness of the test case. In addition to the total number of output requests (corresponding to the expected request merge number in the above embodiment), a verification point can also be added to verify whether the addresses of the output requests satisfy such a constraint (corresponding to the address constraint item in the above embodiment): (((memory_addr % L) == 0) || (((memory_addr + L) % R) == 0)) && (memory_addr < W * M * L), where % is the modulo operation; it should be noted that here it is assumed that LOOP_id == 0. If LOOP_id is not 0, memory_addr in the constraint condition needs to be replaced with memory_addr – the current LOOP_memory_base_address.

[0185] Given W / M / N / R / L / LOOP_CNT / test_case_id, the pseudo-code for the process of obtaining memory_thread_id and memory_addr from init_thread_id (0 <= init_thread_id < W * M) is as follows. Here, / / represents integer division.

[0186]

[0187]

[0188] In the above test case pattern 1, the pseudo-code for obtaining memory_thread_id and memory_addr from init_thread_id can be converted into the following process:

[0189] Step 1: Determine the starting unmerged request identifier (t_remain_start_id_in_warp), request group identifier (block_id), and in-block request identifier (thread_id_in_block) within the sub-request group based on the configuration parameters (request allocation amount M, request processing threshold N, merge factor test_case_id) and the initial request identifier (init_thread_id).

[0190] Among them, each memory merging module corresponds to a request group, each request group includes at least one sub-request group, the request group includes M requests with a request allocation amount, and the sub-request group includes N requests with a request processing threshold.

[0191] Among them, the request processing threshold N divides the merging factor test_case_id, and then multiplies by the merging factor test_case_id to obtain the starting unmerged request identifier t_remain_start_id_in_warp within the sub-request group; the initial request identifier init_thread_id divides the request allocation amount M to obtain the request group identifier block_id; the initial request identifier init_thread_id takes the remainder of the request allocation amount M to obtain the in-block request identifier thread_id_in_block.

[0192] Step 2, for each loop, based on the loop identifier LOOP_id, the number of memory request modules W, the request allocation amount M, the request length L, and the boundary granularity R, determine the loop base address (LOOP_memory_base_address) of the current loop.

[0193] Among them, round up (W * M * L) / R, then multiply by R, and then multiply by LOOP_id to obtain LOOP_memory_base_address.

[0194] In the case where there are unmerged requests in the current test case (N!= t_remain_start_id_in_warp) and the current request is an unmerged request ((thread_id_in_block % N) >= t_remain_start_id_in_warp), execute Steps 4 to 7; in other cases, execute Steps 8 to 11 (Step 3 corresponds to the determination process of the merging attribute in the above embodiment).

[0195] Exemplarily, N is 9 and test_case_id is 2, indicating that among 9 requests, every two requests are merged together. Therefore, among the 9 requests, the first 8 requests are merged requests and the 9th request is an unmerged request.

[0196] Step 4, based on the request allocation amount M, the request processing threshold N, the starting unmerged request identifier within the sub-request group (t_remain_start_id_in_warp), the request group identifier (block_id), and the in-block request identifier (thread_id_in_block), determine the first allocation amount (t4 + t_remain_start_id_in_warp * (M / / N) + block_id * M) and the second allocation amount (t5).

[0197] Among them, the request identifier thread_id_in_block within the block is divided evenly by the request processing threshold N, and then multiplied by the difference between the request processing threshold N and the starting unmerged request identifier t_remain_start_id_in_warp within the sub-request group to obtain t4; the request identifier thread_id_in_block within the block is taken modulo the request processing threshold N, and then subtracted by the starting unmerged request identifier t_remain_start_id_in_warp within the sub-request group to obtain t5; the request allocation amount M is divided evenly by the request processing threshold N, and then multiplied by the starting unmerged request identifier t_remain_start_id_in_warp within the sub-request group to obtain t_remain_start_id_in_warp*(M / / N); the request group identifier block_id is multiplied by the request allocation amount M to obtain block_id*M.

[0198] Step 5: Sum the first allocation amount and the second allocation amount to obtain the memory request identifier (memory_thread_id).

[0199] Step 6: Based on the first allocation amount and the configuration parameters (boundary granularity R, request length L), determine the current boundary base address (coalesce_base_start_addr) and the boundary end address (coalesce_boundary_end_addr).

[0200] Among them, multiplying the first allocation amount by the request length L gives the boundary base address coalesce_base_start_addr; dividing the boundary base address coalesce_base_start_addr by the boundary granularity R, multiplying the result by the boundary granularity R, and then adding the boundary granularity R gives the boundary end address coalesce_boundary_end_addr.

[0201] Step 7: Based on the memory request identifier and the request length, determine the first request address (memory_addr). Go to Step 12.

[0202] Among them, multiplying the memory request identifier memory_thread_id by the request length L gives the first request address memory_addr.

[0203] Step 8: Based on the request identifier thread_id_in_block within the block, the request allocation amount M, the request processing threshold N, the request group identifier (block_id), and the merge factor test_case_id, determine the first allocation amount (t1 + t2 + block_id*M) and the second allocation amount (t3).

[0204] Among them, the sub-request group includes at least one merge item, and the merge item includes at least one request merged together.

[0205] Exemplarily, N is 9 and test_case_id is 2, indicating that in the sub-request group, every two requests are merged together. This sub-request group includes 5 merge items. Each of the first 4 merge items includes 2 requests, and the last merge item includes 1 request.

[0206] Step 9: Sum the first allocation quantity and the second allocation quantity to obtain the memory request identifier (memory_thread_id).

[0207] Step 10: Determine the current boundary start address (coalesce_base_start_addr) and the boundary end address (coalesce_boundary_end_addr) based on the first allocation quantity and the configuration parameters (boundary granularity R, request length L).

[0208] Among them, multiplying the first allocation quantity by the request length L gives the boundary start address coalesce_base_start_addr; dividing the boundary start address coalesce_base_start_addr by the boundary granularity R, then multiplying by the boundary granularity R, and then adding the boundary granularity R gives the boundary end address coalesce_boundary_end_addr.

[0209] Step 11: Determine the first request address (memory_addr) based on the memory request identifier and the request length. Go to Step 12.

[0210] Among them, multiplying the memory request identifier memory_thread_id by the request length L gives the first request address memory_addr.

[0211] Step 12: In the case where the sum of the first request address and the request length (L) is greater than the boundary end address, use the difference between the boundary end address and the request length as the second request address. In other cases, use the first request address as the second request address.

[0212] Step 13: Sum the second request address and the loop base address of the current loop to obtain the target request address.

[0213] In some embodiments, when LOOP_id = 0, assuming W = 1, M = 1024, N = 8, R = 32, and L_Max = 16, this group of test cases has a total of N * L_Max = 8 * 16 = 128 sub-test cases. For test_case_id = 2 and L = 4, the correspondence between the initial request identifier / memory request identifier / target request address is shown in Table 4. It can be found that in this scenario, the target request address is equal to the memory request identifier * L.

[0214] Table 4

[0215] Initial request identifier Memory request identifier Target request address 0 0 0x0 1 1 0x4 2 256 0x400 3 257 0x404 4 512 0x800 5 513 0x804 6 768 0xc00 7 769 0xc04 8 2 0x8 9 3 0xc 10 258 0x408 11 259 0x40c 12 514 0x808 13 515 0x80c 14 770 0xc08 15 771 0xc0c …

[0216] In some embodiments, for test_case_id = 3 and L = 15, the correspondence between the initial request identifier / memory request identifier / target request address is shown in Table 5. For the row where the initial request identifier is 2, the target request address calculated by multiplying the memory request identifier by L is 0x1e. Since 0x1e + 15 exceeds the combined boundary common to the three threads with initial request identifiers 0, 1, and 2, the final target request address is adjusted so that the three threads with initial request identifiers 0, 1, and 2 can be combined within the same R range. For test_case_id = 3 and N = 8, the initial request identifiers 0, 1, 2 (3, 4, 5; 6, 7; 8, 9, 10; 11, 12, 13; 14, 15) are combined together. It can be found that the target request addresses of 12 and 13 are equal after adjustment, which is also expected.

[0217] Table 5:

[0218] Initial request identifier Memory request identifier Memory request identifier *L Target request address 0 0 0x0 0x0 1 1 0xf 0xf 2 2 0x1e 0x11 3 384 0x1680 0x1680 4 385 0x168f 0x168f 5 386 0x169e 0x1691 6 768 0x2d00 0x2d00 7 769 0x2d0f 0x2d0f 8 3 0x2d 0x2d 9 4 0x3c 0x31 10 5 0x4b 0x31 11 387 0x16ad 0x16ad 12 388 0x16bc 0x16b1 13 389 0x16cb 0x16b1 14 770 0x2d1e 0x2d11 15 771 0x2d2d 0x2d11 …

[0219] Among them, for test cases 1 to N of the above test case pattern 1, as many sub-requests as possible among the N sub-requests are combined; similarly, there is exactly one group of test_case_id sub-requests that can be combined among the N sub-requests, which is the test case pattern 2. Any of the following test cases corresponds to multiple sub-test cases, and the lengths (L) of the requests of different sub-test cases are set differently, where 0 < L <= L_Max. Therefore, this group of test cases has a total of N * L_Max sub-test cases, but the scenarios where test_case_id = 1 or N are exactly the same as those of test_case_id = 1 or N in test case pattern 1.

[0220] Among them, similar to test case pattern 1, if it is found that the memory request address + L of a certain thread crosses the merge boundary of the current merge, then adjust the target request address of this thread to ensure that the addresses + lengths of multiple threads that can be merged together are all limited within the same R range. For the threads in the N sub-requests that cannot be merged with all other sub-requests, make their target request address = memory request identifier * R, so that it can be ensured that this sub-request cannot be merged with all other sub-requests.

[0221] In some embodiments, given W / M / N / R / L / test_case_id, the calculation formula for the number of output requests finally tested by the test address merging module for a test case of test case pattern 2 (corresponding to the second merge relationship in the above embodiments) is W * (M / N) * (N - test_case_id + 1), assuming that M is a positive integer multiple of N. The total number is used as a verification point for verifying the correctness of the test case. In addition to the total number of output requests, a verification point can also be added to verify whether the addresses of the output requests satisfy the following constraints: (((memory_addr % L) == 0) || (((memory_addr + L) % R) == 0)) && (memory_addr < W * M * L), where % is the modulo operation; it should be noted that here it is assumed that the constraint of LOOP_id == 0. If LOOP_id is not 0, memory_addr in the constraint condition needs to be replaced with memory_addr - the base address of the current loop.

[0222] Among them, test case pattern 2 is shown in Table 6.

[0223] Table 6

[0224]

[0225]

[0226] Given W / M / N / R / L / LOOP_CNT / test_case_id, the pseudocode for the process of obtaining memory_thread_id and memory_addr from init_thread_id (0 <= init_thread_id < W * M) is as follows. Among them, / / represents integer division.

[0227]

[0228] In the above test case pattern 2, the pseudocode for obtaining memory_thread_id and memory_addr from init_thread_id can be converted into the following process:

[0229] Step 1: Determine the request group identifier (block_id) and the request identifier within the block (thread_id_in_block) based on the configuration parameter (request allocation amount M) and the initial request identifier (init_thread_id).

[0230] Among them, block_id is obtained by dividing the initial request identifier init_thread_id by the request allocation amount M; thread_id_in_block is obtained by taking the remainder of the initial request identifier init_thread_id divided by the request allocation amount M.

[0231] Step 2: Determine the request identifier within the sub-request group (thread_id_in_N) and the sub-request group identifier (warp_id) based on the request identifier within the block thread_id_in_block and the request processing threshold N.

[0232] Among them, thread_id_in_N is obtained by taking the remainder of the request identifier within the block thread_id_in_block divided by the request processing threshold N; warp_id is obtained by dividing the request identifier within the block thread_id_in_block by N.

[0233] Step 3: For each loop (LOOP_id), determine the loop base address of the current loop (LOOP_memory_base_address) based on the loop identifier LOOP_id, the number of memory request modules W, the request allocation amount M, the request length L, and the boundary granularity R.

[0234] Among them, round up ((W * M * L) / R), then multiply by R, and then multiply by LOOP_id to obtain LOOP_memory_base_address.

[0235] Step 4: When the request identifier within the sub-request group is less than the merge factor (thread_id_in_N < test_case_id), execute Steps 5 to 10; in other cases, execute Steps 11 to 14. (Step 4 corresponds to the determination process of the merge attribute in the above embodiment)

[0236] Step 5: Determine the first allocation amount (warp_id * test_case_id + block_id * M) and the second allocation amount (thread_id_in_N) based on the request allocation amount M, the request identifier within the sub-request group thread_id_in_N, the sub-request group identifier warp_id, the merge factor test_case_id, and the request group identifier block_id.

[0237] Among them, the warp_id * test_case_id is obtained by multiplying the sub-request group identifier warp_id by the merging factor test_case_id; the block_id * M is obtained by multiplying the block_id by M.

[0238] Step 6: Sum the first allocation quantity and the second allocation quantity to obtain the memory request identifier (memory_thread_id).

[0239] Step 7: Determine the current boundary base address (coalesce_base_start_addr) and the boundary end address (coalesce_boundary_end_addr) based on the first allocation quantity and the configuration parameters (boundary granularity R, request length L).

[0240] Among them, the boundary base address coalesce_base_start_addr is obtained by multiplying the first allocation quantity by the request length L; the boundary end address coalesce_boundary_end_addr is obtained by dividing the boundary base address coalesce_base_start_addr by the boundary granularity R, multiplying the result by the boundary granularity R, and then adding the boundary granularity R.

[0241] Step 8: Determine the first request address (memory_addr) based on the memory request identifier and the request length.

[0242] Among them, the memory_addr is obtained by multiplying the memory request identifier memory_thread_id by the request length L.

[0243] Step 9: In the case where the sum of the first request address and the request length (L) is greater than the boundary end address, use the difference between the boundary end address and the request length as the second request address. In other cases, use the first request address as the second request address.

[0244] Step 10: Sum the second request address and the loop base address of the current loop to obtain the target request address.

[0245] Step 11: Determine the first allocation quantity (M / / N * thread_id_in_N + warp_id + block_id * M) and the second allocation quantity (0) based on the request allocation quantity M, the request processing threshold N, the request identifier thread_id_in_N within the sub-request group, the sub-request group identifier warp_id, and the request group identifier block_id.

[0246] Among them, by dividing the requested allocation amount M by the request processing threshold N and then multiplying by the request identifier thread_id_in_N within the sub-request group, M / / N*thread_id_in_N is obtained.

[0247] Step 12: Sum the first allocation amount and the second allocation amount to obtain the memory request identifier (memory_thread_id).

[0248] Step 13: Determine the intermediate request address (memory_addr) based on the memory request identifier and the boundary granularity.

[0249] Among them, by multiplying the memory request identifier memory_thread_id by the boundary granularity R, the intermediate request address memory_addr is obtained.

[0250] Step 14: Sum the intermediate request address and the memory base address of the current loop to obtain the target request address.

[0251] In some embodiments, LOOP_id = 0. Also assume that W = 1, M = 1024, N = 8, R = 32, and L_Max = 16. For test_case_id = 5, L = 8, and the correspondence between the initial request identifier / memory request identifier / target request address is shown in Table 7.

[0252] Table 7

[0253] Initial request identifier Memory request identifier Memory request identifier *L Target request address 0 0 0x0 0x0 1 1 0x8 0x8 2 2 0x10 0x10 3 3 0x18 0x18 4 4 0x20 0x18 5 640 0x5000 0x5000 6 768 0x6000 0x6000 7 896 0x7000 0x7000 8 5 0x28 0x28 9 6 0x30 0x30 10 7 0x38 0x38 11 8 0x40 0x38 12 9 0x48 0x38 13 641 0x5020 0x5020 14 769 0x6020 0x6020 15 897 0x7020 0x7020 …

[0254] In the embodiments of the present application: 1. Define a mapping method for thread IDs, mapping the initial task thread ID to the thread ID for memory access (init_thread_id -> memory_thread_id). The principle of the mapping relationship is that among the N threads processed in the same cycle, the memory_thread_ids of the threads that can be merged together are consecutive, and the memory_thread_ids of the threads that cannot be merged together are not consecutive. 2. Provide two sets of typical address merging test cases, applicable to any W / M / N / R / L parameters. Among them, test case pattern 1 is to merge as many sub-requests as possible among the N sub-requests (merge every test_case_id sub-requests among the N sub-requests); test case pattern 2 is that there is and only one set of test_case_id sub-requests that can be merged among the N sub-requests. 3. If it is found that the memory_addr + L of a certain thread crosses the merge boundary of the current merge, adjust the memory_addr of this thread so that memory_addr + L is within the merge boundary of the current merge; that is, ensure that the addresses + lengths of multiple threads that can be merged together are all limited within the same R range. 4. For the two sets of test cases, given W / M / N / R / L / LOOP_CNT / test_case_id, the pseudo-code for the process of obtaining memory_thread_id and memory_addr from init_thread_id (0 <= init_thread_id < W * M). 5. For the test of the memory merging module, first test the above two typical scenarios, and then overlay random cases to cover all combined scenarios, improving the iteration efficiency of the test.

[0255] In the embodiments of the present application, compared with the related art, a test method adaptable to any W / M / N / R / L / LOOP_CNT parameters is proposed (first test the above two typical scenarios, and then overlay random cases to cover all combined scenarios), improving the iteration efficiency of the test. Two sets of typical address merging test cases are provided, applicable to any W / M / N / R / L / LOOP_CNT parameters.

[0256] Based on the foregoing embodiments, an embodiment of the present application provides a test device, which includes each unit included and each module included in each unit, and can be implemented by a processor in a computer device; of course, it can also be implemented by specific logic circuits; during implementation, the processor can be a central processing unit (CPU), a microprocessor unit (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), etc.

[0257] Figure 6 It is a schematic diagram of the composition structure of a test device provided by an embodiment of the present application. As Figure 6 shown, the test device 600 includes: an acquisition module 610, an input module 620, and a verification module 630, where:

[0258] The acquisition module 610 is configured to acquire test cases and expected merge results of the memory merge module; each request in the test cases and the expected merge results are determined based on the merge relationship between a preset number of requests; the merge relationship is used to characterize whether adjacent requests among the preset number of requests are merged.

[0259] The input module 620 is configured to input each request in the test cases into the memory merge module to obtain an actual merge result.

[0260] The verification module 630 is configured to verify the actual merge result based on the expected merge result to obtain the test result of the memory merge module.

[0261] In some embodiments, the merge relationship between the preset number of requests is used to characterize that the preset number of requests are divided into P sets based on a merge factor K, and requests within one set correspond to one merge result output by the memory merge module; the merge relationship includes at least one of the following: a first merge relationship, which is used to characterize the existence of P - 1 first sets and 1 second set; a second merge relationship, which is used to characterize the existence of 1 first set and N - K third sets; where, the first set includes K requests, the second set includes O requests, 1 ≤ O < K, the third set includes 1 request, and N is the preset number.

[0262] In some embodiments, the test case includes a target request address corresponding to the initial request identifier of each request. The obtaining module 610 is further configured to determine a first allocation amount and a second allocation amount based on the merging relationship, configuration parameters, and the initial request identifier; determine a memory request identifier based on the first allocation amount and the second allocation amount; and determine the target request address corresponding to the initial request identifier based on the configuration parameters, the first allocation amount, and the memory request identifier.

[0263] In some embodiments, the configuration parameters include: a request processing threshold, which is used to represent the maximum number of requests processed by the memory merging module at a single time; a request allocation amount, which is used to represent the number of requests allocated to the memory merging module; a merging factor, which is used to represent the maximum number of requests in a set. The obtaining module 610 is further configured to determine a request identifier within a block based on the initial request identifier and the request allocation amount. The request identifier within the block is used to represent the bit order of the request corresponding to the initial request identifier in the corresponding request group. Each request group corresponds to a memory merging module, and the request group includes all requests allocated to the memory merging module. Determine the merging attribute corresponding to the initial request identifier based on the merging relationship, the merging factor, the request processing threshold, and the request identifier within the block. The merging attribute is used to determine whether the request corresponding to the initial request identifier is merged. Determine the first allocation amount based on the merging attribute, the merging relationship, the initial request identifier, the request identifier within the block, the request processing threshold, and the request allocation amount. Determine the second allocation amount based on the merging attribute, the merging relationship, the request identifier within the block, the request processing threshold, and the request allocation amount.

[0264] In some embodiments, the obtaining module 610 is further configured to determine the second allocation amount based on the merging attribute, the merging relationship, the request identifier within the block, the request processing threshold, and the request allocation amount; and determine the memory request identifier based on the first allocation amount and the second allocation amount.

[0265] In some embodiments, the configuration parameters further include: a boundary granularity, which is used to represent the length between the boundary start address and the boundary end address; a request length, which is used to represent the length of the address of the request. The obtaining module 610 is further configured to determine the boundary base address based on the first allocation amount and the request length; determine the boundary end address based on the boundary base address and the boundary granularity; determine a first request address based on the memory request identifier and the request length; and determine the target request address based on the first request address, the request length, and the boundary end address.

[0266] In some embodiments, the configuration parameters further include a request allocation quantity, the number of memory request modules, and a loop identifier. The request allocation quantity is used to represent the number of requests allocated to the memory merging module. The memory merging module is used to execute a request merging process for at least two loops. The loop identifier is used to represent the current loop sequence corresponding to the request with the initial request identifier. The obtaining module 610 is further configured to determine a loop base address corresponding to the loop identifier based on the loop identifier, the number of memory request modules, the request allocation quantity, the request length, and the boundary granularity; determine a second request address based on the first request address, the request length, and the boundary termination address; and determine the target request address based on the second request address and the loop base address.

[0267] In some embodiments, the expected merging result includes at least one of the following: an expected request merging number and an address constraint item. The expected request merging number is used to represent the number of requests after normal merging of each request in the test case. The address constraint item is used to represent the range of the target request addresses after normal merging of each request in the test case.

[0268] In some embodiments, the configuration parameters include: a request processing threshold, which is used to represent the maximum number of requests processed by the memory merging module at a single time; a request allocation quantity, which is used to represent the number of requests allocated to the memory merging module; the configuration parameters further include the number of memory request modules; the test case corresponds to a merging factor, and the merging factor is used to represent the number of requests merged together. The obtaining module 610 is further configured to determine the expected request merging number based on the number of memory request modules, the request allocation quantity, the request processing threshold, the merging factor, and the merging relationship.

[0269] In some embodiments, the actual merging result includes an actual request merging number. The verification module 630 is further configured to compare the actual request merging number with the expected request merging number to obtain a comparison result. When the comparison result indicates that the actual request merging number is the same as the expected request merging number, the test result is that the test passes. When the comparison result indicates that the actual request merging number is different from the expected request merging number, the test result is that the test fails.

[0270] In some embodiments, the configuration parameters include: the request allocation quantity, which is used to characterize the number of requests allocated to the memory merging module; the request length, which is used to characterize the number of requests merged together; the boundary granularity, which is used to characterize the length between the boundary start address and the boundary end address; the configuration parameters further include the number of memory request modules; the test case includes the target request address corresponding to the initial request identifier of each request; the address constraint items include a first constraint item, a second constraint item, and a third constraint item; the first constraint item is used to characterize the relationship between the target request address of the output request and the request length, the second constraint item is used to characterize the relationship between the target request address of the output request and the boundary granularity, and the third constraint item is used to characterize the relationship between the target request address of the output request and the request allocation quantity; the obtaining module 610 is further configured to determine the first constraint item based on the target request address and the request length; determine the second constraint item based on the target request address, the request length, and the boundary granularity; and determine the third constraint item based on the target request address, the number of memory request modules, the request allocation quantity, and the request length.

[0271] In some embodiments, the actual merging result includes the target request addresses of at least one output request; the verification module 630 is further configured to match each of the target request addresses with the first constraint item, the second constraint item, and the third constraint item respectively to obtain a matching result corresponding to each of the target request addresses; wherein, when the target request address satisfies at least one of the first constraint item and the second constraint item and the target request address satisfies the third constraint item, the matching result is a successful match; in other cases, the matching result is a failed match; when all the matching results are successful matches, the test result is a pass; when at least one of the matching results is a failed match, the test result is a fail.

[0272] The description of the above device embodiments is similar to the description of the above method embodiments and has similar beneficial effects to the method embodiments. In some embodiments, the functions or modules included in the device provided in the embodiments of the present application can be used to execute the methods described in the above method embodiments. For the technical details not disclosed in the device embodiments of the present application, please refer to the description of the method embodiments of the present application for understanding.

[0273] It should be noted that in the embodiments of the present application, if the above-mentioned test method is implemented in the form of software function modules and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence or the part that contributes to the related technology, can be embodied in the form of a software product. The software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (ROMs), magnetic disks, or optical discs. In this way, the embodiments of the present application are not limited to any specific hardware, software, or firmware, or any combination of hardware, software, and firmware.

[0274] The embodiments of the present application provide a computer device, including a memory and a processor. The memory stores a computer program that can run on the processor, and when the processor executes the program, it implements some or all of the steps in the above method.

[0275] The embodiments of the present application provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements some or all of the steps in the above method. The computer-readable storage medium can be transient or non-transient.

[0276] The embodiments of the present application provide a computer program, including computer-readable code. When the computer-readable code runs in a computer device, the processor in the computer device executes to implement some or all of the steps in the above method.

[0277] The embodiments of the present application provide a computer program product. The computer program product includes a non-transient computer-readable storage medium storing a computer program. When the computer program is read and executed by a computer, it implements some or all of the steps in the above method. The computer program product can be specifically implemented in a manner of hardware, software, or a combination thereof. In some embodiments, the computer program product is specifically embodied as a computer storage medium. In other embodiments, the computer program product is specifically embodied as a software product, such as a software development kit (SDK), etc.

[0278] It should be noted here that: the descriptions of the above embodiments tend to emphasize the differences between the embodiments, and their similarities can be referred to each other. The descriptions of the above embodiments of the device, storage medium, computer program and computer program product are similar to the descriptions of the above method embodiments, and have beneficial effects similar to those of the method embodiments. For the technical details not disclosed in the embodiments of the device, storage medium, computer program and computer program product of the present application, please refer to the descriptions of the method embodiments of the present application for understanding.

[0279] Figure 7 The following is a schematic diagram of the hardware entity of a computer device provided by an embodiment of the present application. As Figure 7 shown, the hardware entity of the computer device 700 includes: a processor 701 and a memory 702. Among them, the memory 702 stores a computer program that can run on the processor 701, and when the processor 701 executes the program, it implements the steps in the method of any of the above embodiments.

[0280] The memory 702 stores a computer program that can run on the processor. The memory 702 is configured to store instructions and applications executable by the processor 701, and can also cache data to be processed or already processed by the processor 701 and each module in the computer device 700 (for example, image data, audio data, voice communication data, and video communication data), and can be implemented by flash memory (FLASH) or random access memory (Random Access Memory, RAM).

[0281] When the processor 701 executes the program, it implements the steps of the test method of any of the above. The processor 701 generally controls the overall operation of the computer device 700.

[0282] An embodiment of the present application provides a computer storage medium. The computer storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of the test method of any of the above embodiments.

[0283] It should be noted here that: the descriptions of the above storage medium and device embodiments are similar to the descriptions of the above method embodiments, and have beneficial effects similar to those of the method embodiments. For the technical details not disclosed in the embodiments of the storage medium and device of the present application, please refer to the descriptions of the method embodiments of the present application for understanding.

[0284] The above-mentioned processor may be at least one of an Application Specific Integrated Circuit (ASIC), a Digital Signal Processor (DSP), a Digital Signal Processing Device (DSPD), a Programmable Logic Device (PLD), a Field Programmable Gate Array (FPGA), a Central Processing Unit (CPU), a controller, a microcontroller, and a microprocessor. It can be understood that the electronic device implementing the functions of the above-mentioned processor may also be other devices, and the embodiments of the present application do not make specific limitations.

[0285] The above-mentioned computer storage medium / memory may be a Read Only Memory (ROM), a Programmable Read-Only Memory (PROM), an Erasable Programmable Read-Only Memory (EPROM), an Electrically Erasable Programmable Read-Only Memory (EEPROM), a Ferromagnetic Random Access Memory (FRAM), a Flash Memory, a magnetic surface memory, an optical disc, or a Compact Disc Read-Only Memory (CD-ROM), etc.; it may also be various terminals including one or any combination of the above-mentioned memories, such as a mobile phone, a computer, a tablet device, a personal digital assistant, etc.

[0286] As described above, it is only the implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application.

Claims

1. A testing method, characterized in that: The method comprises: Obtaining a test case and an expected merge result of a memory merge module; each request in the test case and the expected merge result are determined based on a merge relationship between a preset number of requests; the merge relationship is used to characterize whether adjacent requests in the preset number of requests are merged; Input each request in the test case into the memory merging module to obtain an actual merging result; The actual merging result is verified based on the expected merging result to obtain a test result of the memory merging module.

2. The method according to claim 1, characterized in that The merging relationship between the preset number of requests is used to characterize that the preset number of requests are divided into P sets based on a merging factor K, and the requests in one set correspond to a merging result output by a memory merging module; the merging relationship includes at least one of the following: a first merging relationship, used to characterize the existence of P-1 first sets and 1 second set; a second merging relationship, used to characterize the existence of 1 first set and NK third sets; wherein the first set includes K requests, the second set includes O requests, 1≤O<K, the third set includes 1 request, and N is the preset number.

3. The method according to claim 1, characterized in that: The test case includes a target request address corresponding to an initial request identifier of each request, and the test case for obtaining the memory merging module includes: Determining a first allocation amount and a second allocation amount based on the merging relationship, the configuration parameters and the initial request identifier; Determining a memory request identifier based on the first allocation amount and the second allocation amount; Based on the configuration parameters, the first allocation amount, and the memory request identifier, a target request address corresponding to the initial request identifier is determined.

4. The method according to claim 3, characterized in that The configuration parameters include: a request processing threshold, which is used to characterize the maximum number of requests that the memory merging module can process at a time; a request allocation amount, which is used to characterize the number of requests that the memory merging module is allocated; a merging factor, which is used to characterize the maximum number of requests in a set; and determining the first allocation amount and the second allocation amount based on the merging relationship, the configuration parameters and the initial request identifier, including: Based on the initial request identifier and the request allocation amount, determine the intra-block request identifier; the intra-block request identifier is used to represent the bit sequence of the request of the initial request identifier in the corresponding request group; wherein each request group corresponds to a memory merging module, and the request group includes all requests allocated to the memory merging module; Based on the merging relationship, the merging factor, the request processing threshold and the intra-block request identifier, determining a merging attribute corresponding to the initial request identifier; the merging attribute is used to determine whether the request corresponding to the initial request identifier is to be merged; Determine the first allocation amount based on the merging attribute, the merging relationship, the initial request identifier, the intra-block request identifier, the request processing threshold, and the request allocation amount; A second allocation amount is determined based on the merging attribute, the merging relationship, the intra-block request identifier, the request processing threshold, and the request allocation amount.

5. The method according to claim 3, characterized in that: The configuration parameters also include: a boundary granularity, which is used to characterize the length between a boundary start address and a boundary end address; a request length, which is used to characterize the length of the requested address; and determining the target request address corresponding to the initial request identifier based on the configuration parameters, the first allocation amount, and the memory request identifier, includes: determining the boundary base address based on the first allocation and the request length; Determining the boundary termination address based on the boundary base address and the boundary granularity; Determine a first request address based on the memory request identifier and the request length; The target request address is determined based on the first request address, the request length, and the boundary termination address.

6. The method according to claim 5, characterized in that The configuration parameters also include a request allocation amount, a number of memory request modules, and a loop identifier, wherein the request allocation amount is used to represent the number of requests allocated to the memory merging module, the memory merging module is used to perform at least two loop request merging processes, and the loop identifier is used to represent the current loop order corresponding to the request of the initial request identifier; The determining the target request address based on the first request address, the request length, and the boundary termination address includes: Determine a loop base address corresponding to the loop identifier based on the loop identifier, the number of memory request modules, the requested allocation amount, the request length, and the boundary granularity; Determine a second request address based on the first request address, the request length, and the boundary termination address; The target request address is determined based on the second request address and the loop base address.

7. The method according to any one of claims 1 to 6, characterized in that: The expected merging result includes at least one of the following: an expected number of request merging and an address constraint item; the expected number of request merging is used to characterize the number of requests after each request in the test case is normally merged, and the address constraint item is used to characterize the range of target request addresses after each request in the test case is normally merged.

8. The method according to claim 7, characterized in that The configuration parameters include: a request processing threshold, which is used to characterize the maximum number of requests that the memory merging module can process at one time; a request allocation amount, which is used to characterize the number of requests that are allocated to the memory merging module; the configuration parameters also include the number of memory request modules; the test case corresponds to a merging factor, which is used to characterize the number of requests merged together; obtaining the expected merging result of the memory merging module includes: The expected number of request merges is determined based on the number of memory request modules, the request allocation amount, the request processing threshold, the merge factor, and the merge relationship.

9. The method according to claim 8, characterized in that The actual merge result includes the actual number of merge requests, and the actual merge result is verified based on the expected merge result to obtain the test result of the memory merge module, including: Compare the actual number of requests to be merged with the expected number of requests to be merged to obtain a comparison result; If the comparison result indicates that the actual number of requests to be merged is the same as the expected number of requests to be merged, the test result is that the test passed; When the comparison result indicates that the actual number of requests to be merged is different from the expected number of requests to be merged, the test result is that the test fails.

10. The method according to claim 7, characterized in that The configuration parameters include: request allocation amount, which is used to characterize the number of requests allocated to the memory merging module; request length, which is used to characterize the number of requests merged together; boundary granularity, which is used to characterize the length between the boundary start address and the boundary end address; the configuration parameters also include the number of memory request modules; the test case includes the target request address corresponding to the initial request identifier of each request; the address constraint item includes a first constraint item, a second constraint item and a third constraint item; the first constraint item is used to characterize the relationship between the target request address of the output request and the request length, the second constraint item is used to characterize the relationship between the target request address of the output request and the boundary granularity, and the third constraint item is used to characterize the relationship between the target request address of the output request and the request allocation amount; obtaining the expected merging result of the memory merging module includes: Determining a first constraint item based on the target request address and the request length; Determining a second constraint item based on the target request address, the request length, and the boundary granularity; A third constraint item is determined based on the target request address, the number of memory request modules, the request allocation amount, and the request length.

11. The method according to claim 10, characterized in that The actual merge result includes a target request address of at least one output request; and the actual merge result is verified based on the expected merge result to obtain a test result of the memory merge module, including: Match each of the target request addresses with the first constraint item, the second constraint item, and the third constraint item respectively to obtain a matching result corresponding to each of the target request addresses; wherein, if the target request address satisfies at least one of the first constraint item and the second constraint item and the target request address satisfies the third constraint item, the matching result is a successful match; in other cases, the matching result is a failed match; When all matching results are successful, the test result is test passed; In the case where at least one matching result is a matching failure, the test result is a test failure.

12. A testing device, characterized in that: The device comprises: An acquisition module, used to acquire a test case and an expected merging result of a memory merging module; each request in the test case and the expected merging result are determined based on a merging relationship between a preset number of requests; the merging relationship is used to characterize whether adjacent requests in the preset number of requests are merged; An input module, used for inputting each request in the test case into the memory merging module to obtain an actual merging result; A verification module is used to verify the actual merging result based on the expected merging result to obtain a test result of the memory merging module.

13. A computer device comprising a memory and a processor, wherein the memory stores a computer program executable on the processor, wherein: When the processor executes the program, the steps in the method according to any one of claims 1 to 11 are implemented.

14. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps in the method according to any one of claims 1 to 11 are implemented.

15. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instruction is executed by a processor, the steps in the method according to any one of claims 1 to 11 are implemented.