Coordinate determination method and device, electronic equipment, chip and storage medium

By calculating the coordinate increment of the second thread based on the coordinates of the first thread in the AI ​​chip, the problem of high computational complexity is solved, and more efficient coordinate determination and resource utilization are achieved.

CN119939075AActive Publication Date: 2025-05-06BEIJING X RING TECHNOLOGY CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202411999868.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-06
Estimated Expiration
2044-12-31

AI Technical Summary

Technical Problem

During the task scheduling process of AI chips, when quickly calculating the coordinates of the first thread in each warp to meet the throughput requirements, the existing technology has problems of high computing complexity and high resource consumption.

Method used

By determining the coordinate increment of the second thread based on the coordinates of the first thread and calculating the coordinates of the second thread using this increment, multiple division and multiplication operations are avoided, and the calculation complexity is reduced.

Benefits of technology

It improves the efficiency of coordinate determination, saves computing resources, reduces the use of hardware resources, and improves the performance of chip microarchitecture.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119939075A_ABST
    Figure CN119939075A_ABST
Patent Text Reader

Abstract

The invention provides a coordinate determination method and device, electronic equipment, a chip and a storage medium, and the method comprises the steps that the coordinate increment of a second thread is determined according to the coordinate of a first thread, the first thread and the second thread belong to a first thread block, and the serial number of the first thread is smaller than the serial number of the second thread; and determining the coordinate of the second thread according to the coordinate of the first thread and the coordinate increment of the second thread. The coordinate of the second thread can be determined according to the coordinate increment, the calculation complexity is reduced, the calculation efficiency is improved, and calculation resources are saved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing, and in particular to a coordinate determination method, device, electronic device, chip and storage medium. Background Art

[0002] The task scheduling of artificial intelligence (AI) chips is mainly divided into several levels: stream, kernel, block, warp, and thread. The scheduling unit is mainly responsible for the scheduling and information generation of the stream, kernel, and block layers. It splits the block into warps and sends them to the instruction fetch unit to execute the tasks issued. In the step of splitting the block into warps, it is necessary to quickly calculate the coordinates of the first thread in each warp and send them to the execution unit to achieve the throughput requirements. Summary of the invention

[0003] The present disclosure provides a coordinate determination method, device, electronic device, chip and storage medium to solve the problems in the related art.

[0004] The first aspect of the present disclosure proposes a coordinate determination method, the method comprising: determining a coordinate increment of a second thread according to the coordinates of a first thread, the first thread and the second thread belong to a first thread block, and the sequence number of the first thread is smaller than the sequence number of the second thread; determining the coordinates of the second thread according to the coordinates of the first thread and the coordinate increment of the second thread.

[0005] In some embodiments of the present disclosure, determining the coordinate increment of the second thread based on the coordinates of the first thread includes: determining the maximum dimension of the first thread block; determining the coordinate increment of the second thread based on the maximum dimension of the first thread block and the coordinates of the first thread, the coordinate increment including a first coordinate increment, a second coordinate increment, and a third coordinate increment, and the coordinates of the first thread include a first coordinate value, a second coordinate value, and a third coordinate value.

[0006] In some embodiments of the present disclosure, determining the first coordinate increment, the second coordinate increment, and the third coordinate increment of the second thread according to the maximum dimension of the first thread block and the coordinates of the first thread includes: determining the sequence number of the thread warp to which the first thread belongs, and determining the first remainder according to the sequence number of the thread warp and the maximum dimension of the first thread block; determining the second remainder and the third remainder according to the maximum dimension of the first thread block and the number of at least one thread included in the thread warp; determining the value of the first coordinate increment and the value of the second coordinate increment according to at least one of the first remainder, the second remainder, the third remainder and the first coordinate value; determining the value of the third coordinate increment according to the maximum dimension of the first thread block, the number of at least one thread included in the thread warp and the second coordinate increment.

[0007] In some embodiments of the present disclosure, determining the value of the first coordinate increment and the value of the second coordinate increment according to at least one of the first remainder, the second remainder, the third remainder and the first coordinate value includes: when the sum of the first remainder and the second remainder satisfies a first preset condition, determining the value of the first coordinate increment to be a first value; when the sum of the first remainder and the second remainder does not satisfy the first preset condition, determining the value of the first coordinate increment to be a second value; when the sum of the third remainder and the first coordinate value satisfies the second preset condition, determining the value of the second coordinate increment to be the first value; when the sum of the third remainder and the first coordinate value does not satisfy the second preset condition, determining the value of the second coordinate increment to be a second value.

[0008] In some embodiments of the present disclosure, the coordinates of the second thread include a fourth coordinate value, a fifth coordinate value, and a sixth coordinate value. Determining the coordinates of the second thread according to the coordinates of the first thread and the coordinate increment of the second thread includes: determining a fixed increment value according to the maximum dimension of the first thread block and the number of at least one thread included in the thread warp; determining the sum of the fixed increment value, the first coordinate value, and the first coordinate increment as the fourth coordinate value; determining the fifth coordinate value according to the second coordinate increment and the second coordinate value; determining the sixth coordinate value according to the third coordinate value, the third coordinate increment, the first coordinate increment, the maximum dimension of the first thread block, and the number of at least one thread included in the thread warp.

[0009] In some embodiments of the present disclosure, determining the fifth coordinate value based on the second coordinate increment and the second coordinate value includes: when the value of the second coordinate increment is the first value, determining the fifth coordinate value based on the third remainder, the first coordinate value, and the maximum dimension of the first thread block; when the value of the second coordinate increment is the second value, determining the sum of the third remainder and the first coordinate value as the fifth coordinate value.

[0010] In some embodiments of the present disclosure, determining the sixth coordinate value based on the third coordinate value, the third coordinate increment, the first coordinate increment, the maximum dimension of the first thread block, and the number of at least one thread included in the thread warp includes: determining the number of carries based on the maximum dimension of the first thread block, the number of at least one thread included in the thread warp, and the first coordinate increment; determining the sixth coordinate value based on the third coordinate value, the third coordinate increment, and the number of carries.

[0011] In some embodiments of the present disclosure, the method also includes: when the value of the first coordinate increment is a first value, determining a fourth remainder corresponding to the fourth coordinate value based on the maximum dimension of the first thread block, the first remainder, and the second remainder; when the value of the first coordinate increment is a second value, determining the sum of the first remainder and the second remainder as the fourth remainder corresponding to the fourth coordinate value.

[0012] The second aspect of the present disclosure proposes a coordinate determination device, which includes: a first processing unit, used to determine the coordinate increment of a second thread according to the coordinate of a first thread, the first thread and the second thread belong to a first thread block, the sequence number of the first thread is smaller than the sequence number of the second thread, and the coordinate of the first thread includes a first coordinate value, a second coordinate value and a third coordinate value; a second processing unit, used to determine the coordinate of the second thread according to the coordinate of the first thread and the coordinate increment of the second thread, the coordinate of the second thread includes a fourth coordinate value, a fifth coordinate value and a sixth coordinate value.

[0013] The third aspect embodiment of the present disclosure proposes an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method described in the first aspect embodiment of the present disclosure.

[0014] The fourth aspect embodiment of the present disclosure proposes a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute the method described in the first aspect embodiment of the present disclosure.

[0015] The fifth aspect embodiment of the present disclosure proposes a chip, characterized in that it includes at least one processor and a communication interface; the communication interface is used to receive signals input into the chip or signals output from the chip, and the processor communicates with the communication interface and implements the method described in the first aspect embodiment of the present disclosure through logic circuits or execution code instructions.

[0016] In summary, the coordinate determination method proposed in the present invention can determine the coordinate increment of the second thread based on the coordinates of the first thread determined last time, and determine the coordinates of the second thread based on the coordinate increment, which can avoid multiple division, multiplication and other operations, reduce calculation complexity, improve the efficiency of coordinate determination, and save computing resources.

[0017] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute improper limitations on the present disclosure.

[0019] Figure 1 A schematic diagram of a coordinate determination method provided in an embodiment of the present disclosure Figure 1 ;

[0020] Figure 2A schematic diagram of a coordinate determination method provided in an embodiment of the present disclosure Figure 2 ;

[0021] Figure 3 A schematic diagram of a coordinate determination method provided in an embodiment of the present disclosure Figure 3 ;

[0022] Figure 4 A schematic diagram of a coordinate calculation method in a SIMT architecture provided in an embodiment of the present disclosure Figure 1 ;

[0023] Figure 5 A schematic diagram of a coordinate calculation method in a SIMT architecture provided in an embodiment of the present disclosure Figure 2 ;

[0024] Figure 6 A schematic diagram of a coordinate calculation method in a SIMT architecture provided in an embodiment of the present disclosure Figure 3 ;

[0025] Figure 7 A schematic diagram of the structure of a coordinate determination device provided in an embodiment of the present disclosure;

[0026] Figure 8 A schematic diagram of the structure of an electronic device provided by an embodiment of the present disclosure;

[0027] Fig. 9 A schematic diagram of the chip structure provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0028] Embodiments of the present disclosure are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present disclosure, and should not be construed as limiting the present disclosure.

[0029] AI chips with a single instruction multiple thread (SIMT) architecture are mainly composed of several units, including scheduling, instruction fetching, computing, and memory access. The scheduling unit is the entry and starting point of SIMT and is responsible for the task initiation and scheduling of the entire AI chip.

[0030] The task scheduling of AI chips is mainly divided into several levels: stream, kernel, block, warp, and thread. The scheduling unit is mainly responsible for the scheduling and information generation of the stream, kernel, and block layers. It can split the block into warps and send them to the instruction fetch unit to execute the issued tasks. One of the key points in the block warp splitting step is to quickly calculate the coordinates of the first thread in each warp and send them to the execution unit to achieve a certain throughput requirement.

[0031] The specific process of calculating the coordinates of the first thread (thread0) in each warp may include:

[0032] 1. In order to obtain the three-dimensional coordinates of thread0 in the block space, let the maximum three-dimensional value of the block be a, b, c, and the block contains a×b×c threads.

[0033] 2. Each warp occupies 32 threads. For example, the coordinates of thread 0 in the first warp (warp0) are (x, y, z) = (0, 0, 0). The coordinates of thread 0 in warp1 require +32 threads. The obtained coordinate value may exceed a or may exceed a×b, so some calculations need to be done to correct the coordinate value.

[0034] 3. Assuming that the number of threads in the block three-dimensional space is expanded in 1 dimension, the coordinates of thread0 of the Nth warp (warpN) are 32×N.

[0035] 4. Then the coordinates of thead0 of warpN are (x, y, z), where z = [32×N / (a×b)]; the remainder is recorded as z_res; in order to avoid using division, the value of z is determined by comparison, and 1*a*b and 2*a*b…1024×a×b are calculated to compare with 32×N, and the integer divisible value is determined as the value of the z coordinate, and the maximum value of z is 1024.

[0036] 5. Further, y = [z_res / a], the remainder is the value of the x-coordinate, and the comparison method is also used to avoid division.

[0037] 6. After further obtaining y, determine x using the formula x=z_res-y×a.

[0038] Although this method can avoid division operations, it needs to generate many comparison items and consumes a large amount of multiplier resources. In addition, the timing of this method is poor, the critical path of calculation is long, first z, then y, and then x, with layer-by-layer logic dependencies and a large combinational logic depth.

[0039] Therefore, in order to solve the above problems, the present disclosure proposes a coordinate determination method, which can better improve the performance (Performance), power consumption (Power) and area (Area) of the chip microarchitecture.

[0040] The specific contents of this method are as follows.

[0041] Figure 1 A schematic diagram of a coordinate determination method provided in an embodiment of the present disclosure Figure 1 .like Figure 1 As shown, the method may include the following steps.

[0042] Step 101, determining a coordinate increment of a second thread according to the coordinates of a first thread.

[0043] In some embodiments, the first thread and the second thread belong to a first thread block, and the sequence number of the first thread is less than the sequence number of the second thread, that is, the first thread and the second thread belong to the same thread block. The first thread block may include multiple thread warps, and a thread warp may include multiple threads. Optionally, the first thread and the second thread belong to different thread warps.

[0044] Optionally, both the first thread and the second thread are the first threads in a warp, and the sequence number of the first thread is smaller than that of the second thread, and the sequence number of the warp to which the first thread belongs is also smaller than that of the warp to which the second thread belongs.

[0045] Optionally, when determining the coordinates of the first thread in a warp, the coordinates may be determined in order according to the warp numbers, that is, the number of the first thread in a warp with a smaller number may be determined first, that is, the coordinates of the first thread may be determined first and then the coordinates of the second thread.

[0046] In some embodiments, the coordinate increment of the second thread can be determined based on the coordinates of the first thread, and the coordinate increment of the second thread is used to determine the coordinates of the second thread. The coordinates of the first thread can be expressed as (x1, y1, z1), and the coordinates of the second thread can be expressed as (x2, y2, z2). The coordinates of the first thread include a first coordinate value, a second coordinate value, and a third coordinate value, wherein the first coordinate value is the above-mentioned z1, the second coordinate value is the above-mentioned x1, and the third coordinate value is the above-mentioned y1; the coordinates of the second thread include a fourth coordinate value, a fifth coordinate value, and a sixth coordinate value, wherein the fourth coordinate value is z2, the fifth coordinate value is x2, and the sixth coordinate value is y2.

[0047] In some embodiments, the coordinate increment of the second thread includes a first coordinate increment, a second coordinate increment and a third coordinate increment, wherein the first coordinate increment is the increment corresponding to the fourth coordinate value, the second coordinate increment is the increment corresponding to the fifth coordinate value, and the third coordinate increment is the increment corresponding to the sixth coordinate value.

[0048] In some embodiments, the coordinates of the first thread in a thread warp can be saved for calculating the coordinate increment of the first thread in a next thread warp. Determining the coordinate increment of the second thread based on the coordinates of the first thread includes: determining the maximum dimension of the first thread block; determining the coordinate increment of the second thread based on the maximum dimension of the first thread block and the coordinates of the first thread, the coordinate increment including a first coordinate increment, a second coordinate increment and a third coordinate increment, and the coordinates of the first thread include a first coordinate value, a second coordinate value and a third coordinate value.

[0049] In some embodiments, the occupied positions in the first thread block and the distribution of at least one thread can be determined based on the maximum dimension of the first thread block. For example, the arrangement of at least one thread in the first thread block can be determined. For example, there can be a maximum of 3 threads in the x-axis direction of the first thread block, a maximum of 5 threads in the y-direction, and a maximum of 6 threads in the z-direction. Then the maximum number of at least one thread included in the first thread block is 90.

[0050] Step 102, determining the coordinates of the second thread according to the coordinates of the first thread and the coordinate increment of the second thread.

[0051] In some embodiments, determining the coordinates of the second thread according to the coordinates of the first thread and the coordinate increment of the second thread includes: determining a fixed increment value according to the maximum dimension of the first thread block and the number of at least one thread included in the thread warp; determining the sum of the fixed increment value, the first coordinate value, and the first coordinate increment as a fourth coordinate value; determining a fifth coordinate value according to the second coordinate increment and the second coordinate value; determining a sixth coordinate value according to the third coordinate value, the third coordinate increment, the first coordinate increment, the maximum dimension of the first thread block, and the number of at least one thread included in the thread warp.

[0052] Among them, the maximum dimension of the first thread block may include: at least one of the maximum index value of the first thread block in the x-axis dimension, the maximum index value of the first thread block in the y-axis dimension, and the maximum index value of the first thread block in the z-axis dimension, that is, the maximum number of threads that can be arranged in the x-axis, y-axis and z-axis directions of the first thread block can be determined according to the maximum dimension of the first thread block.

[0053] In some embodiments, optionally, the fourth coordinate value and the fifth coordinate value may be determined simultaneously, and after the fourth coordinate value and the fifth coordinate value are determined, the sixth coordinate value is determined according to the third coordinate value, the third coordinate increment, the first coordinate increment, the maximum dimension of the first thread block, and the number of at least one thread included in the thread warp. This can improve the efficiency of coordinate determination, obtain better timing, and obtain a smaller combinational logic depth.

[0054] In summary, the above embodiments of the present disclosure can determine the coordinate increment of the second thread according to the coordinates of the first thread, and determine the coordinates of the second thread according to the coordinate increment, which can avoid multiple division, multiplication and other operations, reduce calculation complexity, improve the efficiency of coordinate determination, and save computing resources.

[0055] Figure 2 A schematic diagram of a coordinate determination method provided in an embodiment of the present disclosure Figure 2 .like Figure 2 As shown, based on Figure 1 In the illustrated embodiment, the method comprises the following steps.

[0056] Step 201: determine the maximum dimension of the first thread block.

[0057] In some embodiments, optionally, the distribution mode of at least one thread in the first thread block may be determined, for example, the maximum number of threads in the x-axis direction, the maximum number of threads in the y-axis direction, and the maximum number of threads in the z-axis direction in the first thread block may be determined. In some embodiments, optionally, the number of at least one thread warp included in the first thread block and the number of threads included in each thread warp may be determined, for example, a thread warp may include 32 threads.

[0058] Step 202: Determine the coordinate increment of the second thread according to the maximum dimension of the first thread block and the coordinates of the first thread.

[0059] In some embodiments, determining the first coordinate increment, the second coordinate increment, and the third coordinate increment of the second thread according to the maximum dimension of the first thread block and the coordinates of the first thread includes: determining the sequence number of the thread warp to which the first thread belongs, and determining the first remainder according to the sequence number of the thread warp and the maximum dimension of the first thread block; determining the second remainder and the third remainder according to the maximum dimension of the first thread block and the number of at least one thread included in the thread warp; determining the value of the first coordinate increment and the value of the second coordinate increment according to at least one of the first remainder, the second remainder, the third remainder and the first coordinate value; determining the value of the third coordinate increment according to the maximum dimension of the first thread block, the number of at least one thread included in the thread warp, and the second coordinate increment.

[0060] In some embodiments, the first remainder is the remainder in the z direction when calculating the coordinates of the first thread, that is, the first remainder is the remainder determined in the last coordinate calculation, that is, the number of threads remaining after the number of threads in the x direction and the y direction reaches the maximum value, which can be expressed as z_res_d. The first remainder can be determined according to the sequence number N of the thread warp to which the first thread belongs and the maximum dimension of the first thread block. For example, when the number of threads in the x direction and the y direction in the first thread block is a and the number of threads in the y direction is b, and the number of threads in the thread warp to which the first thread belongs is 32, the remainder of the formula 32×N / (a×b) plus the previous remainder can be determined as the first remainder. For example, when N=1, the corresponding remainder is 2, when N=2, the corresponding remainder is 2 plus 2, which is 4, when N=3, the corresponding remainder is 6, and so on. Since in the thread block, different thread warps are arranged in a certain arrangement order, and multiple threads in a thread warp are also arranged according to a certain rule, when determining the coordinates of the second thread, it can be determined according to the arrangement rule of the threads in the first thread block. For example, the first thread is split into threads along the z-axis, and the number of threads in each layer is a×b. The coordinates of the first thread in the first thread warp in the first thread block can be (0,0,0). When the number of threads in a thread warp is 32, when determining the coordinates of the first thread in the second thread warp, it is necessary to determine it according to the arrangement of the 32 threads in the first thread block.

[0061] For example, in the first thread block, the threads need to be filled from the origin of the coordinates along the x-axis direction, that is, the coordinates of the first thread are (0,0,0), then the coordinates of the second thread are (1,0,0), and the coordinates of the third thread are (2,0,0), and so on. When the number of threads in the x-axis direction is equal to a, that is, when the coordinates reach (a,0,0), the first column and the first layer in the x-axis direction are filled, and then the y-direction is carried by 1, and the second column of the first layer is filled, that is, the coordinates are expressed as (0,1,0), (1,1,0), and so on, until the number of threads in the y-direction is greater than b, then the first layer is filled, and there are a×b threads in total. At this time, the z-direction is carried by 1, and the second layer is filled. That is, the first remainder is the number of threads remaining in the z-direction after filling the 32 threads in the first thread. For example, when the 32 threads of the first thread warp are filled into the first layer, and the first layer is filled but the remaining threads are not enough to fill the second layer, the number of threads remaining in the second layer is the above-mentioned first remainder. In some embodiments, other methods or formulas may be used to determine the first remainder, which is not limited in the present disclosure.

[0062] In some embodiments, the second remainder is used to indicate the number of remaining threads when the second thread fills the x-axis and the y-axis. The second remainder can be determined according to the maximum dimension of the first thread block and the number of at least one thread included in the thread warp. For example, when the number of threads included in the x-direction of the first thread block is a, the number of threads included in the y-direction is b, and the number of threads included in the thread warp is 32, the second remainder can be expressed as res_of_32_ab, that is, the remainder of (32 / ab) is taken as the second remainder, wherein the first remainder is the total number (cumulative number) of threads remaining when filling the x-axis and the y-axis obtained when calculating the coordinates of the first thread, and the second remainder is the newly generated remainder when calculating the coordinates of the second thread, that is, the number of newly remaining threads. The first remainder and the second remainder are added to obtain the total remainder when calculating the coordinates of the second thread.

[0063] In some embodiments, the third remainder is the number of threads remaining in the x direction after multiple carries in the y direction. The third remainder can be determined according to the maximum dimension of the first thread block and the number of at least one thread contained in the thread warp. For example, when the number of threads contained in the x direction in the first thread block is a and the number of threads contained in the thread warp is 32, the second remainder can be expressed as res_of_32_a, that is, the remainder of (32 / a) is taken as the third remainder. For example, when a=5, 6 carries are required in the y direction. The number of threads remaining in the x direction after the last carry is 2, and the third remainder is 2 at this time. That is, when the carry condition is not met in the x direction, the number of existing threads, for example, when the number of threads in the x direction is 0, the remainder is 0, and when the number of threads in the x direction is 5, the y direction carries, and at this time, the number of threads in the x direction on the next column needs to be determined as the third remainder.

[0064] In some embodiments, according to at least one of the first remainder, the second remainder, the third remainder and the first coordinate value, determining the value of the first coordinate increment and the value of the second coordinate increment includes: when the sum of the first remainder and the second remainder meets the first preset condition, determining the value of the first coordinate increment to be the first value; when the sum of the first remainder and the second remainder does not meet the first preset condition, determining the value of the first coordinate increment to be the second value; when the sum of the third remainder and the first coordinate value meets the second preset condition, determining the value of the second coordinate increment to be the first value; when the sum of the third remainder and the first coordinate value does not meet the second preset condition, determining the value of the second coordinate increment to be the second value. Wherein, the first coordinate increment is the increment corresponding to the fourth coordinate value, the second coordinate increment is the increment corresponding to the fifth coordinate value, and the third coordinate increment is the increment corresponding to the sixth coordinate value.

[0065] In some embodiments, the first coordinate increment may be used to indicate whether the fourth coordinate value still needs to be carried compared to the first coordinate value on the basis of a fixed number of carry times. Optionally, the first coordinate increment is an increment of change in the z-axis direction. The first coordinate increment may be expressed as z_res_incr_flag. The first coordinate increment may be used to indicate whether the sum of the accumulated remainders satisfies a carry condition. For example, when the number of threads included in the x direction in the first thread block is a, the number of threads included in the y direction is b, and the number of threads included in the thread warp is 32, when N=5, the sum of the first remainder and the second remainder is 10, which is exactly equal to the number of threads in one layer a×b. If a carry is required in the z direction, then the value of the first coordinate increment is a first value. The first value may be 1, indicating that a carry is required. The second value may indicate that a carry is not required. For example, the second value may be 0. The first preset condition may be expressed as follows:

[0066] z_res_incr_flag=z_res_d+res_of_32_ab>=ab? 1:0

[0067] That is, when the sum of the first remainder z_res_d and the second remainder res_of_32_ab satisfies the z-axis carry condition, that is, is greater than or equal to the maximum number of threads in one layer a×b, the value of the first coordinate increment z_res_incr_flag is 1, otherwise it is 0.

[0068] In some embodiments, the second coordinate increment may be used to indicate whether the fifth coordinate value needs to be carried in the y direction relative to the second coordinate value when calculating the coordinates of the second thread. For example, when the sum of the third remainder and the first coordinate value satisfies the second preset condition, the value of the second coordinate increment is determined to be the first value. For example, when the number of threads included in the x direction of the first thread block is a, the number of threads included in the y direction is b, and the number of threads included in the thread warp is 32, the second preset condition may be expressed as follows:

[0069] x_incr_flag=x_d+res_of_32_a>=a? 1:0

[0070] That is, the third remainder res_of_32_a is the number of threads remaining in the x-direction after the thread warp to which the first thread belongs is filled. That is, according to the coordinates of the first thread and the number of remaining threads, it can be determined whether to carry in the y-direction when calculating the x-axis coordinate of the second thread.

[0071] In some embodiments, the third coordinate increment is used to indicate the number of times the sixth coordinate is carried in the y-axis direction compared with the third coordinate when calculating the coordinates of the second thread, that is, the carry amount. The value of the third coordinate increment is determined according to the maximum dimension of the first thread block, the number of at least one thread included in the thread warp, and the second coordinate increment. For example, when the number of threads included in the x-direction of the first thread block is a, the number of threads included in the y-direction is b, and the number of threads included in the thread warp is 32, the calculation formula of the third coordinate increment is expressed as follows:

[0072]

[0073] That is, By rounding, we can get the fixed number of carry-overs on the y-axis when calculating the coordinates of the second thread. For example, the coordinates of the first thread are (0,0,0,), the number of threads in the first thread block in the x-direction is a=5, and the number of threads in the thread warp is 32. After the thread warp to which the first thread belongs is filled, the x-axis can be filled 6 times. At this time, the y-direction needs to be carried 6 times, which is the fixed number of carry-overs on the y-axis. The second coordinate increment x_incr_flag actually indicates whether the number of threads in the x-axis direction meets the y-axis carry-over condition, that is, the second coordinate increment can be used as the variable increment part of the third coordinate increment, that is, in addition to the fixed number of carry-overs, when the accumulated number of remaining threads is greater than the maximum number of threads on the x-axis a, an additional carry-over is required in the y-axis direction.

[0074] In some embodiments, the method also includes: when the value of the first coordinate increment is a first value, determining a fourth remainder corresponding to the fourth coordinate value based on the maximum dimension of the first thread block, the first remainder, and the second remainder; when the value of the first coordinate increment is a second value, determining the sum of the first remainder and the second remainder as the fourth remainder corresponding to the fourth coordinate value.

[0075] In some embodiments, optionally, the fourth remainder is the cumulative remainder in the z-axis direction determined when calculating the coordinates of the second thread, and the fourth remainder and the first remainder are both cumulative remainders in the z-axis direction. In the first thread block, the number of threads included in the x direction is a, and the number of threads included in the y direction is b. When the value of the remainder is greater than or equal to a×b, a carry is required in the z-axis direction, wherein the first remainder is the cumulative remainder determined when calculating the coordinates of the first thread, and the second thread is the cumulative remainder when calculating the second thread.

[0076] In some embodiments, for example, when the number of threads included in the x direction in the first thread block is a, the number of threads included in the y direction is b, and the number of threads included in the thread warp is 32, the fourth remainder may be determined according to the following formula:

[0077] z_res=z_res_incr_dlag? (res_add-ab):(res_add)

[0078] In other words, when the value of the first coordinate increment is the first value, it is necessary to carry in the z-axis direction. The above res_add is the sum of the first remainder and the second remainder, that is, the remainder in the z-axis direction determined when determining the coordinates of the second thread. When a carry is needed, the remainder needs to be subtracted from the number of threads included in a layer, that is, ab, to obtain a new remainder as the fourth remainder; when the value of the first coordinate increment is the second value, the remainder does not meet the requirement of z-axis carry. At this time, the sum of the first remainder and the second remainder can be directly determined as the fourth remainder to update the accumulated remainder in the z-axis direction.

[0079] In summary, the above-mentioned embodiments of the present application can determine the coordinate increment of the second thread according to the coordinates of the first thread and the distribution pattern of threads in the thread block, so as to facilitate determining the coordinates of the second thread according to the coordinate increment, and can avoid the problems of complex calculations caused by traversing to determine the coordinates of the second thread, low calculation efficiency affecting business execution efficiency, and the like.

[0080] Figure 3 A schematic diagram of a coordinate determination method provided in an embodiment of the present disclosure Figure 3 .like Figure 3 As shown, based on Figure 1 In the illustrated embodiment, the method comprises the following steps.

[0081] Step 301 : determining a fixed increment value according to a maximum dimension of a first thread block and the number of at least one thread included in a warp.

[0082] In some embodiments, optionally, the fixed increment value is the fixed increment value corresponding to the fourth coordinate, that is, the number of fixed carry times of the fourth coordinate relative to the first coordinate, that is, the number of carry times in the z-axis direction. Optionally, the fixed increment value can be determined according to the maximum dimension of the first thread block and the number of at least one thread contained in the thread warp. For example, when the number of threads contained in the x-direction of the first thread block is a and the number of threads contained in the y-direction is b, and the number of threads contained in the thread warp is 32, the fixed increment value can be expressed as [32 / ab], that is, 32 / ab is rounded. For example, when a=2 and b=5, when 32 threads in the thread warp to which the first thread belongs are filled, 3 layers of a×b can be filled, that is, 3 carry times are required in the z-axis direction, and the fixed increment value is 3 at this time.

[0083] Step 302: determine the sum of the fixed increment value, the first coordinate value, and the first coordinate increment as the fourth coordinate value.

[0084] In some embodiments, the fourth coordinate value is a coordinate value z2 of the second thread in the z-axis direction. Optionally, the sum of the fixed increment value, the first coordinate value, and the first coordinate increment may be determined as the fourth coordinate value. When the number of threads included in the thread warp is 32, the fourth coordinate value may be determined by the following formula:

[0085] z=z_d+[32 / ab]+z_res_incr_flag

[0086] That is, the coordinates of the second thread can be obtained by adding the fixed increment [32 / ab] and the variable increment z_res_incr_dlag to the z-axis coordinate z_d of the first thread.

[0087] Step 303: determine the fifth coordinate value according to the second coordinate increment and the second coordinate value.

[0088] In some embodiments, determining the fifth coordinate value based on the second coordinate increment and the second coordinate value includes: when the value of the second coordinate increment is the first value, determining the fifth coordinate value based on the third remainder, the first coordinate value, and the maximum dimension of the first thread block; when the value of the second coordinate increment is the second value, determining the sum of the third remainder and the first coordinate value as the fifth coordinate value.

[0089] In some embodiments, when the value of the second coordinate increment is the first value, it indicates that the number of threads in the x-axis reaches the maximum value, and the y-axis needs to be carried once. When the number of threads in the x-direction of the first thread block is a, the number of threads in the y-direction is b, and the number of threads in the thread warp is 32, the fifth coordinate value can be determined according to the following formula:

[0090] x=x_incr_dlag? res_of_32_a+x_d-a:res_of_32_a+x_d

[0091] That is, when calculating the x-axis coordinate of the second thread, the x-axis coordinate x_d of the first thread is added to the number of threads remaining on the x-axis res_of_32_a after filling all the threads in the thread warp where the first thread is located, and a new x-axis coordinate value can be obtained. When the coordinate value exceeds the maximum number a of the x-axis in the first thread block, it is necessary to carry in the y-axis direction, that is, it is necessary to subtract a, and the x-axis coordinate of the second thread after the carry can be obtained; when no carry is required, the above sum can be directly determined as the x-axis coordinate of the second thread.

[0092] Step 304 : determining a sixth coordinate value according to the third coordinate value, the third coordinate increment, the first coordinate increment, the maximum dimension of the first thread block, and the number of at least one thread included in the thread warp.

[0093] In some embodiments, determining the sixth coordinate value based on the third coordinate value, the third coordinate increment, the first coordinate increment, the maximum dimension of the first thread block, and the number of at least one thread included in the thread warp includes: determining the number of carries based on the maximum dimension of the first thread block, the number of at least one thread included in the thread warp, and the first coordinate increment; determining the sixth coordinate value based on the third coordinate value, the third coordinate increment, and the number of carries.

[0094] In some embodiments, the number of carries is the number of y-axis coordinate offsets caused by not considering the carry on the z-axis when calculating the y-axis coordinate of the second thread. Since only the number of carries on the y-axis is considered when determining the third coordinate increment, and the carry on the z-axis is not considered, the actual number of carries on the y-axis is less than the number of carries determined by the third coordinate increment. When the y-axis coordinate of the second thread is determined based on the number of carries determined by the third coordinate increment, the y-axis coordinate is inaccurate. Therefore, it is necessary to consider the number of carries on the z-axis to determine the number of carries. In the first thread block, the number of threads included in the x-direction is a, the number of threads included in the y-direction is b, and the number of threads included in the thread warp is 32. The number of carries can be expressed as follows:

[0095]

[0096] in, is the actual number of carry times on the z-axis, where is the fixed increment of the z-axis, z_res_incr_flag is the first coordinate increment mentioned above, because each carry of the z-axis indicates that the y-axis coordinate value needs to be reduced by b, that is, after each carry, the y-axis coordinate value can be changed from b to 0, and the coordinate value is reduced by b. Therefore, the number of carries is the amount of coordinate value erroneously increased due to not considering the carry of the z-axis, that is, the offset of the y-axis coordinate of the second thread.

[0097] In some embodiments, the sixth coordinate value may be determined by the third coordinate value, the third coordinate increment, and the number of carry positions. The specific formula is as follows:

[0098]

[0099] In other words, the y-axis coordinate y_d of the first thread can be added with the carry amount, and then subtracted from the offset to obtain the actual y-axis coordinate value of the second thread, wherein the carry amount is the third coordinate increment y_incr.

[0100] In summary, the above embodiments of the present disclosure can determine the coordinates of the second thread according to the first coordinate increment, the second coordinate increment and the third coordinate increment, can implement the coordinates of the first thread determined last time, determine the coordinate increment of the second thread, and determine the coordinates of the second thread according to the coordinate increment, can avoid multiple division, multiplication and other operations, reduce the computational complexity, improve the efficiency of coordinate determination, and save computing resources. In addition, in terms of circuit structure, the use of more dividers, multipliers, and comparators can be reduced, the use area is better, and hardware resources can be saved.

[0101] The technical solution of the present disclosure is further described in detail below in conjunction with specific application examples.

[0102] The following is a method for coordinate calculation in a SIMT architecture provided by an embodiment of the present disclosure. The method can determine the coordinates (x, y, z) of the first thread in the current warp through the coordinates (x_d, y_d, z_d) of the first thread in the previous warp. The specific content of the method is as follows.

[0103] like Figure 4 As shown in the figure, it is a specific process to determine the z value in the coordinates of the first thread in the current warp. x_d, y_d, z_d are the coordinates of thread0 of the previous warp in the three-dimensional space of the block. The circuit is used for storage as the coordinate starting point of the next warp calculation. z_res_d is the remainder of z determined when the coordinates of thread0 in the previous warp are calculated. z=[32×N / (a×b)], z is an integer divisor, z_res is the remainder, and N is the warp number.

[0104] To avoid using division, we can use incremental calculation. When the number of threads in a warp is 32, we need to add 32 threads to calculate thread0 in the new warp. By considering the number of carries to z each time, we can get the new z.

[0105] The number of carries for z each time 32 is added can be composed of two parts, one is the quotient of 32 divided by ab, and the other is the remainder of 32 divided by a×b. Add the previous z_res_d to get a new result to see if there is a carry. If there is a carry, it is increased by 1. The specific formula for calculating z is as follows:

[0106] z=z_d+[32 / ab]+z_res_incr_flag

[0107] z_res_incr_flag=z_res_d+res_of_32_ab>=ab? 1:0

[0108] z_res=z_res_incr_flag? (res_add-ab):(res_add)

[0109] Among them, z_res_incr_flag is used to indicate the amount of carry change, that is, whether to carry z, and res_of_32_ab means taking the remainder of 32 / ab.

[0110] like Figure 5 As shown, to determine the specific process of the x value in the coordinates of the first thread in the current warp, the solution for calculating x is to consider the incremental calculation method, x_d plus the remainder of 32 divided by a, and optionally, the quotient of 32 divided by a is the number of carries to y.

[0111] At the same time, we need to consider whether the remainder of 32 divided by a plus x_d will carry. If it does, the new x will be subtracted from a. The calculation formula is as follows:

[0112] x_incr_flag=x_d+res_of_32_a>=a? 1:0

[0113] x=x_incr_flag? res_of_32_a+x_d-a:res_of_32_a+x_d

[0114] like Figure 6 As shown in the figure, to determine the specific process of the y value in the coordinates of the first thread in the current warp, the way to calculate y is to calculate z and x after calculating. First, y_d needs to add the carry amount after x plus 32, and then subtract the number of carries to z.

[0115] y_incr is the carry amount to y after adding 32 to x, which can be obtained when calculating x. However, the carry may carry to z, resulting in the y coordinate to subtract multiple b. The specific number of carries consists of two parts, which can be obtained in the process of calculating z, including the quotient of 32 / ab, and z_res-incr_flag. These two parts are the number of carries of the y coordinate to z. That is, in the process of calculating x and z, the variables required to calculate y are obtained, and then y is calculated, which better controls the depth of the combinational logic. The specific formula is as follows:

[0116]

[0117] In summary, the above examples of the present disclosure have a smaller combinational logic depth, better timing, and reduce the number of calculations such as division, multiplication, and comparison. The use of this method can have certain advantages for the chip architecture, such as better usage area and reduced use of dividers, multipliers, comparators, etc.

[0118] Figure 7FIG. 7 is a schematic diagram of a coordinate determination device 700 provided in an embodiment of the present disclosure. Figure 7 As shown, the device includes: a first processing unit 710, used to determine the coordinate increment of the second thread according to the coordinate of the first thread, the first thread and the second thread belong to a first thread block, the sequence number of the first thread is smaller than the sequence number of the second thread, and the coordinate of the first thread includes a first coordinate value, a second coordinate value and a third coordinate value; a second processing unit 720, used to determine the coordinate of the second thread according to the coordinate of the first thread and the coordinate increment of the second thread, and the coordinate of the second thread includes a fourth coordinate value, a fifth coordinate value and a sixth coordinate value.

[0119] In some embodiments, the first processing unit is also used to determine the maximum dimension of the first thread block; determine the coordinate increment of the second thread based on the maximum dimension of the first thread block and the coordinates of the first thread, the coordinate increment includes a first coordinate increment, a second coordinate increment and a third coordinate increment, and the coordinates of the first thread include a first coordinate value, a second coordinate value and a third coordinate value.

[0120] In some embodiments, the first processing unit is also used to determine the serial number of the thread warp to which the first thread belongs, determine the first remainder according to the serial number of the thread warp and the maximum dimension of the first thread block; determine the second remainder and the third remainder according to the maximum dimension of the first thread block and the number of at least one thread included in the thread warp; determine the value of the first coordinate increment and the value of the second coordinate increment according to the first remainder, the second remainder, the third remainder and at least one of the first coordinate value; determine the value of the third coordinate increment according to the maximum dimension of the first thread block, the number of at least one thread included in the thread warp and the second coordinate increment.

[0121] In some embodiments, the first processing unit is also used to determine that the value of the first coordinate increment is a first value when the sum of the first remainder and the second remainder meets the first preset condition; determine that the value of the first coordinate increment is a second value when the sum of the first remainder and the second remainder does not meet the first preset condition; determine that the value of the second coordinate increment is a first value when the sum of the third remainder and the first coordinate value meets the second preset condition; and determine that the value of the second coordinate increment is a second value when the sum of the third remainder and the first coordinate value does not meet the second preset condition.

[0122] In some embodiments, the second processing unit is also used for the coordinates of the second thread to include a fourth coordinate value, a fifth coordinate value and a sixth coordinate value. According to the coordinates of the first thread and the coordinate increment of the second thread, determining the coordinates of the second thread includes: determining a fixed increment value according to the maximum dimension of the first thread block and the number of at least one thread included in the thread warp; determining the sum of the fixed increment value, the first coordinate value and the first coordinate increment as the fourth coordinate value; determining the fifth coordinate value according to the second coordinate increment and the second coordinate value; determining the sixth coordinate value according to the third coordinate value, the third coordinate increment, the first coordinate increment, the maximum dimension of the first thread block, and the number of at least one thread included in the thread warp.

[0123] In some embodiments, the second processing unit is also used to determine the fifth coordinate value based on the third remainder, the first coordinate value, and the maximum dimension of the first thread block when the value of the second coordinate increment is the first value; and when the value of the second coordinate increment is the second value, determine the sum of the third remainder and the first coordinate value as the fifth coordinate value.

[0124] In some embodiments, the second processing unit is further used to determine the number of carries based on the maximum dimension of the first thread block, the number of at least one thread included in the thread warp, and the first coordinate increment; and determine the sixth coordinate value based on the third coordinate value, the third coordinate increment, and the number of carries.

[0125] In some embodiments, the coordinate determination device also includes a third processing unit, which is used to determine a fourth remainder corresponding to the fourth coordinate value based on the maximum dimension of the first thread block, the first remainder, and the second remainder when the value of the first coordinate increment is a first value; and to determine that the sum of the first remainder and the second remainder is the fourth remainder corresponding to the fourth coordinate value when the value of the first coordinate increment is a second value.

[0126] In summary, the coordinate determination device 700 can determine the coordinate increment of the second thread based on the coordinates of the first thread determined last time, and determine the coordinates of the second thread based on the coordinate increment, which can avoid multiple division, multiplication and other operations, reduce calculation complexity, improve the efficiency of coordinate determination, and save computing resources.

[0127] In the embodiments provided in the present application, the methods and devices provided in the embodiments of the present application are introduced. In order to implement the functions in the methods provided in the embodiments of the present application, the electronic device may include a hardware structure and a software module, and implement the functions in the form of a hardware structure, a software module, or a hardware structure plus a software module. A function of the functions may be executed in the form of a hardware structure, a software module, or a hardware structure plus a software module.

[0128] Figure 88 is a block diagram of an electronic device 800 for implementing the above method according to an exemplary embodiment. For example, the electronic device 800 may be a mobile phone, a computer, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0129] Reference Figure 8 , the electronic device 800 may include one or more of the following components: a processing component 802 , a memory 804 , a power component 806 , a multimedia component 808 , an audio component 810 , an input / output (I / O) interface 812 , a sensor component 814 , and a communication component 816 .

[0130] The processing component 802 generally controls the overall operation of the electronic device 800, such as operations associated with display, phone calls, data communications, camera operations, and recording operations. The processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the above-mentioned method. In addition, the processing component 802 may include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.

[0131] The memory 804 is configured to store various types of data to support operations on the electronic device 800. Examples of such data include instructions for any application or method operating on the electronic device 800, contact data, phone book data, messages, pictures, videos, etc. The memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEP ROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0132] The power supply component 806 provides power to the various components of the electronic device 800. The power supply component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 800.

[0133] The multimedia component 808 includes a screen that provides an output interface between the electronic device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor may not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the electronic device 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera may receive external multimedia data. Each front camera and rear camera may be a fixed optical lens system or have a focal length and optical zoom capability.

[0134] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC), and when the electronic device 800 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode, the microphone is configured to receive an external audio signal. The received audio signal can be further stored in the memory 804 or sent via the communication component 816. In some embodiments, the audio component 810 also includes a speaker for outputting audio signals.

[0135] I / O interface 812 provides an interface between processing component 802 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include but are not limited to: home button, volume button, start button, and lock button.

[0136] The sensor assembly 814 includes one or more sensors for providing various aspects of status assessment for the electronic device 800. For example, the sensor assembly 814 can detect the open / closed state of the electronic device 800, the relative positioning of components, such as the display and keypad of the electronic device 800, and the sensor assembly 814 can also detect the position change of the electronic device 800 or a component of the electronic device 800, the presence or absence of contact between the user and the electronic device 800, the orientation or acceleration / deceleration of the electronic device 800, and the temperature change of the electronic device 800. The sensor assembly 814 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 814 may also include an optical sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 814 may also include an accelerometer, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0137] The communication component 816 is configured to facilitate wired or wireless communication between the electronic device 800 and other devices. The electronic device 800 can access a wireless network based on a communication standard, such as WiFi, 2G or 3G, 4G LTE, 5G NR (NewRadio) or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.

[0138] In an exemplary embodiment, the electronic device 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above methods.

[0139] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions, and the above instructions can be executed by the processor 820 of the electronic device 800 to complete the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0140] The embodiments of the present disclosure further provide a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute the method described in the above embodiments of the present disclosure.

[0141] An embodiment of the present disclosure further proposes a communication system, which includes a terminal and a network device. The terminal is used to implement the method described in the embodiment of the first aspect of the present disclosure, and the network device can be used to implement the method described in the embodiment of the second aspect of the present disclosure.

[0142] Fig. 9 FIG. 1 is a schematic diagram of a chip 900 for implementing the above method according to an exemplary embodiment. Fig. 9 The chip 900 includes a communication interface 901 and at least one processor 902. The communication interface 901 is used to receive signals input into the chip 900 or signals output from the above chip 900. The processor 902 communicates with the communication interface 901 and implements the method described in the above embodiment of the present disclosure through logic circuits or executing code instructions.

[0143] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0144] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "illustrative embodiments", "examples", "specific examples" or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiments or examples are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any at least one embodiment or example in a suitable manner.

[0145] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code that includes one or more executable instructions for implementing the steps of a specific logical function or process, and the scope of the preferred embodiments of the present invention includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present invention belong.

[0146] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processing module, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in combination with these instruction execution systems, devices or apparatuses. For the purposes of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in combination with these instruction execution systems, devices or apparatuses. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection having at least one wiring (control method), a portable computer disk box (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or otherwise processing in a suitable manner if necessary, and then stored in a computer memory.

[0147] It should be understood that the various parts of the embodiments of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0148] A person skilled in the art may understand that all or part of the steps in the method for implementing the above-mentioned embodiment may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, which, when executed, includes one or a combination of the steps of the method embodiment.

[0149] In addition, each functional unit in each embodiment of the present invention may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium. The above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc.

[0150] Although the embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and are not to be construed as limitations of the present invention. A person skilled in the art may make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.

Claims

1. A coordinate determination method, characterized in that: The method comprises: Determining a coordinate increment of a second thread according to a coordinate of a first thread, wherein the first thread and the second thread belong to a first thread block, and a sequence number of the first thread is smaller than a sequence number of the second thread; The coordinates of the second thread are determined according to the coordinates of the first thread and the coordinate increment of the second thread.

2. The method according to claim 1, characterized in that Determining the coordinate increment of the second thread according to the coordinate of the first thread includes: determining a maximum dimension of the first thread block; According to the maximum dimension of the first thread block and the coordinates of the first thread, the coordinate increment of the second thread is determined, the coordinate increment includes a first coordinate increment, a second coordinate increment and a third coordinate increment, and the coordinates of the first thread include a first coordinate value, a second coordinate value and a third coordinate value.

3. The method according to claim 2, characterized in that Determining the first coordinate increment, the second coordinate increment, and the third coordinate increment of the second thread according to the maximum dimension of the first thread block and the coordinates of the first thread comprises: Determine a sequence number of a thread warp to which the first thread belongs, and determine a first remainder according to the sequence number of the thread warp and a maximum dimension of the first thread block; Determine a second remainder and a third remainder according to the maximum dimension of the first thread block and the number of at least one thread included in the thread warp; Determine the value of the first coordinate increment and the value of the second coordinate increment according to at least one of the first remainder, the second remainder, the third remainder and the first coordinate value; The value of the third coordinate increment is determined according to the maximum dimension of the first thread block, the number of at least one thread included in the thread warp, and the second coordinate increment.

4. The method according to claim 3, characterized in that Determining the value of the first coordinate increment and the value of the second coordinate increment according to at least one of the first remainder, the second remainder, the third remainder and the first coordinate value comprises: When the sum of the first remainder and the second remainder satisfies a first preset condition, determining the value of the first coordinate increment to be a first value; When the sum of the first remainder and the second remainder does not satisfy the first preset condition, determining the value of the first coordinate increment to be a second value; When the sum of the third remainder and the first coordinate value satisfies a second preset condition, determining the value of the second coordinate increment to be a first value; When the sum of the third remainder and the first coordinate value does not satisfy the second preset condition, the value of the second coordinate increment is determined to be a second value.

5. The method according to claim 3, characterized in that: The coordinates of the second thread include a fourth coordinate value, a fifth coordinate value, and a sixth coordinate value, and determining the coordinates of the second thread according to the coordinates of the first thread and the coordinate increment of the second thread includes: Determine a fixed increment value according to a maximum dimension of the first thread block and the number of at least one thread included in the thread warp; Determine the sum of the fixed increment value, the first coordinate value, and the first coordinate increment as the fourth coordinate value; Determine the fifth coordinate value according to the second coordinate increment and the second coordinate value; The sixth coordinate value is determined according to the third coordinate value, the third coordinate increment, the first coordinate increment, a maximum dimension of the first thread block, and the number of at least one thread included in the thread warp.

6. The method according to claim 4, characterized in that Determining the fifth coordinate value according to the second coordinate increment and the second coordinate value comprises: When the value of the second coordinate increment is the first value, determining the fifth coordinate value according to the third remainder, the first coordinate value, and the maximum dimension of the first thread block; When the value of the second coordinate increment is a second value, the sum of the third remainder and the first coordinate value is determined as the fifth coordinate value.

7. The method according to claim 4, characterized in that The determining the sixth coordinate value according to the third coordinate value, the third coordinate increment, the first coordinate increment, the maximum dimension of the first thread block, and the number of at least one thread included in the thread warp comprises: Determining a carry number according to a maximum dimension of the first thread block, a number of at least one thread included in the thread warp, and the first coordinate increment; The sixth coordinate value is determined according to the third coordinate value, the third coordinate increment, and the carry quantity.

8. The method according to claim 4, characterized in that The method further comprises: When the value of the first coordinate increment is a first value, determining a fourth remainder corresponding to the fourth coordinate value according to the maximum dimension of the first thread block, the first remainder, and the second remainder; When the value of the first coordinate increment is a second value, the sum of the first remainder and the second remainder is determined to be a fourth remainder corresponding to the fourth coordinate value.

9. A coordinate determination device, the device comprising: a first processing unit, configured to determine a coordinate increment of a second thread according to a coordinate of a first thread, wherein the first thread and the second thread belong to a first thread block, a sequence number of the first thread is smaller than a sequence number of the second thread, and the coordinate of the first thread includes a first coordinate value, a second coordinate value, and a third coordinate value; The second processing unit is used to determine the coordinates of the second thread according to the coordinates of the first thread and the coordinate increment of the second thread, wherein the coordinates of the second thread include a fourth coordinate value, a fifth coordinate value and a sixth coordinate value.

10. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 8.

11. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-8.

12. A chip, characterized in that: It includes at least one processor and a communication interface; the communication interface is used to receive a signal input to the chip or a signal output from the chip, and the processor communicates with the communication interface and implements the method as described in any one of claims 1 to 8 through a logic circuit or executing code instructions.

Citation Information

Patent Citations

  • Configurable thread ordering for a data processing apparatus

    CN105765536A

  • NDT point cloud registration algorithm and device based on GPU, and electronic equipment

    CN112837354A

  • Computer-executed feature map convolution processing method and device, and electronic equipment

    CN113468469A

  • Thread configuration method, equipment, device, storage medium and program product

    CN114860341A

  • Matrix coordinate determination method and device, electronic equipment and storage medium

    CN117520729A