Arithmetic device, processor, electronic device, arithmetic method

By setting up control modules, computing modules and data transmission modules in the processor and adjusting thread configuration parameters, the hardware cost increase caused by inter-thread voids is solved, more efficient computing and lower hardware costs are achieved, and multiple data types of computing are supported.

CN120123067BActive Publication Date: 2025-07-22MOORE THREADS TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510607481.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-07-22
Estimated Expiration
2045-05-12

AI Technical Summary

Technical Problem

In the prior art, when there are inter-thread holes in the processor, it is difficult to ensure that the program is executed correctly without increasing hardware costs.

Method used

By setting up a control module, an operation module and a data transmission module in the computing device, adjusting the configuration parameters of the thread, making them active and performing target type operations, then returning to the initial state, and outputting the calculation results to avoid storing new state information.

Benefits of technology

It realizes that in the presence of inter-thread voids, reduces hardware costs, improves processor performance, simplifies operations and improves computing efficiency, and supports multiple data types of operations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123067B_ABST
    Figure CN120123067B_ABST
Patent Text Reader

Abstract

The present disclosure relates to the field of chip technologies, and provides an arithmetic device, a processor, an electronic device, and an arithmetic method. In the device, a control module is configured to adjust configuration parameters of a first thread to make the first thread meet a first condition when the configuration parameters of the first thread are in an initial state; an arithmetic module is configured to perform an arithmetic operation of a target type on first data stored in the arithmetic device to obtain an arithmetic result corresponding to the first thread when the first thread meets the first condition; the control module is further configured to adjust the configuration parameters of the first thread to restore the configuration parameters of the first thread to the initial state; and a data transmission module is configured to output the arithmetic result corresponding to the first thread when the first thread meets a second condition. In the case of inter-thread holes, the arithmetic device according to the embodiments of the present disclosure can ensure the correct operation of a program at a lower hardware cost. When the arithmetic device is applied to a processor, the hardware cost of the processor can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of chip technology, and in particular to a computing device, a processor, an electronic device, and a computing method. Background Art

[0002] The processor executes tasks through threads. For a processor that can execute multiple tasks in parallel, it may include multiple threads. The multiple threads corresponding to the processor may be called a thread warp.

[0003] The processor uses active threads to execute tasks. However, some threads in a warp cannot always remain active. For example, conditional branches cause warp splitting, discard operations in pixel shaders, and warp size does not meet the expected hardware scale. This will cause some threads to be inactive. At this time, warps will have gaps between threads.

[0004] Ensuring that the program can still be correctly executed when there are holes between threads has become a technical problem that needs to be solved in this field. Although the processor proposed in the prior art can solve this problem, the hardware cost is relatively high. Summary of the invention

[0005] In view of this, the present disclosure proposes a computing device, a processor, an electronic device, and a computing method. In the case of a hole between threads, the computing device of the present disclosure embodiment can ensure the correct operation of the program with a lower hardware cost, and when the computing device is applied to a processor, the hardware cost of the processor can be reduced.

[0006] According to one aspect of the present disclosure, a computing device is provided, the device corresponding to a first thread in a thread bundle, the device being used to implement data computing between the first thread and other threads in the thread bundle, the device comprising a control module, a computing module, and a data transmission module, the control module being used to adjust the configuration parameters of the first thread to make the first thread active when the configuration parameters of the first thread are in an initial state; the computing module being used to perform a target type of computing on first data stored in the computing device to obtain a computing result corresponding to the first thread when the first thread is active; the control module being used to adjust the configuration parameters of the first thread to restore the configuration parameters of the first thread to an initial state after obtaining the computing result corresponding to the first thread; the data transmission module being used to output the computing result corresponding to the first thread when the first thread is active when the configuration parameters are in the initial state.

[0007] In a possible implementation, the configuration parameters of the first thread include a first identifier, a second identifier, and an active mask. When the configuration parameters are in the initial state and the result of the AND operation between the second identifier and the active mask is 1, the first thread is active; when the configuration parameters are not in the initial state and the result of the AND operation between the first identifier and the active mask is 1, the first thread is active.

[0008] In a possible implementation, the control module is specifically configured to adjust at least one of the first identifier and the active mask of the first thread by executing a control instruction.

[0009] In a possible implementation, the data transmission module is further configured to determine the first data according to the configuration parameters. Specifically, when the first thread is not active when the configuration parameters are in the initial state, the identity element of the operation corresponding to the operation of the target type is used as the first data; when the first thread is active when the configuration parameters are in the initial state, the data to be operated on by the first thread is used as the first data. The operation module is specifically configured to perform an operation of the target type on the first data after the first data determined by the data transmission module is stored in the arithmetic device.

[0010] In a possible implementation, when the target type is a reduction type, the operation of the target type is an addition operation, and the identity element is 0; when the target type is a prefix type, the operation of the target type is a multiplication operation, and the identity element is 1.

[0011] In a possible implementation, when the warp includes X threads, the operation of the target type includes log2X operations. The i-th operation realizes the data operation between the first thread and the second thread in the warp, where X > 0 and X is an integer power of 2; 0 < i ≤ log2X and i is an integer. The device further includes a first storage module and a second storage module. The first storage module is used to store the first data. The data transmission module is further configured to, during the i-th operation, obtain the first data stored in the arithmetic device corresponding to the second thread as the second data and write it into the second storage module. The operation module is specifically configured to, during the i-th operation, perform the operation of the target type on the first data stored in the first storage module and the second data stored in the second storage module, and write the result of the operation as the new first data into the first storage module. After the log2X-th operation ends, the data stored in the first storage module is used as the operation result corresponding to the first thread.

[0012] In a possible implementation, the target type is a reduction type, and the operation of the target type is an addition operation; during the i-th operation, the X threads are divided into X / 2 groups, each group includes 2 threads, each thread belongs to only one group, and the span between the 2 threads in each group is 2 i-1 ; the second thread and the first thread belong to the same group during the i-th operation, and the X computing devices corresponding to the X threads perform the i-th operation synchronously.

[0013] In a possible implementation, the target type is a prefix type, and the operation of the target type is a multiplication operation; during the i-th operation, the second thread is before the first thread, and the span between the second thread and the first thread is 2 i-1 , and the X computing devices corresponding to the X threads perform the i-th operation synchronously.

[0014] According to another aspect of the present disclosure, a processor is provided, including the computing device described above.

[0015] According to another aspect of the present disclosure, an electronic device is provided, including the processor described above.

[0016] According to another aspect of the present disclosure, an operation method is provided, which is applied to a computing device. The device corresponds to the first thread in a warp, and the device is used to implement data operations between the first thread and other threads in the warp. The device includes a control module, an operation module, and a data transmission module. The method includes: when the configuration parameters of the first thread are in the initial state, using the control module to adjust the configuration parameters of the first thread to make the first thread active; when the first thread is active, using the operation module to perform an operation of the target type on the first data stored in the computing device to obtain an operation result corresponding to the first thread; after obtaining the operation result corresponding to the first thread, using the control module to adjust the configuration parameters of the first thread to restore the configuration parameters of the first thread to the initial state; when the first thread is active when the configuration parameters are in the initial state, using the data transmission module to output the operation result corresponding to the first thread.

[0017] In a possible implementation, the configuration parameters of the first thread include a first identifier, a second identifier, and an active mask. When the second identifier and the active mask perform a logical AND operation and the result is 1, the first thread is active when the configuration parameters are in the initial state; when the configuration parameters are not in the initial state, the first thread is active when the first identifier and the active mask perform a logical AND operation and the result is 1.

[0018] In a possible implementation, adjusting the configuration parameters of the first thread by using the control module includes: adjusting at least one of the first identifier and the active mask of the first thread by executing a control instruction.

[0019] In a possible implementation, the method further includes: using the data transmission module to determine the first data according to the configuration parameters, where when the first thread is inactive when the configuration parameters are in the initial state, using the identity element of the operation corresponding to the operation of the target type as the first data; when the first thread is active when the configuration parameters are in the initial state, using the data to be operated on by the first thread as the first data; using the operation module to perform an operation of the target type on the first data stored in the operation device includes: after the first data determined by the data transmission module is stored in the operation device, using the operation module to perform an operation of the target type on the first data.

[0020] In a possible implementation, when the target type is a reduction type, the operation of the target type is an addition operation, and the identity element of the operation is 0; when the target type is a prefix type, the operation of the target type is a multiplication operation, and the identity element of the operation is 1.

[0021] In a possible implementation, when the warp includes X threads, the operation of the target type includes log2X operations. The i-th operation realizes the data operation between the first thread and the second thread in the warp, X>0, and X is an integer power of 2; 0 < i ≤ log2X, and i is an integer; the device further includes a first storage module and a second storage module, and the first storage module is used to store the first data; the method further includes: at the i-th operation, using the data transmission module to obtain the first data stored in the operation device corresponding to the second thread from the operation device corresponding to the second thread, and writing it as the second data into the second storage module; at the i-th operation, using the operation module to perform the operation of the target type on the first data stored in the first storage module and the second data stored in the second storage module, and writing the result of the operation as the new first data into the first storage module; after the log2X-th operation ends, the data stored in the first storage module is used as the operation result corresponding to the first thread.

[0022] In a possible implementation, the target type is a reduction type, and the operation of the target type is an addition operation; at the i-th operation, the X threads are divided into X / 2 groups, each group includes 2 threads, each thread belongs to only one group, and the span between the 2 threads in each group is 2 i-1; The second thread and the first thread belong to the same group during the i-th operation, and the X computing devices corresponding to the X threads perform the i-th operation synchronously.

[0023] In a possible implementation, the target type is a prefix type, and the operation of the target type is a multiplication operation; during the i-th operation, the second thread is before the first thread, and the span between the second thread and the first thread is 2 i-1 , and the X computing devices corresponding to the X threads perform the i-th operation synchronously.

[0024] According to the computing device of the embodiments of the present disclosure, by setting a control module, an arithmetic module, and a data transmission module, the control module is used to adjust the configuration parameters of the first thread to make the first thread active when the configuration parameters of the first thread are in an initial state. The arithmetic module is used to perform an operation of a target type on the first data stored in the computing device when the first thread is active, to obtain an operation result corresponding to the first thread. Therefore, this device does not need to store new state information, and can realize data operations between the first thread and other threads of the thread bundle by adjusting the configuration parameters; the control module is further used to adjust the configuration parameters of the first thread after obtaining the operation result corresponding to the first thread, so that the configuration parameters of the first thread are restored to the initial state. The data transmission module is used to output the operation result corresponding to the first thread when the first thread is active when the configuration parameters are in the initial state. Therefore, the output of the operation result can be realized by adjusting the configuration parameters. Since there is no need to store new state information, the hardware implementation cost is smaller; since there is no need to obtain state information during data operation, the operation is simpler. In summary, the computing device of the embodiments of the present disclosure can ensure the correct operation of the program in a simpler manner and with a smaller hardware cost. When the computing device is applied to a processor, the processor can enable all computing devices corresponding to the threads to participate in the operation of the target type without destroying the initial state of the thread bundle and without requiring additional space to temporarily store the state of the thread bundle, and only the operation results corresponding to specific threads that meet the conditions are output, improving the performance of the processor.

[0025] The above control module, arithmetic module, and data transmission module can be existing modules in a general-purpose graphics processor. By simply adjusting the functions of each module, the above functions can be added to each module, further reducing the hardware implementation cost.

[0026] Since the computing device is not customized for a certain type of data operation, the computing device does not limit the data types supported by the data operation, and has stronger compatibility with different processors.

[0027] The arithmetic device according to the embodiments of the present disclosure can implement inter-thread arithmetic with logarithmic complexity, effectively utilize the parallel processing ability of the processor, reduce the complexity of the arithmetic, and improve the arithmetic efficiency.

[0028] The arithmetic device according to the embodiments of the present disclosure can be at least used to implement reduction arithmetic and prefix arithmetic, and improve the arithmetic performance of the processor for reduction arithmetic and prefix arithmetic.

[0029] According to the following detailed description of the exemplary embodiments with reference to the accompanying drawings, other features and aspects of the present disclosure will become clear. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] The drawings included in and constituting a part of the specification, together with the specification, illustrate the exemplary embodiments, features, and aspects of the present disclosure, and are used to explain the principles of the present disclosure.

[0031] Figure 1 An exemplary application scenario of the arithmetic device according to the embodiments of the present disclosure is shown.

[0032] Figure 2a A schematic diagram showing the structure of the arithmetic device according to the embodiments of the present disclosure is shown.

[0033] Figure 2b An exemplary working process of the arithmetic device according to the embodiments of the present disclosure is shown.

[0034] Figure 3 A schematic diagram showing multiple arithmetic devices corresponding to a warp working in parallel according to the embodiments of the present disclosure is shown.

[0035] Figure 4 A schematic diagram showing the reduction type addition operation completed in a butterfly data exchange manner according to the embodiments of the present disclosure is shown.

[0036] Figure 5 A schematic diagram showing the prefix type multiplication operation completed in a butterfly data exchange manner according to the embodiments of the present disclosure is shown. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0037] The following will detail various exemplary embodiments, features, and aspects of the present disclosure with reference to the accompanying drawings. The same reference numerals in the drawings denote elements having the same or similar functions. Although various aspects of the embodiments are shown in the drawings, the drawings do not have to be drawn to scale unless otherwise specified.

[0038] As used herein, the terms "comprising", "including", "having", or variations thereof are open-ended and include one or more stated features, wholes, elements, steps, components, or functions, but do not exclude the existence or addition of one or more other features, wholes, elements, steps, components, functions, or groups thereof.

[0039] When an element is referred to as being "connected," "coupled," "responsive," or variations thereof, to another element, it may be directly connected, coupled, or responsive to the other element or intervening elements may be present.

[0040] Although the terms first, second, third, etc. may be used to describe various elements / operations in this article, these elements / operations should not be limited by these terms. These terms are only used to distinguish one element / operation from another element / operation. Therefore, without departing from the teachings of the present invention, the first element / operation in some embodiments may be referred to as the second element / operation in other embodiments.

[0041] The word “exemplary” is used exclusively herein to mean “serving as an example, example, or illustration.” Any embodiment described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments.

[0042] In addition, in order to better illustrate the present disclosure, numerous specific details are given in the following specific embodiments. It should be understood by those skilled in the art that the present disclosure can also be implemented without certain specific details. In some examples, methods, means, components and circuits well known to those skilled in the art are not described in detail in order to highlight the subject matter of the present disclosure.

[0043] The following is an introduction to the terms that appear in this article.

[0044] Reduction operation: An operation on a tensor (or data structure such as an array or vector) to obtain a certain result. This operation usually involves accumulating or aggregating the elements in the data structure to produce a single result.

[0045] Prefix operation: A special form of operation in which the operator is placed before the operand. Supported operations include addition, subtraction, multiplication, logical negation, etc.

[0046] Predicate: A special function that returns either 1 or 0.

[0047] As mentioned above, ensuring that the program can still execute correctly when there are holes between threads has become a technical problem that needs to be solved urgently in this field.

[0048] The processors proposed by the prior art can solve this problem by using software solutions. Taking reduction operation as an example, in the first solution, the data to be operated on by each thread is saved in the shared memory, and at the same time, active information indicating whether each thread is active is recorded in a dedicated storage space. All threads are adjusted to the active state. Each time a thread is selected, it is checked whether the state of the thread is active. The data to be operated on by the thread is read from the shared memory, and the operation instruction is used to operate on the data to be operated on by the thread and the latest operation result. The operation result is broadcast to all threads. The states of the threads are adjusted to the states indicated by the recorded active information, and the operation results at the threads are output by the active threads.

[0049] Taking reduction operation as an example, in the second solution, active information indicating whether each thread is active is recorded in a dedicated storage space. All threads are adjusted to the active state. A thread is selected, and through a dedicated data exchange instruction, the thread obtains the data to be operated on by other threads from other threads. The operation instruction is used to operate on the obtained data and the latest operation result of the thread. After the operation is completed, the operation result is broadcast to all threads. The states of the threads are adjusted to the states indicated by the recorded active information, and the operation results at the threads are output by the active threads.

[0050] Both of the above two solutions are operations with linear time complexity, that is, the number of operations is equal to the total number of threads. A larger number of operations means a larger number of instructions required to execute the program, a longer execution time of the program, more complex operations, and a waste of the parallel computing ability of the processor.

[0051] Some processors choose to set up dedicated hardware, such as a set of tree-shaped arithmetic units, to implement the reduction operation of specific data types between threads. However, such hardware cannot implement the operations between threads of other data types and other operation types, such as the prefix operation between threads, etc. To implement the operations between threads of other data types and other operation types, software solutions still need to be adopted. This increases the hardware cost of the processor.

[0052] In addition, both the software solution and the hardware solution of the prior art require additional storage space to record the active information indicating whether each thread is active, which further increases the hardware cost.

[0053] In view of this, the present disclosure proposes an arithmetic device, a processor, an electronic device, and an arithmetic method. In the case of holes between threads, the arithmetic device of the embodiments of the present disclosure can ensure the correct operation of the program at a lower hardware cost. When the arithmetic device is applied to a processor, the hardware cost of the processor can be reduced.

[0054] The arithmetic device according to the embodiments of the present disclosure can implement inter-thread arithmetic with logarithmic complexity, effectively utilize the parallel processing ability of the processor, reduce the complexity of the arithmetic, and improve the arithmetic efficiency.

[0055] The arithmetic device according to the embodiments of the present disclosure can be at least used to implement reduction arithmetic and prefix arithmetic, and improve the arithmetic performance of the processor for reduction arithmetic and prefix arithmetic.

[0056] The arithmetic device according to the embodiments of the present disclosure supports reduction arithmetic and prefix arithmetic of various data types, and has stronger compatibility with different processors.

[0057] Figure 1 An exemplary application scenario of the arithmetic device according to the embodiments of the present disclosure is shown.

[0058] As Figure 1 shown, the processor can implement parallel processing of tasks through warps. A warp can include multiple threads, and the processor can set multiple arithmetic devices, and the multiple arithmetic devices correspond to the multiple threads one by one.

[0059] The processor can be a graphics processor or other processors that support reduction arithmetic and / or prefix arithmetic. The specific type of the processor is not limited in the embodiments of the present disclosure.

[0060] The processor may further include a configuration parameter register group for storing the configuration parameters of each arithmetic device (including the active mask, the first identifier, the second identifier, etc. described below). The specific types and quantities of the configuration parameters are not limited in the embodiments of the present disclosure. The arithmetic device can change the active state of the corresponding thread by adjusting the configuration parameters.

[0061] When the processor executes a program for reduction arithmetic / prefix arithmetic, it can use the arithmetic device to complete the corresponding arithmetic. The arithmetic device can avoid the influence of thread holes on data arithmetic by adjusting the configuration parameters. And use the arithmetic device corresponding to a specific thread to output the arithmetic result corresponding to the thread. The processor can store the arithmetic results output by each arithmetic device, or can continue to perform other types of arithmetic on the arithmetic results output by the arithmetic device, or output the arithmetic results output by the arithmetic device to the outside of the processor. The use of the arithmetic results output by the arithmetic device is not limited in the embodiments of the present disclosure.

[0062] Figure 2a A schematic diagram showing the structure of the arithmetic device according to the embodiments of the present disclosure is shown. Figure 2b An exemplary working process of the arithmetic device according to the embodiments of the present disclosure is shown.

[0063] As Figure 2a and Figure 2bAs shown, in a possible implementation, the present disclosure provides an arithmetic device. The device corresponds to the first thread in a warp. The device is used to implement data arithmetic between the first thread and other threads in the warp. The device includes a control module, an arithmetic module, and a data transmission module.

[0064] The control module is used to adjust the configuration parameters of the first thread to make the first thread active when the configuration parameters of the first thread are in the initial state.

[0065] The arithmetic module is used to perform an arithmetic operation of a target type on the first data stored in the arithmetic device when the first thread is active, and obtain an arithmetic result corresponding to the first thread.

[0066] The control module is further used to adjust the configuration parameters of the first thread to restore the configuration parameters of the first thread to the initial state after obtaining the arithmetic result corresponding to the first thread.

[0067] The data transmission module is used to output the arithmetic result corresponding to the first thread when the first thread is active when the configuration parameters are in the initial state.

[0068] For example, Figure 2a The arithmetic device shown can correspond to the first thread in a warp. The first thread can be any thread in the warp. When the processor uses the first thread to execute a task, the data arithmetic between the first thread and other threads in the warp to be completed can be implemented by the arithmetic device.

[0069] Such as Figure 2b As shown, the working process of the arithmetic device can be divided into four stages. The first stage and the third stage implement the adjustment of the state of the first thread, the second stage implements data arithmetic, and the fourth stage implements the output of the arithmetic result. Such as Figure 1 In the related description, the processor can store the configuration parameters of the first thread. The configuration parameters of the first thread can indicate whether the first thread is active. In the second stage of implementing data arithmetic, the first thread must be active for the arithmetic device to participate in data arithmetic; if the first thread is not active, the arithmetic device will not be able to participate in data arithmetic. In the fourth stage of implementing the output of the arithmetic result, the first thread must be active for the arithmetic device to output the arithmetic result; if the first thread is not active, the arithmetic device does not have to output the arithmetic result.

[0070] Thread holes between threads may cause some threads in the warp to be inactive. If the arithmetic devices corresponding to the inactive threads cannot participate in the arithmetic, it will affect the arithmetic results of the arithmetic devices corresponding to the active threads. Therefore, to avoid the influence of thread holes between threads on the arithmetic between threads, before starting data arithmetic, the configuration parameters of the first thread can be adjusted to make the first thread active.

[0071] Exemplarily, the computing device may include a control module. The control module undertakes the state control function and is used to adjust the configuration parameters of the first thread to make the first thread active (the first stage) when the configuration parameters of the first thread are in the initial state. In this case, the influence of the configuration parameters on the state of the first thread is masked. Examples of adjusting the configuration parameters are given later.

[0072] Figure 3 A schematic diagram showing multiple computing devices corresponding to a warp working in parallel according to an embodiment of the present disclosure.

[0073] As Figure 3 shown, the warp may include Thread 1 - Thread 8. When the configuration parameters of the first thread are in the initial state, Thread 1 - Thread 4 are active and Thread 5 - Thread 8 are inactive. After the computing devices corresponding to each thread use the control module to adjust the configuration parameters of the thread, all threads become active.

[0074] The computing device may include an arithmetic module. The arithmetic module undertakes the data arithmetic function and is used to perform an arithmetic operation of a target type on the first data stored in the computing device when the first thread is active, to obtain an arithmetic result corresponding to the first thread (the second stage). The arithmetic operation of the target type performed by the arithmetic module includes multiple operations between the first data and the data stored in other threads. Exemplary arithmetic processes are given later.

[0075] As Figure 3 shown, since all threads are active, the computing devices corresponding to all threads use the arithmetic module to complete the data arithmetic and obtain the arithmetic results corresponding to the threads. In Figure 3 the example, the arithmetic results corresponding to Thread 1 - Thread 8 are a - h respectively.

[0076] The control module is further used to adjust the configuration parameters of the first thread to restore the configuration parameters of the first thread to the initial state after obtaining the arithmetic result corresponding to the first thread (the third stage). In this case, the adjustment of the state of the first thread by the computing device will not affect the effect of the first thread executing other tasks subsequently. Examples of adjusting the configuration parameters are given later.

[0077] As Figure 3 shown, after the control module adjusts the configuration parameters again, Thread 1 - Thread 4 are active and Thread 5 - Thread 8 are inactive, which is the same as the initial state.

[0078] The computing device may include a data transfer module. The data transfer module undertakes the data transfer function and is used to output the arithmetic result corresponding to the first thread when the first thread is active when the configuration parameters are in the initial state (the fourth stage).

[0079] As Figure 3As shown, since Thread 1 - Thread 4 are active when the configuration parameters are in the initial state, and Thread 5 - Thread 8 are not active when the configuration parameters are in the initial state, only the operation results corresponding to Thread 1 - Thread 4 need to be output. In this case, the operation results corresponding to the threads that are not active when the configuration parameters are in the initial state will not be output, ensuring that the data output situation of the operation device conforms to the state of the threads.

[0080] According to the operation device of the present disclosure embodiment, by setting a control module, an operation module, and a data transmission module, the control module is used to adjust the configuration parameters of the first thread to make the first thread active when the configuration parameters of the first thread are in the initial state. The operation module is used to perform an operation of a target type on the first data stored in the operation device when the first thread is active, to obtain an operation result corresponding to the first thread. Therefore, the device does not need to store new state information, and can achieve data operation between the first thread and other threads in the thread bundle by adjusting the configuration parameters. The control module is further used to, after obtaining the operation result corresponding to the first thread, adjust the configuration parameters of the first thread to restore the configuration parameters of the first thread to the initial state. The data transmission module is used to output the operation result corresponding to the first thread when the first thread is active when the configuration parameters are in the initial state. Therefore, the output of the operation result can be achieved by adjusting the configuration parameters. Since there is no need to store new state information, the hardware implementation cost is smaller. Since there is no need to obtain state information during data operation, the operation is simpler. In summary, the operation device of the present disclosure embodiment can ensure the correct operation of the program in a simpler manner and with a smaller hardware cost. When the operation device is applied to a processor, the processor can enable all operation devices corresponding to the threads to participate in the operation of the target type without destroying the initial state of the thread bundle and without requiring additional space to temporarily store the state of the thread bundle, and only the operation results corresponding to specific threads that meet the conditions are output, improving the performance of the processor.

[0081] The above control module, operation module, and data transmission module can be existing modules in a general - purpose graphics processor. By simply adjusting the functions of each module, the above functions can be added to each module, further reducing the hardware implementation cost. The operation device can be implemented at an extremely low cost on most general - purpose graphics processors and single - instruction - multiple - data stream processors, and has no impact on the original program and functions, with strong hardware versatility.

[0082] Since the operation device is not customized for a certain type of data operation, the operation device does not limit the data types supported by the data operation, and has stronger compatibility with different processors. As long as the calculation types and data types supported by the processor, including user - defined data structures, etc., the operation device of the present disclosure embodiment can provide operations without additional hardware support.

[0083] The following describes an example of configuration parameters, and how to determine whether the first thread is active based on the configuration parameters and how to adjust the configuration parameters in combination with the example of the configuration parameters.

[0084] In a possible implementation, the configuration parameters of the first thread include a first identifier, a second identifier, and an active mask. When the configuration parameters are in the initial state, if the logical AND operation result of the second identifier and the active mask is 1, the first thread is active; when the configuration parameters are not in the initial state, if the logical AND operation result of the first identifier and the active mask is 1, the first thread is active.

[0085] In a possible implementation, the control module is specifically configured to adjust at least one of the first identifier and the active mask of the first thread by executing a control instruction.

[0086] For example, the configuration parameters can include three types, namely a first identifier, a second identifier, and an active mask. Correspondingly, the configuration parameter register group can include a first register, a second register, and a mask register corresponding to the first thread. The first register stores the first identifier of the first thread, the second register stores the second identifier of the first thread, and the mask register stores the active mask of the first thread. The embodiments of the present disclosure do not limit the specific storage manner of the configuration parameters in the configuration parameter register group.

[0087] Among them, the first identifier can be the output of the first predicate for the first thread, and the second identifier can be the output of the second predicate for the first thread. In one example, the first predicate can be: whether to allow the thread to participate in data operations. If the thread is allowed to participate in data operations, the first identifier can be 1; if the thread is not allowed to participate in data operations, the first identifier can be 0. The second predicate can be: whether to allow the thread to output the operation result. If the thread is allowed to output the operation result, the second identifier can be 1; if the thread is not allowed to output the operation result, the second identifier can be 0. The active mask can indicate whether the thread can enter the active state. When the thread can enter the active state, the active mask can be 1; when the thread cannot enter the active state, the active mask can be 0.

[0088] It should be understood that only when the first thread can enter the active state and is allowed to participate in data operations, can it be active in the stage of implementing data operations, that is, when the configuration parameters are not in the initial state, if the logical AND operation result of the first identifier and the active mask is 1, the first thread is active. Only when the first thread can enter the active state and is allowed to output the operation result, can it be active in the stage of implementing the output of the operation result, that is, when the configuration parameters are in the initial state, if the logical AND operation result of the second identifier and the active mask is 1, the first thread is active.

[0089] When the configuration parameters of the first thread are in the initial state, the result of the AND operation between the first identifier and the active mask may be 0 or 1. To make the first thread active, the control module can adjust at least one of the first identifier and the active mask so that the result of the AND operation between the first identifier and the active mask is 1. Correspondingly, to restore the configuration parameters of the first thread to the initial state, the control module can adjust at least one of the first identifier and the active mask to restore the first identifier and the active mask to the initial state.

[0090] For example, assume that the first identifier of the first thread is 0 in the initial state and the active mask is 0 in the initial state. Then, adjusting both the first identifier and the active mask of the first thread to 1 can make the first thread active.

[0091] It should be understood that if the first identifier and the active mask of the first thread are already 1 in the initial state, the first thread is already active, and there is no need to adjust the configuration parameters of the first thread.

[0092] Exemplarily, the control module can adjust the configuration parameters to make the first thread active through control instruction C0. The control module can adjust the configuration parameters to restore the configuration parameters of the first thread to the initial state through control instruction C1. Control instruction C1 can cancel the execution effect of control instruction C0.

[0093] In this case, by simply adding the execution logic of control instruction C0 and control instruction C1 to the control module, the arithmetic device can be made to have the function of shielding the influence of the hole between threads on the data operation of the thread, with less modification to the arithmetic device.

[0094] Those skilled in the art should understand that the setting and adjustment methods of the configuration parameters should not be limited to the above examples. For example, the configuration parameters can also be set to include a third identifier, which indicates that the first thread is active when the third identifier is 1 and indicates that the first thread is inactive when the third identifier is 0. As long as the control module can make the first thread active in the second stage and make the state of the first thread in the fourth stage the same as before entering the first stage by adjusting the configuration parameters.

[0095] In a possible implementation, the data transmission module is further configured to determine the first data according to the configuration parameters, where, when the first thread is inactive when the configuration parameters are in the initial state, the identity element of the operation corresponding to the target type of operation is used as the first data; when the first thread is active when the configuration parameters are in the initial state, the data to be operated on by the first thread is used as the first data;

[0096] The arithmetic module is specifically configured to perform the target type of operation on the first data after the first data determined by the data transmission module is stored in the arithmetic device.

[0097] For example, each thread has corresponding data to be computed. For a thread that is inactive when the configuration parameter is in the initial state, the data to be computed for that thread does not need to participate in the data computation of the target type among the threads. Therefore, whether the data to be computed can be used as the first data can be determined according to the configuration parameter.

[0098] Among them, when the first thread is inactive when the configuration parameter is in the initial state (for example, the AND operation result of the second identifier and the active mask is 0), the data transfer module can use the identity element of the operation corresponding to the operation of the target type as the first data; when the first thread is active when the configuration parameter is in the initial state (for example, the AND operation result of the second identifier and the active mask is 1), the data transfer module can use the data to be computed of the first thread as the first data. In this way, the accuracy of the operation result can be guaranteed.

[0099] In a possible implementation, when the target type is a reduction type, the operation of the target type is an addition operation, and the identity element of the operation corresponding to the operation of the target type is 0. Adding any data to 0 does not change the value of the data, thereby ensuring the accuracy of the operation result.

[0100] In a possible implementation, when the target type is a prefix type, the operation of the target type is a multiplication operation, and the identity element of the operation corresponding to the operation of the target type is 1. Multiplying any data by 1 does not change the value of the data, thereby ensuring the accuracy of the operation result.

[0101] Those skilled in the art should understand that the target type can also be other types, and the operation of the target type can also be other types. The embodiments of the present disclosure do not limit the target type and the type of the target operation. The embodiments of the present disclosure do not limit the specific value of the identity element corresponding to the operation of the target type, as long as the value of the data will not be changed after performing the operation of the target type on any data and the identity element.

[0102] The data transfer module can store the determined first data in the computing device. After that, during the stage of implementing data computation, when the first thread is active, the computing module performs the operation of the target type on the first data.

[0103] The computing device may include a first storage module for storing the first data. The storage of the first data can be completed by the data transfer module. Exemplarily, the data transfer module can complete the storage of the first data by executing a data transfer instruction.

[0104] In one example, when the data transfer module stores the first data, it can first determine the first data and store the first data in the first storage module by executing a data transfer instruction M0 once. In this way, the number of instructions can be further reduced.

[0105] In another example, when the first thread is active when the configuration parameter is in the initial state, the data transfer module may first execute a data transfer instruction M1 to write the operation identity element corresponding to the operation of the target type into the first storage module, and then when the first thread is active when the configuration parameter is not in the initial state, execute a data transfer instruction M2 again to write the data to be operated into the first storage module. In this way, the logic complexity of the data transfer module can be reduced.

[0106] Those skilled in the art should understand that as long as the first data stored in the first storage module is correct when the operation module starts to operate, the specific manner in which the data transfer module stores the first data in the first storage module is not limited in the embodiments of the present disclosure.

[0107] The following describes an exemplary manner in which the arithmetic device performs an operation of the target type on the first data.

[0108] In a possible implementation, when the thread bundle includes X threads, the operation of the target type includes log2X operations. The i-th operation implements the data operation between the first thread and the second thread in the thread bundle, where X>0 and X is an integer power of 2; 0 < i ≤ log2X and i is an integer.

[0109] The device further includes a first storage module and a second storage module. The first storage module is used to store the first data;

[0110] The data transfer module is further configured to, during the i-th operation, obtain the first data stored in the corresponding arithmetic device of the second thread as the second data and write it into the second storage module;

[0111] The arithmetic module is specifically configured to, during the i-th operation, perform an operation of the target type on the first data stored in the first storage module and the second data stored in the second storage module, and write the result of the operation as the new first data into the first storage module;

[0112] After the log2X-th operation ends, the data stored in the first storage module is used as the operation result corresponding to the first thread.

[0113] For example, the arithmetic module can complete data operations in a butterfly data exchange manner. When the thread bundle includes X threads, the operation of the target type includes log2X operations. The i-th operation implements the data operation between the first thread and the second thread in the thread bundle.

[0114] The arithmetic device may include a second storage module for storing second data. The second data is the first data stored in the first storage module included in the arithmetic device corresponding to the second thread during the i-th operation. The second data is obtained by the data transmission module of the arithmetic device executing a data exchange instruction and written into the second storage module. Accordingly, during the i-th operation, the arithmetic device corresponding to the second thread also writes the first data stored in the arithmetic device corresponding to the first thread into the second storage module included therein, thereby realizing butterfly data exchange.

[0115] In this case, specifically, during the i-th operation, the arithmetic module is configured to perform an operation of a target type on the first data stored in the first storage module and the second data stored in the second storage module in the arithmetic device to which it belongs. The data transmission module executes a data transfer instruction and writes the result of the operation as new first data into the first storage module of the arithmetic device to which it belongs.

[0116] Accordingly, during the i-th operation, the arithmetic device corresponding to the second thread also performs an operation of a target type on the first data stored in the first storage module and the second data stored in the second storage module in the arithmetic device to which it belongs. The data transmission module executes a data transfer instruction and writes the result of the operation as new first data into the first storage module of the arithmetic device to which it belongs.

[0117] After the log2X-th operation ends, in the arithmetic device corresponding to the first thread, the data stored in the first storage module can be used as the arithmetic result corresponding to the first thread and output by the data transmission module in the arithmetic device by executing a data transfer instruction.

[0118] Those skilled in the art should understand that the first storage module and the second storage module may also be integrated into one storage module, and the first data and the second data are respectively stored in different storage areas of the storage module. As long as the first data and the second data are respectively stored in the arithmetic device, the specific storage manner of the first data and the second data in the embodiments of the present disclosure is not limited.

[0119] Those skilled in the art should understand that the data transmission module may also be further refined into a data transfer unit and a data exchange unit. The data transfer unit executes the data transfer instruction described above, and the data exchange unit executes the data exchange instruction described above. The specific structure of the data transmission module in the embodiments of the present disclosure is not limited.

[0120] The arithmetic device of the embodiments of the present disclosure supports performing a reduction type of addition operation in a butterfly data exchange manner. Figure 4 A schematic diagram showing the completion of a reduction type of addition operation in a butterfly data exchange manner according to an embodiment of the present disclosure is shown.

[0121] In a possible implementation, the target type is a specification type, and the operation of the target type is an addition operation. During the i-th operation, X threads are divided into X / 2 groups, each group includes 2 threads, each thread belongs to only one group, and the span between the 2 threads in each group is 2 i-1 ; The second thread and the first thread belong to the same group during the i-th operation, and the X computing devices corresponding to the X threads perform the i-th operation synchronously.

[0122] As Figure 4 shown, before the computing devices corresponding to Thread 1 - Thread 8 start operating, the first data stored are 0, 1, 0, 3, 4, 0, 6, 0 respectively. Since the total number of threads X = 8, the number of operations is log28 = 3.

[0123] During the 1st operation, the span between the 2 threads in the group is 2 0 = 1. Therefore, Thread 1 and Thread 2 belong to the same group, Thread 3 and Thread 4 belong to the same group, Thread 5 and Thread 6 belong to the same group, and Thread 7 and Thread 8 belong to the same group.

[0124] Taking Thread 1 as the first thread as an example, during the 1st operation, with Thread 2 as the second thread, in the computing device corresponding to Thread 1, the data transfer module obtains (transfers) the first data stored in the computing device corresponding to Thread 2. Therefore, the data obtained is 1, which is written as the second data into the second storage module. At this time, in the computing device corresponding to Thread 1, the first data stored in the first storage module is 0. The operation module performs an addition operation on the first data 0 stored in the first storage module and the second data 1 stored in the second storage module, and writes the result 1 of the operation as the new first data into the first storage module.

[0125] When the computing device corresponding to Thread 1 performs the 1st operation, the computing devices corresponding to Thread 2 - Thread 8 also perform the 1st operation respectively. The details of the 1st operation of the computing devices corresponding to Thread 2 - Thread 8 will not be elaborated here. After the computing devices corresponding to Thread 1 - Thread 8 complete the 1st operation, the first data stored in the first storage module are 1, 1, 3, 3, 4, 4, 6, 6 respectively.

[0126] During the 2nd operation, the span between the 2 threads in the group is 2 1 = 2. Therefore, Thread 1 and Thread 3 belong to the same group, Thread 2 and Thread 4 belong to the same group, Thread 5 and Thread 7 belong to the same group, and Thread 6 and Thread 8 belong to the same group.

[0127] Still taking Thread 1 as the first thread as an example, during the second operation, Thread 3 is the second thread. In the computing device corresponding to Thread 1, the data transmission module obtains (transports) the first data stored in the computing device corresponding to Thread 3. Therefore, the obtained data is 3, which is written as the second data into the second storage module. At this time, in the computing device corresponding to Thread 1, the first data stored in the first storage module is 1. The computing module performs an addition operation on the first data 1 stored in the first storage module and the second data 3 stored in the second storage module, and writes the operation result 4 as the new first data into the first storage module.

[0128] When the computing device corresponding to Thread 1 performs the second operation, the computing devices corresponding to Threads 2 - 8 also perform the second operation respectively. Details of the second operation of the computing devices corresponding to Threads 2 - 8 will not be elaborated here. After the computing devices corresponding to Threads 1 - 8 complete the second operation, the first data stored in the first storage module are 4, 4, 4, 4, 10, 10, 10, 10 respectively.

[0129] During the third operation, the span between two threads in a group is 2 2 = 4. Therefore, Thread 1 and Thread 5 belong to the same group, Thread 2 and Thread 6 belong to the same group, Thread 3 and Thread 7 belong to the same group, and Thread 4 and Thread 8 belong to the same group.

[0130] Still taking Thread 1 as the first thread as an example, during the third operation, Thread 5 is the second thread. In the computing device corresponding to Thread 1, the data transmission module obtains (transports) the first data stored in the computing device corresponding to Thread 5. Therefore, the obtained data is 10, which is written as the second data into the second storage module. At this time, in the computing device corresponding to Thread 1, the first data stored in the first storage module is 4. The computing module performs an addition operation on the first data 4 stored in the first storage module and the second data 10 stored in the second storage module, and writes the operation result 14 as the new first data into the first storage module.

[0131] When the computing device corresponding to Thread 1 performs the third operation, the computing devices corresponding to Threads 2 - 8 also perform the third operation respectively. Details of the third operation of the computing devices corresponding to Threads 2 - 8 will not be elaborated here. After the computing devices corresponding to Threads 1 - 8 complete the third operation, the first data stored in the first storage module are 14, 14, 14, 14, 14, 14, 14, 14 respectively.

[0132] In this way, the reduction - type addition operation can be correctly implemented with logarithmic complexity.

[0133] The computing device of the present disclosure embodiment supports completing the prefix - type multiplication operation in a butterfly data exchange manner.Figure 5 A schematic diagram showing the multiplication operation of the prefix type in the form of butterfly data exchange according to an embodiment of the present disclosure.

[0134] In a possible implementation, the target type is the prefix type, and the operation of the target type is multiplication operation;

[0135] In the i-th operation, the second thread is before the first thread, and the span between the second thread and the first thread is 2 i-1 , and the X computing devices corresponding to X threads perform the i-th operation synchronously.

[0136] As Figure 5 shown, before the computing devices corresponding to Thread 1 - Thread 8 start the operation, the first data stored are 1, 1, 3, 1, 1, 6, 1, 8 respectively. Since the total number of threads X = 8, the number of operations is log28 = 3.

[0137] In the 1st operation, the span between the second thread and the first thread is 2 0 = 1. Since the second thread is before the first thread, when Thread 1 is the first thread, there is no corresponding second thread. When Thread 2 is the first thread, the corresponding second thread is Thread 1. When Thread 3 is the first thread, the corresponding second thread is Thread 2. And so on, when Thread 8 is the first thread, the corresponding second thread is Thread 7.

[0138] Taking Thread 8 as the first thread as an example, in the 1st operation, Thread 7 is the second thread. In the computing device corresponding to Thread 8, the data transmission module obtains (transports) the first data stored in the computing device corresponding to Thread 7, so the obtained data is 1, which is written as the second data into the second storage module. At this time, in the computing device corresponding to Thread 8, the first data stored in the first storage module is 8. The computing module performs a multiplication operation on the first data 8 stored in the first storage module and the second data 1 stored in the second storage module, and writes the operation result 8 as the new first data into the first storage module.

[0139] When the computing device corresponding to Thread 8 performs the 1st operation, the computing devices corresponding to Threads 2 - 7 also perform the 1st operation respectively. The details of the 1st operation of the computing devices corresponding to Threads 2 - 7 will not be elaborated here. After the computing devices corresponding to Threads 2 - 8 complete the 1st operation, the first data stored in the first storage module are 1, 1, 3, 3, 1, 6, 6, 8 respectively.

[0140] In the 2nd operation, the span between the second thread and the first thread is 2 1= 2. Since the second thread is before the first thread, when Thread 1 and Thread 2 are the first threads, there is no corresponding second thread. When Thread 3 is the first thread, the corresponding second thread is Thread 1. When Thread 4 is the first thread, the corresponding second thread is Thread 2. And so on. When Thread 8 is the first thread, the corresponding second thread is Thread 6.

[0141] Still taking Thread 8 as the first thread as an example, in the second operation, Thread 6 is the second thread. In the computing device corresponding to Thread 8, the data transfer module obtains (transports) the first data stored in the computing device corresponding to Thread 6. Therefore, the obtained data is 6, which is written as the second data into the second storage module. At this time, in the computing device corresponding to Thread 8, the first data stored in the first storage module is 8. The computing module performs a multiplication operation on the first data 8 stored in the first storage module and the second data 6 stored in the second storage module, and writes the operation result 48 as the new first data into the first storage module.

[0142] When the second operation is performed on the computing device corresponding to Thread 8, the computing devices corresponding to Threads 3 - 7 also perform the second operation respectively. The details of the second operation on the computing devices corresponding to Threads 3 - 7 will not be elaborated here. After the computing devices corresponding to Threads 3 - 8 complete the second operation, the first data stored in the first storage module are 1, 1, 3, 3, 3, 18, 6, 48 respectively.

[0143] In the third operation, the span between the second thread and the first thread is 2 2 = 4. Since the second thread is before the first thread, when Threads 1 - 4 are the first threads, there is no corresponding second thread. When Thread 5 is the first thread, the corresponding second thread is Thread 1. When Thread 6 is the first thread, the corresponding second thread is Thread 2. And so on. When Thread 8 is the first thread, the corresponding second thread is Thread 4.

[0144] Still taking Thread 8 as the first thread as an example, in the third operation, Thread 4 is the second thread. In the computing device corresponding to Thread 8, the data transfer module obtains (transports) the first data stored in the computing device corresponding to Thread 4. Therefore, the obtained data is 3, which is written as the second data into the second storage module. At this time, in the computing device corresponding to Thread 8, the first data stored in the first storage module is 48. The computing module performs a multiplication operation on the first data 48 stored in the first storage module and the second data 3 stored in the second storage module, and writes the operation result 144 as the new first data into the first storage module.

[0145] When the computing device corresponding to thread 8 performs the third computation, the computing devices corresponding to threads 5 - 7 also perform the third computation respectively. The details of the third computation of the computing devices corresponding to threads 5 - 7 will not be elaborated here. After the computing devices corresponding to threads 5 - 8 complete the third computation, the first data stored in the first storage module are 1, 1, 3, 3, 3, 18, 18, and 144 respectively.

[0146] Those skilled in the art should understand that during the i-th computation, if the first thread has no corresponding second thread, in the computing device corresponding to the first thread, the data transfer module may not perform data transfer and may not write data to the second storage module, and the computing module may not perform multiplication. Alternatively, the data transfer module may directly write the operation unit 1 as the second data to the second storage module, and the computing module performs computations on the data stored in the first storage module and the second storage module. The embodiments of the present disclosure do not limit whether the data transfer module in the computing device corresponding to the first thread still writes data to the second storage module and whether the computing module still completes the computation when the first thread has no corresponding second thread.

[0147] In this way, the prefix type of multiplication can be correctly implemented with logarithmic complexity.

[0148] The embodiments of the present disclosure also propose a processor including the above-described computing device. A schematic diagram of the structure of the processor can be seen in Figure 1 . The processor can be a graphics processor or other processors with inter-thread data computation requirements. The embodiments of the present disclosure do not limit the specific type of the processor.

[0149] The embodiments of the present disclosure also propose an electronic device including the above-described processor. The electronic device can be a server or a terminal device. The embodiments of the present disclosure do not limit the specific type of the electronic device.

[0150] The embodiments of the present disclosure also propose an operation method. A schematic diagram of the process of the method can be seen in Figure 2b .

[0151] In a possible implementation, the method is applied to an arithmetic device corresponding to the first thread in a warp. The device is used to implement data arithmetic between the first thread and other threads in the warp. The device includes a control module, an arithmetic module, and a data transmission module. The method includes: when the configuration parameters of the first thread are in the initial state, using the control module to adjust the configuration parameters of the first thread to make the first thread active; when the first thread is active, using the arithmetic module to perform an arithmetic operation of a target type on the first data stored in the arithmetic device to obtain an arithmetic result corresponding to the first thread; after obtaining the arithmetic result corresponding to the first thread, using the control module to adjust the configuration parameters of the first thread to restore the configuration parameters of the first thread to the initial state; when the configuration parameters are in the initial state and active, using the data transmission module to output the arithmetic result corresponding to the first thread.

[0152] In a possible implementation, the configuration parameters of the first thread include a first identifier, a second identifier, and an active mask. When the configuration parameters are in the initial state and the logical AND operation result of the second identifier and the active mask is 1, the first thread is active; when the configuration parameters are not in the initial state and the logical AND operation result of the first identifier and the active mask is 1, the first thread is active.

[0153] In a possible implementation, using the control module to adjust the configuration parameters of the first thread includes: by executing a control instruction, adjusting at least one of the first identifier and the active mask of the first thread.

[0154] In a possible implementation, the method further includes:

[0155] Using the data transmission module to determine the first data according to the configuration parameters, where, when the first thread is not active when the configuration parameters are in the initial state, using the arithmetic unit element corresponding to the arithmetic operation of the target type as the first data; when the first thread is active when the configuration parameters are in the initial state, using the data to be arithmetically operated by the first thread as the first data; using the arithmetic module to perform an arithmetic operation of a target type on the first data stored in the arithmetic device includes: after the first data determined by the data transmission module is stored in the arithmetic device, using the arithmetic module to perform an arithmetic operation of a target type on the first data.

[0156] In a possible implementation, when the target type is a reduction type, the operation of the target type is an addition operation, and the identity element of the operation is 0; when the target type is a prefix type, the operation of the target type is a multiplication operation, and the identity element of the operation is 1.

[0157] In a possible implementation, when the warp includes X threads, the operation of the target type includes log2X operations. The i-th operation implements the data operation between the first thread and the second thread in the warp, where X > 0 and X is an integer power of 2; 0 < i ≤ log2X and i is an integer. The apparatus further includes a first storage module and a second storage module, and the first storage module is used to store the first data. The method further includes: in the i-th operation, using the data transmission module, obtaining the first data stored in the corresponding arithmetic device of the second thread as the second data and writing it into the second storage module; in the i-th operation, using the arithmetic module, performing the operation of the target type on the first data stored in the first storage module and the second data stored in the second storage module, and writing the result of the operation as the new first data into the first storage module; after the log2X-th operation ends, the data stored in the first storage module is used as the arithmetic result corresponding to the first thread.

[0158] In a possible implementation, the target type is a reduction type, and the operation of the target type is an addition operation; in the i-th operation, the X threads are divided into X / 2 groups, each group includes 2 threads, each thread belongs to only one group, and the span between the 2 threads in each group is 2 i-1 ; the second thread and the first thread belong to the same group in the i-th operation, and the X arithmetic devices corresponding to the X threads perform the i-th operation synchronously.

[0159] In a possible implementation, the target type is a prefix type, and the operation of the target type is a multiplication operation; in the i-th operation, the second thread is located before the first thread, and the span between the second thread and the first thread is 2 i-1 ; the X arithmetic devices corresponding to the X threads perform the i-th operation synchronously.

[0160] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0161] The embodiments of the present disclosure have been described above. The above description is exemplary and not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art in the technical field without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or the improvement of technologies in the market, or to enable other ordinary skill in the art in the technical field to understand the embodiments disclosed herein.

Claims

1. An operation device, characterized in that, The device corresponds to the first thread in a warp, and the device is used to implement data operations between the first thread and other threads in the warp. The device includes a control module, an arithmetic module, and a data transmission module. The control module is used to adjust the configuration parameters of the first thread to make the first thread active when the configuration parameters of the first thread are in an initial state. The arithmetic module is used to perform an operation of a target type on the first data stored in the arithmetic device to obtain an operation result corresponding to the first thread when the first thread is active. The control module is further used to adjust the configuration parameters of the first thread to restore the configuration parameters of the first thread to the initial state after obtaining the operation result corresponding to the first thread. The data transmission module is used to output the operation result corresponding to the first thread when the first thread is active when the configuration parameters are in the initial state.

2. The device according to claim 1, characterized in that, The configuration parameters of the first thread include a first identifier, a second identifier, and an active mask. When the AND operation result of the second identifier and the active mask is 1 in the case where the configuration parameters are in the initial state, the first thread is active. When the configuration parameters are not in the initial state, the first thread is active when the AND operation result of the first identifier and the active mask is 1.

3. The device according to claim 2, characterized in that The control module is specifically used for adjusting at least one of the first identifier and the active mask of the first thread by executing a control instruction.

4. The device according to claim 1, wherein the data transmission module is further used to determine the first data according to the configuration parameters. Among them, when the first thread is not active when the configuration parameters are in the initial state, the operation identity element corresponding to the operation of the target type is used as the first data; when the first thread is active when the configuration parameters are in the initial state, the data to be operated on by the first thread is used as the first data. The arithmetic module is specifically used for performing an operation of a target type on the first data after the first data determined by the data transmission module is stored in the arithmetic device.

5. The device according to claim 4, wherein when the target type is a reduction type, the operation of the target type is an addition operation, and the operation identity element is 0. when the target type is a prefix type, the operation of the target type is a multiplication operation, and the operation identity element is 1.

6. The device according to any one of claims 1-5, characterized in that When the warp includes X threads, the operation of the target type includes log2X operations. The i-th operation implements data operations between the first thread and the second thread in the warp, X > 0, X is an integer power of 2; 0 < i ≤ log2X, and i is an integer. The device further includes a first storage module and a second storage module. The first storage module is used to store the first data. The data transmission module is further used to obtain the first data stored in the arithmetic device corresponding to the second thread as the second data and write it into the second storage module at the i-th operation. Specifically, in the i-th operation, the operation module performs an operation of the target type on the first data stored in the first storage module and the second data stored in the second storage module, and writes the result of the operation as the new first data into the first storage module; After the log2X-th operation ends, the data stored in the first storage module is used as the operation result corresponding to the first thread.

7. The device according to claim 6, characterized in that, The target type is a reduction type, and the operation of the target type is an addition operation; During the i-th operation, the X threads are divided into X / 2 groups, each group includes 2 threads, each thread belongs to only one group, and the span between the 2 threads in each group is 2 i-1 ; In the i-th operation, the second thread and the first thread belong to the same group, and the X computing devices corresponding to the X threads perform the i-th operation synchronously.

8. The device according to claim 6, characterized in that, The target type is a prefix type, and the operation of the target type is a multiplication operation; At the i-th operation, the second thread is before the first thread, and the span between the second thread and the first thread is 2 i-1 , and the X computing devices corresponding to the X threads perform the i-th operation synchronously.

9. A processor, characterized in that, Including the computing device according to any one of claims 1-8.

10. An electronic device, characterized in that, Including the processor according to claim 9.

11. An operation method, characterized in that, Applied to a computing device, the device corresponds to the first thread in a warp, the device is used to implement data operations between the first thread and other threads in the warp, the device includes a control module, an operation module, and a data transmission module, and the method includes: When the configuration parameter of the first thread is in the initial state, use the control module to adjust the configuration parameter of the first thread to make the first thread active; When the first thread is active, use the operation module to perform an operation of the target type on the first data stored in the computing device to obtain an operation result corresponding to the first thread; After obtaining the operation result corresponding to the first thread, use the control module to adjust the configuration parameter of the first thread to restore the configuration parameter of the first thread to the initial state; When the first thread is active when the configuration parameter is in the initial state, use the data transmission module to output the operation result corresponding to the first thread.

Citation Information

Patent Citations

  • Method and apparatus for requesting processing order-preserving in distributed storage protocol

    CN107277128A

  • Data processor proceeding of accelerated synchronization between central processing unit and graphics processing unit

    KR1020180099420A