Communication error detection method, device, equipment and medium
By using multi-threading and optimized reduction algorithms in a general graphics processor for data processing and detecting the consistency of memory addresses, the DUE problem caused by data access address errors is solved, and the reliability of the system is improved.
Patent Information
- Application Number
- CN202510192360.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2045-02-21
AI Technical Summary
During data processing, detectable unrecoverable type errors (DUEs) caused by data access address errors are common problems. The prior art is difficult to detect and avoid these errors in a timely manner, resulting in a decrease in system reliability.
By using several threads in a general graphics processor to obtain the pending data blocks based on their own memory access addresses, and multiple iterative processing is performed using an optimized reduction algorithm. In each iteration process, the receiving thread calculates the consistency of the target memory address to detect whether there is a communication error.
It realizes timely discovering communication errors during data processing, avoids DUEs caused by data access address errors, and improves the reliability of the general graphics processor.
Smart Images

Figure CN119697069B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of communication technology, and in particular to a communication error detection method, device, equipment and medium. Background Art
[0002] GPGPU (General-purpose computing on graphics processing units) is widely used in the fields of neural networks and high-performance computing due to its highly parallel computing capabilities, and is gradually expanding to application scenarios with higher precision, reliability and timeliness, such as autonomous driving.
[0003] Soft errors are one of the important factors that threaten system reliability. Although soft errors are not permanent faults, and the written erroneous values can be corrected by rewriting the affected values, they may propagate when the written erroneous values have been read and used. Such propagation may lead to incorrect calculation results or even program interruption; among them, the detectable unrecoverable type error (DUE, Detected Uncorrectable Error) can be detected, but there is no way to recover, so it is likely to cause the program to stop or become unresponsive, and the DUE caused by data access address errors during data processing is a relatively common type of error. Based on this, how to detect communication errors in time during data processing to avoid DUE caused by data access address errors is a problem that needs to be optimized. Summary of the invention
[0004] In view of this, the purpose of the present invention is to provide a communication error detection method, device, equipment and medium, which can detect communication errors in a timely manner during data processing to avoid DUE caused by data access address errors and improve the reliability of general graphics processors. The specific scheme is as follows:
[0005] In a first aspect, the present application provides a communication error detection method, which is applied to a general-purpose graphics processor, comprising:
[0006] Obtaining corresponding to-be-processed data blocks from the to-be-processed array by a plurality of threads based on their own memory access addresses to the to-be-processed array; each thread is configured with a corresponding thread identifier, and the thread identifiers of the plurality of threads are identifiers determined based on continuous numerical values; and the memory access addresses of the plurality of threads are memory addresses increasing at equal intervals;
[0007] The optimized reduction algorithm is used by the several threads to perform multiple rounds of iterative processing on the corresponding data blocks to be processed to obtain the final processing result; the optimized reduction algorithm is an algorithm obtained by optimizing the parallel reduction algorithm based on the end reversal method; the end reversal method is to control each sending thread in the threads participating in the current round of iterative processing to transmit its own current round of iterative processing result to the corresponding receiving thread; the sum of the thread identifiers of the sending thread and the corresponding receiving thread is equal to the target value; the thread identifier of each sending thread is greater than or less than the corresponding receiving thread; the target value is the sum of the maximum thread identifier and the minimum thread identifier among the thread identifiers of the threads participating in the current round of iterative processing;
[0008] During each round of iterative processing, each receiving thread is controlled to add its own memory access address and the corresponding memory access address of the sending thread to calculate the target memory address. Based on the consistency of the target memory address calculated by each receiving thread, it is detected whether there is a communication error in this round of iterative processing.
[0009] Optionally, before the acquiring corresponding data blocks to be processed from the array to be processed by the plurality of threads based on the memory access addresses of the array to be processed by the threads themselves, the method further includes:
[0010] Obtaining a computing task for an array to be processed, and determining a number of threads required for the computing task based on a data type and a data size of the array to be processed; the total number of threads of the number of threads is a power of two;
[0011] A corresponding thread identifier is configured for each thread based on a continuous numerical value starting from zero, and a memory access address of each thread to the array to be processed is determined according to a predetermined address spacing, a starting address of the array to be processed and a thread identifier of each thread; wherein the address spacing is a spacing predetermined based on a data type and a data size of the array to be processed.
[0012] Optionally, determining a number of threads required for the computing task based on the data type and data size of the array to be processed includes:
[0013] Determine the total number of threads based on the data type and data size of the array to be processed, and determine the number of thread warps required for the computing task according to the total number of threads and the number of threads corresponding to one thread warp;
[0014] A corresponding number of continuous thread warps are determined from the thread block using the number of thread warps, and a corresponding number of continuous threads are determined from the continuous thread warps using the total number of threads, so as to obtain a number of threads required for the computing task.
[0015] Optionally, the process of determining each sending thread and the corresponding receiving thread includes:
[0016] In each round of iterative processing, determining the current round thread number of threads participating in the current round of iterative processing, and converting a value obtained by subtracting one from the current round thread number into a first binary number;
[0017] Determine a preset number of threads with the largest or smallest thread identifiers from the threads participating in the current round of iterative processing to obtain each of the sending threads; the preset number is half of the number of threads in the current round;
[0018] The thread identifier of the sending thread is converted into a second binary number, and an XOR operation is performed on the first binary number and the second binary number to convert the XOR result into a decimal number to obtain a target thread identifier, and then the thread corresponding to the target thread identifier is determined as the receiving thread corresponding to the sending thread.
[0019] Optionally, the performing multiple rounds of iterative processing on the corresponding data blocks to be processed by the plurality of threads using an optimized reduction algorithm to obtain a final processing result includes:
[0020] Starting multiple rounds of iterative processing of the corresponding data blocks to be processed by the plurality of threads, taking the first round of iterative processing as the current round of iterative processing, and taking the plurality of threads as threads participating in the first round of iterative processing;
[0021] In the current round of iterative processing, the threads participating in the current round of iterative processing process process their respective data blocks to be processed to obtain corresponding results of the current round of iterative processing, and control the sending threads in the threads participating in the current round of iterative processing to transmit their own results of the current round of iterative processing to the corresponding receiving threads based on the end reversal mode, and after the transmission is completed, use the receiving threads as threads participating in the next round of iterative processing, and then enter the next round of iterative processing;
[0022] The next round of iterative processing process is used as a new current round of iterative processing process, and in the current round of iterative processing process, the threads participating in this round of iterative processing determine the results of this round of iterative processing based on their own previous round of iterative processing results and the received previous round of iterative processing results, and then jump again to the step of controlling each sending thread in the threads participating in this round of iterative processing based on the end reversal method to transmit its own current round of iterative processing results to the corresponding receiving thread, until there is only one thread participating in this round of iterative processing, and the latest result of this round of iterative processing is used as the final processing result.
[0023] Optionally, in each round of iterative processing, determining the number of threads in the current round of iterative processing includes:
[0024] In each round of iterative processing, the current thread mask is updated based on the threads not participating in the current round of iterative processing, so as to modify the mask identifiers corresponding to the threads not participating in the current round of iterative processing in the current thread mask from the first preset identifier to the second preset identifier; wherein the current thread mask includes the mask identifiers of the plurality of threads;
[0025] The number of threads in this round of iterative processing is determined according to the number of the first preset identifiers in the current thread mask.
[0026] Optionally, the detecting whether there is a communication error in the current round of iterative processing based on the consistency of the target memory address calculated by each receiving thread includes:
[0027] If the target memory addresses calculated by each receiving thread are consistent, it is determined that there is no communication error in this round of iterative processing;
[0028] If the target memory addresses calculated by each of the receiving threads are inconsistent, it is determined that a communication error exists in this round of iterative processing.
[0029] In a second aspect, the present application provides a communication error detection device, which is applied to a general-purpose graphics processor, comprising:
[0030] An acquisition module is used to acquire corresponding to-be-processed data blocks from the to-be-processed array based on memory access addresses of the to-be-processed array by a plurality of threads; each thread is configured with a corresponding thread identifier, and the thread identifiers of the plurality of threads are identifiers determined based on continuous numerical values; and the memory access addresses of the plurality of threads are memory addresses that increase in equal intervals;
[0031] A processing module, used for performing multiple rounds of iterative processing on the corresponding data blocks to be processed by the plurality of threads using an optimized reduction algorithm to obtain a final processing result; the optimized reduction algorithm is an algorithm obtained by optimizing a parallel reduction algorithm based on an end-reversal method; the end-reversal method is to control each sending thread in the threads participating in this round of iterative processing to transmit its own iterative processing result of this round to the corresponding receiving thread; the sum of the thread identifiers of the sending thread and the corresponding receiving thread is equal to a target value; the thread identifier of each sending thread is greater than or less than the corresponding receiving thread; the target value is the sum of the maximum thread identifier and the minimum thread identifier among the thread identifiers of the threads participating in this round of iterative processing;
[0032] The detection module is used to control each receiving thread to add its own memory access address and the corresponding memory access address of the sending thread in each round of iterative processing to calculate the target memory address, and based on the consistency of the target memory address calculated by each receiving thread, detect whether there is a communication error in this round of iterative processing.
[0033] In a third aspect, the present application provides an electronic device, including:
[0034] Memory, used to store computer programs;
[0035] A processor is used to execute the computer program to implement the aforementioned communication error detection method.
[0036] In a fourth aspect, the present application provides a computer-readable storage medium for storing a computer program, wherein the computer program implements the aforementioned communication error detection method when executed by a processor.
[0037] In the present application, a general graphics processor obtains corresponding to-be-processed data blocks from an array to be processed through several threads based on its own memory access address to the array to be processed; each thread is configured with a corresponding thread identifier, and the thread identifiers of the several threads are identifiers determined based on continuous numerical values; the memory access addresses of the several threads are memory addresses that increase at equal intervals; the several threads use an optimized reduction algorithm to perform multiple rounds of iterative processing on the corresponding to-be-processed data blocks to obtain a final processing result; the optimized reduction algorithm is an algorithm obtained by optimizing a parallel reduction algorithm based on an end-inversion method; the end-inversion method is to control each sending thread in the threads participating in this round of iterative processing to automatically The result of this round of iterative processing is transmitted to the corresponding receiving thread; the sum of the thread identifiers of the sending thread and the corresponding receiving thread is equal to the target value; the thread identifier of each sending thread is greater than or less than the corresponding receiving thread; the target value is the sum of the maximum thread identifier and the minimum thread identifier among the thread identifiers of the threads participating in this round of iterative processing; in each round of iterative processing, each receiving thread is controlled to add its own memory access address and the memory access address of the corresponding sending thread to calculate the target memory address, and based on the consistency of the target memory address calculated by each receiving thread, it is detected whether there is a communication error in this round of iterative processing.
[0038] It can be seen that the present application configures corresponding thread identifiers for several threads required to process the array to be processed based on continuous numerical values, and configures the memory access addresses of several threads to the array to be processed as memory addresses that increase at equal intervals. Based on this, when several threads use the optimized reduction algorithm to perform multiple rounds of iterative processing on the array to be processed, since the optimized reduction algorithm limits the sum of the thread identifiers of the sending thread and the corresponding receiving thread in each round of iterative processing to be equal to the target value, and the target value is the sum of the maximum thread identifier and the minimum thread identifier among the thread identifiers of the threads participating in this round of iterative processing, therefore, the addition result of the memory access address of the sending thread and the corresponding receiving thread in each round of iterative processing will be the same, that is, the target memory address calculated by each receiving thread will be the same, so that based on the consistency of the target memory address calculated by each receiving thread, it is possible to detect whether there is a communication error in this round of iterative processing, so as to timely discover communication errors in the data processing process to avoid DUE caused by data access address errors, thereby improving the reliability of the general graphics processor. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0040] Figure 1 A flow chart of a communication error detection method disclosed in this application;
[0041] Figure 2 A schematic diagram of a multi-round iterative process disclosed in this application;
[0042] Figure 3 A schematic diagram of the structure of a communication error detection device disclosed in this application;
[0043] Figure 4 This is a structural diagram of an electronic device disclosed in this application. DETAILED DESCRIPTION
[0044] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0045] Although detectable unrecoverable type errors (DUE) can be detected, there is no way to recover, so there is a high probability that the program will stop or become unresponsive. DUEs caused by data access address errors during data processing are a relatively common type of error. To this end, the present application provides a communication error detection method, so that when several threads use an optimized reduction algorithm to perform multiple rounds of iterative processing on the array to be processed, for each round of iterative processing, based on the consistency of the target memory address calculated by each receiving thread, it is detected whether there is a communication error in the current round of iterative processing, so as to timely detect communication errors in the data processing process to avoid DUEs caused by data access address errors, thereby improving the reliability of general-purpose graphics processors.
[0046] See also Figure 1 As shown, an embodiment of the present invention discloses a communication error detection method, which is applied to a general-purpose graphics processor, comprising:
[0047] Step S11, obtaining corresponding data blocks to be processed from the array to be processed by a number of threads based on their own memory access addresses to the array to be processed; each thread is configured with a corresponding thread identifier, and the thread identifiers of the several threads are identifiers determined based on continuous numerical values; the memory access addresses of the several threads are memory addresses that increase at equal intervals.
[0048] The general-purpose graphics processor in this embodiment adopts a general-purpose SIMT (Single Instruction Multiple Threads) architecture represented by CUDA (Compute Unified Device Architecture).
[0049] Furthermore, the general graphics processor first obtains a computing task for the array to be processed, and then determines the number of threads required for the computing task based on the data type and data size of the array to be processed; wherein the total number of threads is a power of two. Then, a corresponding thread identifier is configured for each thread based on a continuous numerical value starting from zero, that is, different threads are configured with different thread identifiers, and the memory access address of each thread to the array to be processed is determined based on a predetermined address spacing, the starting address of the array to be processed and the thread identifier of each thread; wherein the address spacing is a spacing predetermined based on the data type and data size of the array to be processed. After determining the memory access address of each thread to the array to be processed, the corresponding data block to be processed is obtained from the array to be processed by the threads based on their own memory access address of the array to be processed.
[0050] Taking the case where the number of threads required for the computing task is K threads and the address spacing is d, a corresponding thread ID is configured for each thread based on a continuous value starting from zero, that is, the thread IDs of the K threads are 0, 1, 2, ..., K-1 in sequence; then, according to the address spacing d, the starting address addr0 of the array to be processed and the thread ID of each thread, the memory access address of each thread to the array to be processed is determined; that is, the memory access address of the array to be processed by the thread with thread ID k is addr(k)=addr0+k×d. It can be found that the memory access addresses of several threads are memory addresses that increase in equal spacing, and the combination of the data blocks to be processed obtained from the array to be processed by several threads based on their own memory access addresses to the array to be processed is the array to be processed.
[0051] Among them, for determining the number of threads required for the computing task, first determine the total number of threads based on the data type and data size of the array to be processed, and the total number of threads is a power of two. Then determine the number of thread bundles required for the computing task based on the total number of threads and the number of threads corresponding to a thread bundle. The number of threads corresponding to a thread bundle can also be understood as the number of threads contained in a thread bundle, and each thread bundle contains the same number of threads. Then, use the number of thread bundles to determine the corresponding number of continuous thread bundles from the thread block, and use the total number of threads to determine the corresponding number of continuous threads from the continuous thread bundle to obtain the number of threads required for the computing task. The relationship between thread blocks, thread bundles and threads is that a thread block includes multiple thread bundles, and a thread bundle includes multiple threads; generally, a thread bundle includes 32 threads.
[0052] Step S12, performing multiple rounds of iterative processing on the corresponding data blocks to be processed by the several threads using an optimized reduction algorithm to obtain a final processing result; the optimized reduction algorithm is an algorithm obtained by optimizing a parallel reduction algorithm based on an end-reversal method; the end-reversal method is to control each sending thread in the threads participating in this round of iterative processing to transmit its own iterative processing result of this round to the corresponding receiving thread; the sum of the thread identifiers of the sending thread and the corresponding receiving thread is equal to a target value; the thread identifier of each sending thread is greater than or less than the corresponding receiving thread; the target value is the sum of the maximum thread identifier and the minimum thread identifier among the thread identifiers of the threads participating in this round of iterative processing.
[0053] In this embodiment, after several threads obtain the corresponding data blocks to be processed from the array to be processed based on their own memory access addresses to the array to be processed, the several threads use an optimized reduction algorithm to perform multiple rounds of iterative processing on the corresponding data blocks to be processed to obtain the final processing result. Among them, the optimized reduction algorithm is an algorithm obtained by optimizing the parallel reduction algorithm based on the end reversal method. It should be noted that the parallel reduction algorithm calculates partial results of the corresponding data blocks by each thread, and transmits and interacts the partial results between two threads to perform multiple rounds of iterative operations, wherein the number of threads is continuously halved in each round of iterative operations until only one thread is left to obtain the final result. The end reversal method optimizes the data transmission and interaction process in the parallel reduction algorithm, and proposes a corresponding communication error detection method for the optimized transmission and interaction process, so as to promptly detect communication errors in the data processing process.
[0054] Since several threads need to use an optimized reduction algorithm to perform multiple rounds of iterative processing on the corresponding data blocks to be processed, in each round of iterative processing, the end reversal method is specifically as follows: each sending thread among the threads participating in this round of iterative processing is controlled to transmit its own iterative processing result of this round to the corresponding receiving thread; wherein the sum of the thread identifiers of the sending thread and the corresponding receiving thread is equal to the target value; the thread identifier of each sending thread is greater than or less than the thread identifier of the corresponding receiving thread; and the target value is the sum of the maximum thread identifier and the minimum thread identifier among the thread identifiers of the threads participating in this round of iterative processing.
[0055] Taking the example that the threads participating in this round of iterative processing include four threads with thread identifiers 0, 1, 2, and 3, at this time, the target value 3 can be calculated first according to the maximum thread identifier 3 and the minimum thread identifier 0. Since the sum of the thread identifiers of the sending thread and the corresponding receiving thread is equal to the target value 3, when the thread identifiers of each sending thread are greater than the thread identifier of the corresponding receiving thread, the sending thread and the corresponding receiving thread can be a thread with thread identifier 3 and a thread with thread identifier 0, or a thread with thread identifier 2 and a thread with thread identifier 1. Then the thread with thread identifier 3 is controlled to transmit its own result of this round of iterative processing to the thread with thread identifier 0, and the thread with thread identifier 2 is controlled to transmit its own result of this round of iterative processing to the thread with thread identifier 1.
[0056] Based on this, the sending thread and the corresponding receiving thread can also be understood as the Nth thread in reverse order and the Nth thread in forward order among the threads participating in this round of iterative processing; N is a positive integer, and N is not greater than half of the number of threads in this round; the number of threads in this round is the number of threads participating in this round of iterative processing.
[0057] In another specific implementation, the process of determining each sending thread and the corresponding receiving thread in the threads participating in this round of iterative processing includes: in each round of iterative processing, determining the thread number of the threads participating in this round of iterative processing, and converting the value obtained by subtracting one from the thread number of this round of processing into a first binary number; then determining a preset number of threads with the largest or smallest thread identifiers from the threads participating in this round of iterative processing to obtain each sending thread; wherein the preset number is half of the thread number of this round; finally, converting the thread identifier of the sending thread into a second binary number, and performing an XOR operation on the first binary number and the second binary number to convert the XOR result into a decimal number to obtain a target thread identifier, and then determining the thread corresponding to the target thread identifier as the receiving thread corresponding to the sending thread.
[0058] Taking the example that the threads participating in this round of iterative processing include four threads with thread identifiers 0, 1, 2, and 3, the number of threads in this round is 4, the preset number is 2, and then the value obtained by subtracting one from the number of threads in this round is converted into a first binary number, that is, 3 is converted into a first binary number 0011; if the two threads with the largest thread identifiers are determined as sending threads from the four threads participating in this round of iterative processing, then each sending thread includes a thread with thread identifier 3 and a thread with thread identifier 2. For the sending thread being the thread with thread identifier 3, 3 is converted into a second binary number 0011, and an XOR operation is performed on the first binary number 0011 and the second binary number 0011 to obtain an XOR result of 0000, and the XOR result 0000 is converted into a decimal number 0 to obtain a target thread identifier, and then the thread corresponding to the target thread identifier is determined as the receiving thread corresponding to the sending thread, that is, the receiving thread corresponding to the thread with thread identifier 3 is the thread with thread identifier 0. For the sending thread being the thread with thread ID 2, 2 is converted into the second binary number 0010, and an XOR operation is performed on the first binary number 0011 and the second binary number 0010 to obtain an XOR result 0001, and the XOR result 0001 is converted into a decimal number 1 to obtain the target thread ID, and then the thread corresponding to the target thread ID is determined as the receiving thread corresponding to the sending thread, that is, the receiving thread corresponding to the thread with thread ID 2 is the thread with thread ID 1.
[0059] For the current round thread number of threads participating in the current round of iterative processing, a specific method is: in each round of iterative processing, directly count the number of threads participating in the current round of iterative processing to obtain the current round thread number.
[0060] Another specific method for determining the number of threads in the current round of the iterative processing of the threads participating in the current round of the iterative processing is as follows: in each round of the iterative processing, the current thread mask is updated based on the threads not participating in the current round of the iterative processing, so as to modify the mask identifiers corresponding to the threads not participating in the current round of the iterative processing in the current thread mask from the first preset identifier to the second preset identifier; wherein the current thread mask includes the mask identifiers of a plurality of threads. Then, the number of threads in the current round of the threads participating in the current round of the iterative processing is determined according to the number of the first preset identifiers in the current thread mask.
[0061] Take the case where the number of threads required for the computing task is four threads, and the thread identifiers of the four threads are 0, 1, 2, and 3 in sequence. Since the current thread mask includes the mask identifiers of the several threads, and the mask identifiers of the several threads are all the first preset identifier, such as 1, at the initial time, the current thread mask is 1111. In the first round of iterative processing, since all four threads participate in this round of iterative processing, it is not necessary to update the current thread mask. At this time, the number of threads participating in this round of iterative processing is determined to be 4 according to the number of 1s in the current thread mask. In the second round of iterative processing, since only threads with thread identifiers 0 and 1 participate in this round of iterative processing, it is necessary to update the current thread mask based on the threads not participating in this round of iterative processing, that is, the mask identifier corresponding to the threads not participating in this round of iterative processing in the current thread mask is modified from the first preset identifier 1 to the second preset identifier 0, and the updated current thread mask is 1100. Then, according to the number of 1s in the current thread mask, it is determined that the number of threads participating in this round of iterative processing is 2.
[0062] In this embodiment, after several threads obtain corresponding data blocks to be processed from the array to be processed based on their own memory access addresses to the array to be processed, multiple rounds of iterative processing processes for the corresponding data blocks to be processed by the several threads are started, so that the first round of iterative processing process is used as the current round of iterative processing process, and the several threads are used as threads participating in the first round of iterative processing; and in the current round of iterative processing, the threads participating in this round of iterative processing (that is, several threads at this time) process their respective data blocks to be processed to obtain corresponding results of this round of iterative processing, and then control each sending thread among the threads participating in this round of iterative processing to transmit its own results of this round of iterative processing to the corresponding receiving thread based on the end-inversion method, and after the transmission is completed, each receiving thread is used as a thread participating in the next round of iterative processing, and then enters the next round of iterative processing.
[0063] The next round of iterative processing is used as the new current round of iterative processing. In the current round of iterative processing, the threads participating in the current round of iterative processing determine the results of the current round of iterative processing based on their own results of the previous round of iterative processing and the received results of the previous round of iterative processing. Then the process jumps again to the step of controlling each sending thread in the threads participating in the current round of iterative processing based on the end reversal method to transmit its own results of the current round of iterative processing to the corresponding receiving thread, until there is only one thread participating in the current round of iterative processing, and the latest result of the current round of iterative processing is used as the final processing result.
[0064] Step S13: In each round of iterative processing, each receiving thread is controlled to add its own memory access address and the corresponding memory access address of the sending thread to calculate the target memory address, and based on the consistency of the target memory address calculated by each receiving thread, it is detected whether there is a communication error in this round of iterative processing.
[0065] In this embodiment, since the thread identifiers of several threads are identifiers determined based on continuous numerical values, the memory access addresses of several threads are memory addresses that increase at equal intervals, and the sum of the thread identifiers of each sending thread and the corresponding receiving thread in each round of iterative processing is equal to the target value, therefore, the addition results of the memory access addresses of each sending thread and the corresponding receiving thread in each round of iterative processing should be the same.
[0066] Taking the case where the number of threads required for the computing task is K threads and the address spacing is d, a corresponding thread ID is configured for each thread based on a continuous numerical value starting from zero, that is, the thread IDs of the K threads are 0, 1, 2, ..., K-1 respectively; then, according to the address spacing d, the starting address addr0 of the array to be processed and the thread ID of each thread, the memory access address of each thread to the array to be processed is determined; that is, the memory access address of the array to be processed by the thread with thread ID k is addr(k)=addr0+k×d. Based on this, in each round of iterative processing, assuming that the thread identifier of the sending thread is m, and the thread identifier of the corresponding receiving thread is n, since the sum of the thread identifiers of the sending thread and the corresponding receiving thread is equal to the target value p, the addition result of the memory access addresses of the sending thread and the corresponding receiving thread is addr(m)+addr(n)=addr0+m×d+addr0+n×d=2addr0+p×d; it can be found that since addr0 and d are constants, and p is also a constant during this round of iterative processing, therefore, during this round of iterative processing, the addition result of the memory access addresses of the sending thread and the corresponding receiving thread is a constant that is independent of the thread identifiers of the sending thread and the corresponding receiving thread.
[0067] Taking the example that the threads participating in this round of iterative processing include four threads with thread identifiers 0, 1, 2, and 3, the sending thread and the corresponding receiving thread at this time should be the thread with thread identifier 3 and the thread with thread identifier 0, and the thread with thread identifier 2 and the thread with thread identifier 1. After the thread with control thread identifier 3 transmits its own result of this round of iterative processing to the thread with thread identifier 0, and the thread with control thread identifier 2 transmits its own result of this round of iterative processing to the thread with thread identifier 1, the thread with control thread identifier 0 adds its own memory access address and the memory access address of the thread with thread identifier 3 to calculate the corresponding target memory address, and the thread with control thread identifier 1 adds its own memory access address and the memory access address of the thread with thread identifier 2 to calculate the corresponding target memory address. According to the memory access address of the array to be processed by the thread with thread identification k, addr(k)=addr0+k×d, it can be known that under the correct circumstances, the two target memory addresses should both be 2addr0+3d, but if the memory access address of the thread with thread identification 2 is wrong, it changes from the correct addr0+2d to the wrong addr0+3d, and the two target memory addresses will be inconsistent at this time; therefore, based on the consistency of the target memory addresses calculated by each receiving thread, it is possible to detect whether there is a communication error in this round of iterative processing.
[0068] Based on this, in each round of iterative processing, each receiving thread is controlled to add its own memory access address and the corresponding sending thread's memory access address to calculate the target memory address, and then based on the consistency of the target memory address calculated by each receiving thread, it is detected whether there is a communication error in this round of iterative processing. If the target memory addresses calculated by each receiving thread are consistent, it is determined that there is no communication error in this round of iterative processing; if the target memory addresses calculated by each receiving thread are inconsistent, it is determined that there is a communication error in this round of iterative processing.
[0069] It should be noted that when each thread needs to process multiple arrays at the same time, the communication error detection method proposed in this application can also be used to detect communication errors in the data processing process. Assume that each thread needs to process array A and array B at the same time, and the threads required to process array A and array B are both K threads, and the thread identifiers of the K threads are 0, 1, 2, ..., K-1, respectively, the starting address of array A is addrA0, the starting address of array B is addrB0, and the address spacing is d. At this time, the memory access address of the thread with thread identifier k to array A and array B can be expressed as addr(k)=addrA0+addrB0+2×k×d. Based on this, in each round of iterative processing, assuming that the thread identifier of the sending thread is m, and the thread identifier of the corresponding receiving thread is n, since the sum of the thread identifiers of the sending thread and the corresponding receiving thread is equal to the target value p, the sum of the memory access addresses of the sending thread and the corresponding receiving thread is addr(m)+addr(n)=2(addrA0+addrB0)+2×p×d, and the sum of the memory access addresses of the sending thread and the corresponding receiving thread is also a constant. Therefore, even if each thread needs to process multiple arrays at the same time, the communication error detection method proposed in this application can be used to detect communication errors in the data processing process.
[0070] It can be seen that the present application configures corresponding thread identifiers for several threads required to process the array to be processed based on continuous numerical values, and configures the memory access addresses of several threads to the array to be processed as memory addresses that increase at equal intervals. Based on this, when several threads use the optimized reduction algorithm to perform multiple rounds of iterative processing on the array to be processed, since the optimized reduction algorithm limits the sum of the thread identifiers of the sending thread and the corresponding receiving thread in each round of iterative processing to be equal to the target value, and the target value is the sum of the maximum thread identifier and the minimum thread identifier among the thread identifiers of the threads participating in this round of iterative processing, therefore, the addition result of the memory access address of the sending thread and the corresponding receiving thread in each round of iterative processing will be the same, that is, the target memory address calculated by each receiving thread will be the same, so that based on the consistency of the target memory address calculated by each receiving thread, it is possible to detect whether there is a communication error in this round of iterative processing, so as to timely discover communication errors in the data processing process to avoid DUE caused by data access address errors, thereby improving the reliability of the general graphics processor.
[0071] See also Figure 2 As shown, an embodiment of the present invention discloses a communication error detection method, which is applied to a general-purpose graphics processor, comprising:
[0072] Obtain a computing task for an array to be processed, and determine the total number of threads N+1 and the address spacing d based on the data type and data size of the array to be processed, and then determine the N+1 continuous threads required for the computing task from the multiple thread bundles included in the thread block to obtain a number of threads. Afterwards, configure a corresponding thread ID for each thread based on a continuous numerical value starting from zero, that is, the thread IDs of the N+1 threads are 0, 1, 2, ..., N in sequence, and determine the memory access address of each thread to the array to be processed based on the address spacing d, the starting address addr0 of the array to be processed, and the thread ID of each thread, that is, the memory access address of the array to be processed by the thread with thread ID k is addr(k)=addr0+k×d.
[0073] Furthermore, several threads obtain corresponding data blocks to be processed from the array to be processed based on their own memory access addresses to the array to be processed, and start multiple rounds of iterative processing of the corresponding data blocks to be processed by several threads.
[0074] In the first round of iterative processing process T0, the threads participating in this round of iterative processing (that is, N+1 threads at this time) process their respective data blocks to be processed to obtain the corresponding results of this round of iterative processing, and then control each sending thread in the threads participating in this round of iterative processing to transmit its own results of this round of iterative processing to the corresponding receiving thread based on the end reversal method; wherein the sum of the thread identifiers of the sending thread and the corresponding receiving thread is equal to N, and the thread identifier of the sending thread is greater than the corresponding receiving thread. After each sending thread completes the transmission, each receiving thread is controlled to add its own memory access address and the memory access address of the corresponding sending thread to calculate the target memory address, and the target memory address at this time should be 2addr0+N×d. If the target memory address calculated by each receiving thread is not 2addr0+N×d, it is determined that there is a communication error in this round of iterative processing, and the multiple rounds of iterative processing are terminated, and a communication error prompt is reported. If the target memory address calculated by each receiving thread is 2addr0+N×d, it is determined that there is no communication error in this round of iterative processing, and each receiving thread is used as a thread participating in the second round of iterative processing, and then enters the second round of iterative processing.
[0075] In the second round of iterative processing process T1, the threads participating in this round of iterative processing (i.e., thread 0 to thread (N-1) / 2) determine the corresponding iterative processing result of this round based on their own previous round of iterative processing results and the received previous round of iterative processing results, and then control each sending thread in the threads participating in this round of iterative processing to transmit its own iterative processing result of this round to the corresponding receiving thread based on the end reversal method; wherein the sum of the thread identifiers of the sending thread and the corresponding receiving thread is equal to (N-1) / 2, and the thread identifier of the sending thread is greater than the corresponding receiving thread. After each sending thread completes the transmission, each receiving thread is controlled to add its own memory access address and the memory access address of the corresponding sending thread to calculate the target memory address, and the target memory address at this time should be 2addr0+[(N-1) / 2]×d. If the target memory address calculated by each receiving thread is not 2addr0+[(N-1) / 2]×d, it is determined that there is a communication error in this round of iterative processing, and the multiple rounds of iterative processing are terminated, and a communication error prompt is reported. If the target memory address calculated by each receiving thread is 2addr0+[(N-1) / 2]×d, it is determined that there is no communication error in this round of iterative processing, and each receiving thread is used as a thread participating in the next round of iterative processing, and then enters the next round of iterative processing.
[0076] And so on, until the next round of iterative processing, when there is only one thread participating in this round of iterative processing, the thread participating in this round of iterative processing (that is, thread 0) determines the corresponding result of this round of iterative processing based on its own result of the previous round of iterative processing and the received result of the previous round of iterative processing, and takes the latest result of this round of iterative processing as the final processing result to be written back to the sender of the computing task.
[0077] It can be seen that the present application configures corresponding thread identifiers for several threads required to process the array to be processed based on continuous numerical values, and configures the memory access addresses of several threads to the array to be processed as memory addresses that increase at equal intervals. Based on this, when several threads use the optimized reduction algorithm to perform multiple rounds of iterative processing on the array to be processed, since the optimized reduction algorithm limits the sum of the thread identifiers of the sending thread and the corresponding receiving thread in each round of iterative processing to be equal to the target value, and the target value is the sum of the maximum thread identifier and the minimum thread identifier among the thread identifiers of the threads participating in this round of iterative processing, therefore, the addition result of the memory access address of the sending thread and the corresponding receiving thread in each round of iterative processing will be the same, that is, the target memory address calculated by each receiving thread will be the same, so that based on the consistency of the target memory address calculated by each receiving thread, it is possible to detect whether there is a communication error in this round of iterative processing, so as to timely discover communication errors in the data processing process to avoid DUE caused by data access address errors, thereby improving the reliability of the general graphics processor.
[0078] See also Figure 3 As shown, an embodiment of the present invention discloses a communication error detection device, which is applied to a general-purpose graphics processor, comprising:
[0079] The acquisition module 11 is used to acquire corresponding to-be-processed data blocks from the to-be-processed array based on the memory access addresses of the to-be-processed array by a plurality of threads; each thread is configured with a corresponding thread identifier, and the thread identifiers of the plurality of threads are identifiers determined based on continuous numerical values; the memory access addresses of the plurality of threads are memory addresses increasing at equal intervals;
[0080] The processing module 12 is used to perform multiple rounds of iterative processing on the corresponding data blocks to be processed by the plurality of threads using an optimized reduction algorithm to obtain a final processing result; the optimized reduction algorithm is an algorithm obtained by optimizing a parallel reduction algorithm based on an end-reversal method; the end-reversal method is to control each sending thread in the threads participating in this round of iterative processing to transmit its own iterative processing result of this round to the corresponding receiving thread; the sum of the thread identifiers of the sending thread and the corresponding receiving thread is equal to a target value; the thread identifier of each sending thread is greater than or less than the corresponding receiving thread; the target value is the sum of the maximum thread identifier and the minimum thread identifier among the thread identifiers of the threads participating in this round of iterative processing;
[0081] The detection module 13 is used to control each of the receiving threads to add its own memory access address and the corresponding memory access address of the sending thread in each round of iterative processing to calculate the target memory address, and based on the consistency of the target memory address calculated by each of the receiving threads, detect whether there is a communication error in this round of iterative processing.
[0082] It can be seen that the present application configures corresponding thread identifiers for several threads required to process the array to be processed based on continuous numerical values, and configures the memory access addresses of several threads to the array to be processed as memory addresses that increase at equal intervals. Based on this, when several threads use the optimized reduction algorithm to perform multiple rounds of iterative processing on the array to be processed, since the optimized reduction algorithm limits the sum of the thread identifiers of the sending thread and the corresponding receiving thread in each round of iterative processing to be equal to the target value, and the target value is the sum of the maximum thread identifier and the minimum thread identifier among the thread identifiers of the threads participating in this round of iterative processing, therefore, the addition result of the memory access address of the sending thread and the corresponding receiving thread in each round of iterative processing will be the same, that is, the target memory address calculated by each receiving thread will be the same, so that based on the consistency of the target memory address calculated by each receiving thread, it is possible to detect whether there is a communication error in this round of iterative processing, so as to timely discover communication errors in the data processing process to avoid DUE caused by data access address errors, thereby improving the reliability of the general graphics processor.
[0083] In some specific embodiments, the communication error detection device further includes:
[0084] A thread determination module, used for obtaining a computing task for an array to be processed, and determining a number of threads required for the computing task based on a data type and a data size of the array to be processed; the total number of threads of the number of threads is a power of two;
[0085] An access address determination unit is used to configure a corresponding thread identifier for each thread based on a continuous numerical value starting from zero, and determine the memory access address of each thread to the array to be processed according to a predetermined address spacing, the starting address of the array to be processed and the thread identifier of each thread; wherein the address spacing is a spacing predetermined based on the data type and data size of the array to be processed.
[0086] In some specific embodiments, the thread determination module includes:
[0087] a quantity determination unit, configured to determine the total number of threads based on the data type and data size of the array to be processed, and determine the number of thread warps required for the computing task according to the total number of threads and the number of threads corresponding to one thread warp;
[0088] The thread acquisition unit is used to determine a corresponding number of continuous thread warps from the thread block using the number of thread warps, and to determine a corresponding number of continuous threads from the continuous thread warps using the total number of threads, so as to obtain a number of threads required for the computing task.
[0089] In some specific embodiments, the processing module 12 includes:
[0090] A binary number conversion unit, used for determining the current round thread number of the threads participating in the current round of iterative processing in each round of iterative processing, and converting the value obtained by subtracting one from the current round thread number into a first binary number;
[0091] A sending thread determination unit, used to determine a preset number of threads with the largest or smallest thread identifiers from threads participating in this round of iterative processing, so as to obtain each of the sending threads; the preset number is half of the number of threads in this round;
[0092] A receiving thread determination unit is used to convert the thread identifier of the sending thread into a second binary number, and perform an XOR operation on the first binary number and the second binary number to convert the XOR result into a decimal number to obtain a target thread identifier, and then determine the thread corresponding to the target thread identifier as the receiving thread corresponding to the sending thread.
[0093] In some specific embodiments, the processing module 12 includes:
[0094] An opening unit, used to open multiple rounds of iterative processing of the corresponding data blocks to be processed by the plurality of threads, so as to use the first round of iterative processing as the current round of iterative processing, and use the plurality of threads as threads participating in the first round of iterative processing;
[0095] The iterative processing unit is used to process the respective data blocks to be processed by the threads participating in the current round of iterative processing to obtain the corresponding results of the current round of iterative processing, and control the sending threads among the threads participating in the current round of iterative processing to transmit their own results of the current round of iterative processing to the corresponding receiving threads based on the end reversal mode, and after the transmission is completed, use the receiving threads as threads participating in the next round of iterative processing, and then enter the next round of iterative processing;
[0096] The result determination unit is used to use the next round of iterative processing process as the new current round of iterative processing process, and in the current round of iterative processing process, determine the result of this round of iterative processing based on the thread participating in this round of iterative processing and the received result of the previous round of iterative processing, and then jump again to the step of controlling each sending thread in the threads participating in this round of iterative processing based on the end reversal method to transmit its own result of this round of iterative processing to the corresponding receiving thread, until there is only one thread participating in this round of iterative processing, and the latest result of this round of iterative processing is used as the final processing result.
[0097] In some specific embodiments, the binary number conversion unit includes:
[0098] A mask updating unit, configured to update the current thread mask based on threads not participating in the current round of iterative processing in each round of iterative processing, so as to modify the mask identifier corresponding to the threads not participating in the current round of iterative processing in the current thread mask from a first preset identifier to a second preset identifier; wherein the current thread mask includes the mask identifiers of the plurality of threads;
[0099] The thread number determination unit is used to determine the thread number of this round of threads participating in this round of iterative processing according to the number of the first preset identifiers in the current thread mask.
[0100] In some specific embodiments, the detection module 13 includes:
[0101] A communication error detection unit is used to determine that there is no communication error in this round of iterative processing if the target memory addresses calculated by each of the receiving threads are consistent; if the target memory addresses calculated by each of the receiving threads are inconsistent, it is determined that there is a communication error in this round of iterative processing.
[0102] Furthermore, the present application also discloses an electronic device. Figure 4 This is a structural diagram of an electronic device 20 according to an exemplary embodiment. The content in the diagram cannot be regarded as any limitation on the scope of use of the present application.
[0103] Figure 4 A schematic diagram of the structure of an electronic device 20 provided in an embodiment of the present application. The electronic device 20 may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 is used to store a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the communication error detection method disclosed in any of the aforementioned embodiments. In addition, the electronic device 20 in this embodiment may specifically be an electronic computer.
[0104] In this embodiment, the power supply 23 is used to provide working voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and the external device, and the communication protocol it follows is any communication protocol that can be applied to the technical solution of the present application, and is not specifically limited here; the input and output interface 25 is used to obtain external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs and is not specifically limited here.
[0105] In addition, the memory 22, as a carrier for storing resources, can be a read-only memory, a random access memory, a disk or an optical disk, etc. The resources stored thereon can include an operating system 221, a computer program 222, etc., and the storage method can be temporary storage or permanent storage.
[0106] The operating system 221 is used to manage and control the hardware devices and computer programs 222 on the electronic device 20, and can be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program that can be used to complete the communication error detection method performed by the electronic device 20 disclosed in any of the aforementioned embodiments, the computer program 222 can further include a computer program that can be used to complete other specific tasks.
[0107] Furthermore, the present application also discloses a computer-readable storage medium for storing a computer program; wherein the computer program, when executed by a processor, implements the communication error detection method disclosed above. For the specific steps of the method, reference may be made to the corresponding contents disclosed in the above embodiments, and no further description will be given here.
[0108] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.
[0109] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0110] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0111] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.
[0112] The technical solution provided by the present application is introduced in detail above. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for general technicians in this field, according to the idea of the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. A communication error detection method, characterized in that: Applicable to general-purpose graphics processors, including: Obtaining corresponding to-be-processed data blocks from the to-be-processed array by a plurality of threads based on their own memory access addresses to the to-be-processed array; each thread is configured with a corresponding thread identifier, and the thread identifiers of the plurality of threads are identifiers determined based on continuous numerical values; and the memory access addresses of the plurality of threads are memory addresses increasing at equal intervals; The optimized reduction algorithm is used by the several threads to perform multiple rounds of iterative processing on the corresponding data blocks to be processed to obtain the final processing result; the optimized reduction algorithm is an algorithm obtained by optimizing the parallel reduction algorithm based on the end reversal method; the end reversal method is to control each sending thread in the threads participating in the current round of iterative processing to transmit its own current round of iterative processing result to the corresponding receiving thread; the sum of the thread identifiers of the sending thread and the corresponding receiving thread is equal to the target value; the thread identifier of each sending thread is greater than or less than the corresponding receiving thread; the target value is the sum of the maximum thread identifier and the minimum thread identifier among the thread identifiers of the threads participating in the current round of iterative processing; During each round of iterative processing, each receiving thread is controlled to add its own memory access address and the corresponding memory access address of the sending thread to calculate the target memory address. Based on the consistency of the target memory address calculated by each receiving thread, it is detected whether there is a communication error in this round of iterative processing.
2. The communication error detection method according to claim 1, characterized in that: Before obtaining corresponding to-be-processed data blocks from the to-be-processed array based on the memory access addresses of the to-be-processed array by the plurality of threads, the method further includes: Obtaining a computing task for an array to be processed, and determining a number of threads required for the computing task based on a data type and a data size of the array to be processed; the total number of threads of the number of threads is a power of two; A corresponding thread identifier is configured for each thread based on a continuous numerical value starting from zero, and a memory access address of each thread to the array to be processed is determined according to a predetermined address spacing, a starting address of the array to be processed and a thread identifier of each thread; wherein the address spacing is a spacing predetermined based on a data type and a data size of the array to be processed.
3. The communication error detection method according to claim 2, characterized in that: The determining of the number of threads required for the computing task based on the data type and data size of the array to be processed includes: Determine the total number of threads based on the data type and data size of the array to be processed, and determine the number of thread warps required for the computing task according to the total number of threads and the number of threads corresponding to one thread warp; A corresponding number of continuous thread warps are determined from the thread block using the number of thread warps, and a corresponding number of continuous threads are determined from the continuous thread warps using the total number of threads, so as to obtain a number of threads required for the computing task.
4. The communication error detection method according to claim 2, characterized in that: The process of determining each sending thread and the corresponding receiving thread includes: In each round of iterative processing, determining the current round thread number of threads participating in the current round of iterative processing, and converting a value obtained by subtracting one from the current round thread number into a first binary number; Determine a preset number of threads with the largest or smallest thread identifiers from the threads participating in the current round of iterative processing to obtain each of the sending threads; the preset number is half of the number of threads in the current round; The thread identifier of the sending thread is converted into a second binary number, and an XOR operation is performed on the first binary number and the second binary number to convert the XOR result into a decimal number to obtain a target thread identifier, and then the thread corresponding to the target thread identifier is determined as the receiving thread corresponding to the sending thread.
5. The communication error detection method according to claim 4, characterized in that: The performing multiple rounds of iterative processing on the corresponding data blocks to be processed by the plurality of threads using the optimized reduction algorithm to obtain the final processing result includes: Starting multiple rounds of iterative processing of the corresponding data blocks to be processed by the plurality of threads, taking the first round of iterative processing as the current round of iterative processing, and taking the plurality of threads as threads participating in the first round of iterative processing; In the current round of iterative processing, the threads participating in the current round of iterative processing process process their respective data blocks to be processed to obtain corresponding results of the current round of iterative processing, and control the sending threads in the threads participating in the current round of iterative processing to transmit their own results of the current round of iterative processing to the corresponding receiving threads based on the end reversal mode, and after the transmission is completed, use the receiving threads as threads participating in the next round of iterative processing, and then enter the next round of iterative processing; The next round of iterative processing process is used as a new current round of iterative processing process, and in the current round of iterative processing process, the threads participating in this round of iterative processing determine the results of this round of iterative processing based on their own previous round of iterative processing results and the received previous round of iterative processing results, and then jump again to the step of controlling each sending thread in the threads participating in this round of iterative processing based on the end reversal method to transmit its own current round of iterative processing results to the corresponding receiving thread, until there is only one thread participating in this round of iterative processing, and the latest result of this round of iterative processing is used as the final processing result.
6. The communication error detection method according to claim 5, characterized in that: In each round of iterative processing, determining the number of threads participating in the current round of iterative processing includes: In each round of iterative processing, the current thread mask is updated based on the threads not participating in the current round of iterative processing, so as to modify the mask identifiers corresponding to the threads not participating in the current round of iterative processing in the current thread mask from the first preset identifier to the second preset identifier; wherein the current thread mask includes the mask identifiers of the plurality of threads; The number of threads in this round of iterative processing is determined according to the number of the first preset identifiers in the current thread mask.
7. The communication error detection method according to any one of claims 1 to 6, characterized in that: The detecting whether there is a communication error in the current round of iterative processing based on the consistency of the target memory address calculated by each receiving thread includes: If the target memory addresses calculated by each receiving thread are consistent, it is determined that there is no communication error in this round of iterative processing; If the target memory addresses calculated by each of the receiving threads are inconsistent, it is determined that a communication error exists in this round of iterative processing.
8. A communication error detection device, characterized in that: Applicable to general-purpose graphics processors, including: An acquisition module is used to acquire corresponding to-be-processed data blocks from the to-be-processed array based on memory access addresses of the to-be-processed array by a plurality of threads; each thread is configured with a corresponding thread identifier, and the thread identifiers of the plurality of threads are identifiers determined based on continuous numerical values; and the memory access addresses of the plurality of threads are memory addresses that increase in equal intervals; A processing module, used for performing multiple rounds of iterative processing on the corresponding data blocks to be processed by the plurality of threads using an optimized reduction algorithm to obtain a final processing result; the optimized reduction algorithm is an algorithm obtained by optimizing a parallel reduction algorithm based on an end-reversal method; the end-reversal method is to control each sending thread in the threads participating in this round of iterative processing to transmit its own iterative processing result of this round to the corresponding receiving thread; the sum of the thread identifiers of the sending thread and the corresponding receiving thread is equal to a target value; the thread identifier of each sending thread is greater than or less than the corresponding receiving thread; the target value is the sum of the maximum thread identifier and the minimum thread identifier among the thread identifiers of the threads participating in this round of iterative processing; The detection module is used to control each receiving thread to add its own memory access address and the corresponding memory access address of the sending thread in each round of iterative processing to calculate the target memory address, and based on the consistency of the target memory address calculated by each receiving thread, detect whether there is a communication error in this round of iterative processing.
9. An electronic device, characterized in that: include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the communication error detection method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: Used to store a computer program, which, when executed by a processor, implements the communication error detection method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Method and device for concurrently processing data objects in application program
CN117076130A
Multi-thread data processing method and apparatus
WO2022266842A1