Computing request parallel processing method and device, equipment and medium
By dividing the calculation request into subnet segments and performing graph fusion optimization processing, the slices are sent to the calculation segmentation block in parallel, the flexibility and time overhead problems of the calculation request processing scheme in the prior art are solved, and more efficient calculation request processing is achieved.
Patent Information
- Application Number
- CN202510568096.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-07-11
AI Technical Summary
In the prior art, the calculation request processing scheme is fixedly used to execute operators in the subnet segments through the calculation unit in sequence, resulting in poor flexibility and high time overhead.
The calculation request is divided into multiple subnet segments, determine whether to perform graph fusion optimization processing, slice the subnet segments that are optimized for graph fusion optimization, and then send it to a single computing segmentation block in parallel for calculation, and send it to each computing segmentation block for calculation for subnet segmentation without graph fusion optimization. The calculation segmentation blocks divided by hardware resources are used for data transmission.
It reduces the time overhead of computing request processing, improves the flexibility of processing, and avoids the consumption of a large amount of operation time.
Smart Images

Figure CN120301774A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular, to a method, apparatus, device, and medium for parallel processing of calculation requests. Background Art
[0002] With the development of artificial intelligence technology, more and more enterprises begin to perform processing processes such as image processing, text processing, or speech processing through neural network models running on electronic devices. The electronic device can be an artificial intelligence (AI) chip. A plurality of calculation units for performing calculation operations are provided in the electronic device. Calculation requests are generated in the neural network model running in the electronic device. A calculation request can refer to a data calculation task that needs to be executed by the electronic device. The data calculation task can be composed of multiple operators. An operator is a calculation instruction that can be executed by a calculation unit in the electronic device. The calculation instruction can be an instruction for instructing the calculation unit to perform a specified calculation operation on specified data. The calculation instruction can include the specified data that needs to be calculated and information for identifying the type of the specified calculation operation that needs to be executed. A calculation request can be divided into multiple subnet segments. Each subnet segment is composed of some of the operators that need to be executed in the calculation request.
[0003] In the related art, a common calculation request processing scheme is: after dividing the calculation request into multiple subnet segments, storing each subnet segment in the global storage of the electronic device. For each subnet segment, the operators in the subnet segment are sequentially executed by each calculation unit, and the calculation results of the operators in the subnet segment are written back to the global storage of the electronic device. The calculation request processing scheme in the related art fixedly adopts the method of sequentially executing the operators in the subnet segment through the calculation unit to execute each subnet segment, with poor flexibility, consuming a large amount of operation time and having a large time overhead. Summary of the Invention
[0004] The present invention provides a method, apparatus, device, and medium for parallel processing of calculation requests to solve the problem that the calculation request processing scheme in the related art fixedly adopts the method of sequentially executing the operators in the subnet segment through the calculation unit to execute each subnet segment, with poor flexibility, consuming a large amount of operation time and having a large time overhead.
[0005] According to one aspect of the present invention, there is provided a method for parallel processing of calculation requests, including:
[0006] Dividing a calculation request to be processed into multiple subnet segments, and determining whether each subnet segment is subjected to graph fusion optimization processing;
[0007] For each subnet segment undergoing graph fusion optimization processing, after slicing the subnet segment, each slice of the subnet segment is circulated and parallelly sent to a single computing partition block of the electronic device for calculation;
[0008] For each subnet segment not undergoing graph fusion optimization processing, the operators of the subnet segment are sent to each computing partition block of the electronic device for calculation;
[0009] Among them, each computing partition block is a hardware module divided according to hardware resources. Each computing partition block includes multiple computing units and a shared memory for the multiple computing units to use. Data transfer between the shared memories of each computing partition block is carried out through the global memory of the electronic device.
[0010] According to another aspect of the present invention, there is provided a computing request parallel processing device, including:
[0011] A processing and evaluation module, configured to divide a to-be-processed computing request into multiple subnet segments and determine whether each subnet segment undergoes graph fusion optimization processing;
[0012] A first processing module, configured to, for each subnet segment undergoing graph fusion optimization processing, after slicing the subnet segment, circulate and parallelly send each slice of the subnet segment to a single computing partition block of the electronic device for calculation;
[0013] A second processing module, configured to, for each subnet segment not undergoing graph fusion optimization processing, send the operators of the subnet segment to each computing partition block of the electronic device for calculation;
[0014] Among them, each computing partition block is a hardware module divided according to hardware resources. Each computing partition block includes multiple computing units and a shared memory for the multiple computing units to use. Data transfer between the shared memories of each computing partition block is carried out through the global memory of the electronic device.
[0015] According to another aspect of the present invention, there is provided an electronic device, and the electronic device includes:
[0016] At least one processor;
[0017] And a memory communicatively connected to the at least one processor;
[0018] Among them, the memory stores a computer program executed by the at least one processor. The computer program is executed by the at least one processor so that the at least one processor can execute the computing request parallel processing method according to any embodiment of the present invention.
[0019] According to another aspect of the present invention, there is provided a computer-readable storage medium storing computer instructions for causing a processor to execute the parallel processing method for computing requests according to any embodiment of the present invention when executed.
[0020] According to another aspect of the present invention, there is provided a computer program product including a computer program which, when executed by a processor, implements the parallel processing method for computing requests according to any embodiment of the present invention.
[0021] The technical solution of the embodiments of the present invention divides the computing requests to be processed into multiple subnet segments, and determines whether each subnet segment is to be subjected to graph fusion optimization processing; then, for each subnet segment subjected to graph fusion optimization processing, after slicing the subnet segment, the sliced subnet segments are cyclically and parallelly sent to a single computing partition block of an electronic device for computing; for each subnet segment not subjected to graph fusion optimization processing, the operators of the subnet segment are sent to each computing partition block of the electronic device for computing; wherein, each computing partition block is a hardware module divided according to hardware resources, each computing partition block includes multiple computing units and a shared memory for the multiple computing units to use, and data is transmitted between the shared memories of the respective computing partition blocks through the global memory of the electronic device. This solves the problem that the computing request processing solution in the related art fixedly uses the method of sequentially executing the operators in the subnet segment through the computing units to execute each subnet segment, with poor flexibility, consuming a large amount of operation time and having a large time overhead. After dividing the computing requests into multiple subnet segments, for each subnet segment, it can be determined whether the sliced subnet segments can be cyclically and parallelly sent to a single computing partition block of the electronic device for computing after slicing the subnet segment, so as to reduce the time overhead of executing the subnet segment. Furthermore, for each subnet segment determined to be able to reduce the time overhead, after slicing the subnet segment, the sliced subnet segments are cyclically and parallelly sent to a single computing partition block of the electronic device for computing, thereby improving the flexibility of the computing request processing process, avoiding consuming a large amount of operation time, and reducing the time overhead of the computing request processing process.
[0022] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0024] Figure 1 It is a flowchart of a method for parallel processing of calculation requests provided in Embodiment 1 of the present invention.
[0025] Figure 2 It is a flowchart of a method for parallel processing of calculation requests provided in Embodiment 2 of the present invention.
[0026] Figure 3 It is a schematic structural diagram of a device for parallel processing of calculation requests provided in Embodiment 3 of the present invention.
[0027] Figure 4 It is a schematic structural diagram of an electronic device for implementing the method for parallel processing of calculation requests in the embodiments of the present invention. Detailed implementation manners
[0028] In order to enable those skilled in the art to better understand the solutions of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0029] It should be noted that the terms "target", "first", "second", etc. in the specification and claims of the present invention and the above accompanying drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include", "comprise" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0030] Embodiment 1
[0031] Figure 1The flowchart of a method for parallel processing of computing requests provided in Embodiment 1 of the present invention. This embodiment is applicable to the situation of processing computing requests generated in a neural network model running in an electronic device. This method can be executed by a computing request parallel processing device, which can be implemented in the form of hardware and / or software, and the computing request parallel processing device can be configured in an electronic device. The electronic device can be an AI chip. The neural network model running in the electronic device can be used for processing processes such as image processing, text processing, or speech processing. As Figure 1 shown, the method includes:
[0032] Step 101: Divide the computing request to be processed into multiple subnet segments, and determine whether each subnet segment is subjected to graph fusion optimization processing.
[0033] Optionally, the computing request to be processed may refer to a computing request generated in a neural network model running in an electronic device that needs to be processed at a current moment. A computing request may refer to a data computing task that needs to be executed by the electronic device. The data computing task may be composed of multiple operators. An operator is a computing instruction that can be executed by a computing unit in an electronic device. The computing instruction may be an instruction for instructing the computing unit to perform a specified computing operation on specified data. The computing instruction may include the specified data that needs to be computed and information for identifying the type of the specified computing operation that needs to be executed. The specified computing operation that needs to be executed may be a multiplication operation, an addition operation, or other types of computing operations. A computing request can be divided into multiple subnet segments. Each subnet segment is composed of some of the operators that need to be executed in the computing request. The computing request to be processed can be divided into multiple subnet segments, and a corresponding global storage unit can be allocated for each subnet segment in the idle global storage units in the global storage of the electronic device, and the subnet segment can be stored in the corresponding global storage unit.
[0034] Optionally, the global storage of the electronic device may be a memory set in the electronic device for storing data that needs to be processed by the electronic device and the processing results obtained after processing. The global storage includes multiple global storage units. Each global storage unit can be a region divided from the global storage for storing data. Each global storage unit is set with identification information. The identification information of the global storage unit may be a string for identifying the global storage unit. The idle global storage units in the global storage of the electronic device may refer to the global storage units that do not store data.
[0035] Optionally, for each subnet segment, randomly obtain an idle global storage unit in each of the idle global storage units in the global storage of the electronic device as the global storage unit corresponding to the subnet segment, and store the subnet segment in the global storage unit corresponding to the subnet segment.
[0036] Optionally, determining whether to perform graph fusion optimization processing on each subnet segment may refer to, for each subnet segment, determining whether to slice the subnet segment and then cyclically and parallelly distribute the slices of each subnet segment to a single computing partition block of the electronic device for calculation, so as to reduce the time overhead of executing the subnet segment.
[0037] Optionally, each computing partition block in the electronic device is a hardware module divided according to hardware resources. Each computing partition block includes multiple computing units and a shared memory for the multiple computing units to use. Data is transferred between the shared memories of each computing partition block through the global memory of the electronic device. The hardware resources may refer to all the computing units set in the electronic device and the memories that all the computing units can use. A computing unit is a hardware module for performing computing operations. A memory is a hardware module for storing data. The computing units included in each computing partition block are different. A shared memory for the multiple computing units included in the computing partition block may refer to an area specifically provided for the multiple computing units included in the computing partition block divided from the memories that all the computing units can use. The shared memories included in different computing partition blocks are different. Data is not directly transferred between the shared memories of each computing partition block and needs to be transferred through the global memory of the electronic device.
[0038] Optionally, determining whether to perform graph fusion optimization processing on each subnet segment includes: performing the following operations for each subnet segment: determining the overhead before optimization and the overhead after optimization of the subnet segment according to the operator overhead parameter of the subnet segment; determining whether the overhead after optimization is less than the overhead before optimization; if the overhead after optimization is less than the overhead before optimization, determining that the subnet segment performs graph fusion optimization processing. The operator overhead parameters of each subnet segment of the computing request to be processed may be sent by the target user to the electronic device. The target user may be a technical person in charge of managing the electronic device.
[0039] Optionally, the operator overhead parameters of the subnet segment may include the total number of operators in the subnet segment, the operator overhead before optimization of the subnet segment, the operator overhead after optimization of the subnet segment, the synchronization overhead, the total number of loop times of the subnet segment, and the total number of computing partition blocks in the electronic device. The total number of operators in the subnet segment is the total number of operators included in the subnet segment. The operator overhead before optimization of the subnet segment may refer to the estimated duration required for all computing operations performed in the process of executing the subnet segment by sequentially executing the operators in the subnet segment through each computing unit. The operator overhead after optimization of the subnet segment may refer to the estimated duration required for all computing operations performed in the process of executing the subnet segment by slicing the subnet segment and then cyclically and parallelly distributing each sliced subnet segment to a single computing partition block of the electronic device for calculation. The synchronization overhead may refer to the duration required each time the hardware resources of the electronic device perform data synchronization. The total number of loop times of the subnet segment may refer to the number of loop times estimated when cyclically and parallelly distributing each sliced subnet segment of the subnet segment to a single computing partition block of the electronic device for calculation. The total number of loop times of the subnet segment is equal to the total number of sliced subnet segments in the subnet segment. The total number of computing partition blocks in the electronic device is the total number of computing partition blocks included in the electronic device.
[0040] Optionally, the overhead before optimization of the subnet segment may refer to the total duration required for the process of executing the subnet segment by sequentially executing the operators in the subnet segment through each computing unit. The overhead after optimization of the subnet segment may refer to the total duration required for the process of executing the subnet segment by slicing the subnet segment and then cyclically and parallelly distributing each sliced subnet segment to a single computing partition block of the electronic device for calculation.
[0041] Optionally, determining the overhead before optimization and the overhead after optimization of the subnet segment according to the operator overhead parameters of the subnet segment includes: using the following formula for calculating the overhead before optimization to calculate the overhead before optimization of the subnet segment:
[0042] T0 = T S0 + n * T D ,
[0043] where T0 is the overhead before optimization of the subnet segment, T S0 is the operator overhead before optimization of the subnet segment, n is the total number of operators in the subnet segment, and T D is the synchronization overhead; using the following formula for calculating the overhead after optimization to calculate the overhead after optimization of the subnet segment:
[0044] T1 = T S1 * ceil(C / S)+1 * T D ,
[0045] Among them, T1 is the optimized overhead of the subnet segment, and T S1 is the optimized operator overhead of the subnet segment, C is the total number of loops of the subnet segment, S is the total number of computing partition blocks in the electronic device, ceil(C / S) is the ceiling value of the calculation result obtained by dividing the total number of loops of the subnet segment by the total number of computing partition blocks in the electronic device, and T D is the synchronization overhead. The formula for the overhead before optimization can be a pre-set formula for calculating the overhead before optimization of the subnet segment based on the operator overhead before optimization of the subnet segment, the total number of operators in the subnet segment, and the synchronization overhead. The formula for the overhead after optimization can be a pre-set formula for calculating the overhead after optimization of the subnet segment based on the operator overhead after optimization of the subnet segment, the total number of loops of the subnet segment, the total number of computing partition blocks in the electronic device, and the synchronization overhead.
[0046] Optionally, if the optimized overhead of the subnet segment is less than the overhead before optimization, it indicates that by slicing the subnet segment and then cyclically and parallelly distributing each sliced subnet segment to a single computing partition block of the electronic device for calculation, the time overhead of executing the subnet segment can be reduced, and thus it is determined that the subnet segment is to be subjected to graph fusion optimization processing.
[0047] Optionally, after determining whether the optimized overhead is less than the overhead before optimization, it further includes: if the optimized overhead is greater than or equal to the overhead before optimization, it is determined that the subnet segment is not to be subjected to graph fusion optimization processing. If the optimized overhead of the subnet segment is greater than or equal to the overhead before optimization, it indicates that there is no need to slice the subnet segment and then cyclically and parallelly distribute each sliced subnet segment to a single computing partition block of the electronic device for calculation to reduce the time overhead of executing the subnet segment, and thus it is determined that the subnet segment is not to be subjected to graph fusion optimization processing.
[0048] Step 102: For each subnet segment to be subjected to graph fusion optimization processing, slice the subnet segment and then cyclically and parallelly distribute each sliced subnet segment to a single computing partition block of the electronic device for calculation.
[0049] Optionally, the electronic device includes N computing partition blocks; for each subnet segment to be subjected to graph fusion optimization processing, slicing the subnet segment and then cyclically and parallelly distributing each sliced subnet segment to a single computing partition block of the electronic device for calculation includes: performing the following operations for each subnet segment to be subjected to graph fusion optimization processing: slicing the subnet segment to obtain multiple sliced subnet segments; and cyclically and parallelly distributing each sliced subnet segment to N computing partition blocks for calculation according to a preset polling method.
[0050] Optionally, each computing partition block reads the subnet segment slice to be processed from the global storage unit corresponding to the subnet segment slice to be processed, stores the read subnet segment slice to be processed in the shared storage of each computing partition block, performs calculations on the subnet segment slice to be processed, and writes the calculation result back to the global storage unit corresponding to the calculation result of the subnet segment slice to be processed.
[0051] Optionally, N is a positive integer greater than or equal to 2. A subnet segment contains multiple operators. A subnet segment slice can refer to an operator in the subnet segment. The subnet segment can be sliced to obtain multiple subnet segment slices, and a digital serial number is assigned to each subnet segment slice starting from the number 0. Exemplarily, the subnet segment is sliced to obtain 8 subnet segment slices, and a digital serial number is assigned to each subnet segment slice starting from the number 0. The digital numbers of the 8 subnet segment slices are respectively: 0, 1, 2, 3, 4, 5, 6, 7. Then, in the free global storage units in the global storage of the electronic device, a global storage unit corresponding to the subnet segment slice and a global storage unit corresponding to the calculation result of the subnet segment slice are assigned to each subnet segment slice, and the subnet segment slice is stored in the global storage unit corresponding to the subnet segment slice. The global storage unit corresponding to the subnet segment slice can be the global storage unit for storing the subnet segment slice. The global storage unit corresponding to the calculation result of the subnet segment slice can be the global storage unit for storing the calculation result of the subnet segment slice. The calculation result of the subnet segment slice can be the calculation result obtained after the computing unit executes the subnet segment slice.
[0052] Optionally, for each subnet segment slice, two free global storage units are randomly obtained from the respective free global storage units in the global storage of the electronic device, one as the global storage unit corresponding to the subnet segment slice and one as the global storage unit corresponding to the calculation result of the subnet segment slice.
[0053] Optionally, in accordance with a preset round-robin method, each subnet segment slice is cyclically and parallelly distributed to N computing partition blocks for calculation, including: in accordance with the arrangement order of the N computing partition blocks, each subnet segment slice is cyclically and parallelly distributed to the N computing partition blocks for calculation.
[0054] Optionally, the N computing split blocks are arranged in a preset order. According to the arrangement order of the N computing split blocks, each subnet segment slice is circularly and parallelly sent to the N computing split blocks for calculation, including: determining the N subnet segment slices with the smallest digital serial numbers among the un-sent subnet segment slices as the N subnet segment slices to be parallelly sent to the N computing split blocks for calculation this time; parallelly sending the slice storage addresses and slice result storage addresses of the determined N subnet segment slices to the N computing split blocks, so as to parallelly send the determined N subnet segment slices to the N computing split blocks for calculation; wherein, the slice storage address and slice result storage address of the subnet segment slice whose digital serial number ranks first in the ascending order of the digital serial numbers of the determined N subnet segment slices are sent to the first computing split block among the N computing split blocks, the slice storage address and slice result storage address of the subnet segment slice whose digital serial number ranks second in the ascending order of the digital serial numbers of the determined N subnet segment slices are sent to the second computing split block among the N computing split blocks, and so on, the slice storage address and slice result storage address of the subnet segment slice whose digital serial number ranks Nth in the ascending order of the digital serial numbers of the determined N subnet segment slices are sent to the Nth computing split block among the N computing split blocks; after determining that the N computing split blocks have completed the calculation process of the sent subnet segment slices, return to execute the operation of determining the N subnet segment slices with the smallest digital serial numbers among the un-sent subnet segment slices as the N subnet segment slices to be parallelly sent to the N computing split blocks for calculation this time, until all subnet segment slices have been sent to the computing split blocks for calculation or the number of the last remaining un-sent subnet segment slices is less than N.
[0055] Optionally, the slice storage address of the subnet segment slice is the identification information of the global storage unit corresponding to the subnet segment slice. The slice result storage address of the subnet segment slice is the identification information of the global storage unit corresponding to the calculation result of the subnet segment slice.
[0056] Optionally, for each computational partition block, the subnet segment slice to which the slice storage address and the slice result storage address received by the computational partition block belong is the subnet segment slice to be processed by the computational partition block. After the computational partition block receives the slice storage address and the slice result storage address of the subnet segment slice to be processed, it can determine the global storage unit corresponding to the subnet segment slice to be processed according to the slice storage address of the subnet segment slice to be processed, read the subnet segment slice to be processed from the global storage unit corresponding to the subnet segment slice to be processed, and store the read subnet segment slice to be processed in the shared storage of the computational partition block. Then, the computational partition block executes the subnet segment slice to be processed through the computing unit in the computational partition block to complete the computing operation required for the subnet segment slice, thereby performing the computation on the subnet segment slice to be processed, and determining the global storage unit corresponding to the computation result of the subnet segment slice to be processed according to the slice result storage address of the subnet segment slice to be processed, and storing the obtained computation result in the global storage unit corresponding to the computation result of the subnet segment slice to be processed, thereby writing the computation result back to the global storage unit corresponding to the computation result of the subnet segment slice to be processed.
[0057] Optionally, it is possible to detect whether there is data stored in the global storage unit corresponding to the computation result of the already issued subnet segment slice. When it is detected that there is data stored in the global storage unit corresponding to the computation result of the already issued subnet segment slice, it can be determined that the corresponding computational partition block has completed the computation process of the already issued subnet segment slice. When it is detected that there is data stored in the global storage units corresponding to the computation results of the already issued N subnet segment slices, it can be determined that the N computational partition blocks have completed the computation processes of the issued subnet segment slices. Then, after it is determined that the N computational partition blocks have completed the computation processes of the issued subnet segment slices, the operation of determining the N subnet segment slices with the smallest digital serial numbers among the unissued subnet segment slices as the N subnet segment slices to be concurrently issued to the N computational partition blocks for computation is returned and executed until all the subnet segment slices have been issued to the computational partition blocks for computation or the number of the remaining unissued subnet segment slices is less than N.
[0058] Optionally, if all the subnet segment slices have been issued to the computational partition blocks for computation, it is determined that the current issuing process ends.
[0059] Optionally, if the number of each remaining subnet segment slice that has not been distributed is M, where M is a positive integer less than N, then the slice storage addresses and slice result storage addresses of the last remaining M subnet segment slices that have not been distributed are sent in parallel to the first M computing partition blocks among the N computing partition blocks, so as to distribute the last remaining M subnet segment slices that have not been distributed to the M computing partition blocks in parallel for computing, and determine that the current distribution process ends. Among them, the slice storage address and slice result storage address of the subnet segment slice whose digital serial number ranks first in the ascending order of the digital serial numbers of the last remaining M subnet segment slices that have not been distributed are sent to the first computing partition block among the N computing partition blocks, the slice storage address and slice result storage address of the subnet segment slice whose digital serial number ranks second in the ascending order of the digital serial numbers of the last remaining M subnet segment slices that have not been distributed are sent to the second computing partition block among the N computing partition blocks, and so on. The slice storage address and slice result storage address of the subnet segment slice whose digital serial number ranks Mth in the ascending order of the digital serial numbers of the last remaining M subnet segment slices that have not been distributed are sent to the Mth computing partition block among the N computing partition blocks.
[0060] Optionally, in a specific example, N is 2, and the two computing partition blocks are arranged in a preset order: computing partition block 0 and computing partition block 1. The currently processed subnet segment is divided into 8 subnet segment slices. The digital numbers of the 8 subnet segment slices are respectively: 0, 1, 2, 3, 4, 5, 6, 7.
[0061] First, slice the subnet segments numbered 0 and 1 to determine two subnet segment slices that are to be distributed in parallel to two computing partition blocks for computing. Send the slice storage addresses and slice result storage addresses of the determined subnet segment slices numbered 0 and 1 to computing partition block 0 and computing partition block 1 in parallel, so as to distribute the determined two subnet segment slices to two computing partition blocks for computing in parallel. Among them, the slice storage address and slice result storage address of the subnet segment slice numbered 0 are sent to computing partition block 0, which is the first one among the two computing partition blocks, and the slice storage address and slice result storage address of the subnet segment slice numbered 1 are sent to computing partition block 1, which is the second one among the two computing partition blocks. After receiving the slice storage address and slice result storage address of the subnet segment slice numbered 0, computing partition block 0 can determine the global storage unit corresponding to the subnet segment slice numbered 0 according to the slice storage address of the subnet segment slice numbered 0, read the subnet segment slice numbered 0 from the global storage unit corresponding to the subnet segment slice numbered 0, and store the read subnet segment slice numbered 0 in the shared storage of the computing partition block. Then, computing partition block 0 executes the subnet segment slice numbered 0 through the computing unit in computing partition block 0 to complete the computing operations required for the subnet segment slice numbered 0, thereby computing the subnet segment slice numbered 0, and determining the global storage unit corresponding to the computing result of the subnet segment slice numbered 0 according to the slice result storage address of the subnet segment slice numbered 0, and storing the obtained computing result in the global storage unit corresponding to the computing result of the subnet segment slice numbered 0, so as to write the computing result back to the global storage unit corresponding to the computing result of the subnet segment slice numbered 0. After receiving the slice storage address and slice result storage address of the subnet segment slice numbered 1, computing partition block 1 can determine the global storage unit corresponding to the subnet segment slice numbered 1 according to the slice storage address of the subnet segment slice numbered 1, read the subnet segment slice numbered 1 from the global storage unit corresponding to the subnet segment slice numbered 1, and store the read subnet segment slice numbered 1 in the shared storage of the computing partition block.Then, calculate partition block 1. By executing the subnet segment slice numbered 1 through the computing units in partition block 1, perform the computing operations required to complete the subnet segment slice numbered 1, thereby calculating the subnet segment slice numbered 1. Based on the slice result storage address of the subnet segment slice numbered 1, determine the global storage unit corresponding to the calculation result of the subnet segment slice numbered 1, and store the obtained calculation result in the global storage unit corresponding to the calculation result of the subnet segment slice numbered 1, thereby writing back the calculation result to the global storage unit corresponding to the calculation result of the subnet segment slice numbered 1.
[0062] After determining that partition block 0 and partition block 1 have completed the calculation processes for the subnet segment slices numbered 0 and 1, determine the subnet segment slices numbered 2 and 3 as the two subnet segment slices to be parallelly distributed to the two partition blocks for calculation this time. Parallelly send the slice storage addresses and slice result storage addresses of the determined subnet segment slices numbered 2 and 3 to partition block 0 and partition block 1, thereby parallelly distributing the determined two subnet segment slices to the two partition blocks for calculation. Among them, the slice storage address and slice result storage address of the subnet segment slice numbered 2 are sent to partition block 0, which is the first among the two partition blocks, and the slice storage address and slice result storage address of the subnet segment slice numbered 3 are sent to partition block 1, which is the second among the two partition blocks.
[0063] After determining that partition block 0 and partition block 1 have completed the calculation processes for the subnet segment slices numbered 2 and 3, determine the subnet segment slices numbered 4 and 5 as the two subnet segment slices to be parallelly distributed to the two partition blocks for calculation this time. Parallelly send the slice storage addresses and slice result storage addresses of the determined subnet segment slices numbered 4 and 5 to partition block 0 and partition block 1, thereby parallelly distributing the determined two subnet segment slices to the two partition blocks for calculation. Among them, the slice storage address and slice result storage address of the subnet segment slice numbered 4 are sent to partition block 0, which is the first among the two partition blocks, and the slice storage address and slice result storage address of the subnet segment slice numbered 5 are sent to partition block 1, which is the second among the two partition blocks.
[0064] After determining that calculation partition block 0 and calculation partition block 1 have completed the calculation process of subnet segment slices numbered 4 and 5, subnet segment slices numbered 6 and 7 are determined as the two subnet segment slices that are currently dispatched in parallel to the two calculation partition blocks for calculation. The slice storage addresses and slice result storage addresses of the determined subnet segment slices numbered 6 and 7 are sent to calculation partition block 0 and calculation partition block 1 in parallel, so as to dispatch the determined two subnet segment slices to the two calculation partition blocks for calculation in parallel. Among them, the slice storage address and slice result storage address of the subnet segment slice numbered 6 are sent to the first-ranked calculation partition block 0 among the two calculation partition blocks, and the slice storage address and slice result storage address of the subnet segment slice numbered 7 are sent to the second-ranked calculation partition block 1 among the two calculation partition blocks. After determining that calculation partition block 0 and calculation partition block 1 have completed the calculation process of the subnet segment slices numbered 6 and 7, it is determined that all subnet segment slices have been dispatched to the calculation partition blocks for calculation, and it is determined that the current dispatch process ends.
[0065] Calculation partition block 0 is responsible for processing subnet segment slices with even numbers: subnet segment slice numbered 0, subnet segment slice numbered 2, subnet segment slice numbered 4, subnet segment slice numbered 6. Calculation partition block 1 is responsible for processing subnet segment slices with odd numbers: subnet segment slice numbered 1, subnet segment slice numbered 3, subnet segment slice numbered 5, subnet segment slice numbered 7.
[0066] Optionally, according to a preset round-robin method, each subnet segment slice is dispatched to N calculation partition blocks for calculation in parallel, including: according to the arrangement order of the identification information of the N calculation partition blocks, each subnet segment slice is dispatched to the N calculation partition blocks for calculation in parallel. The identification information of each calculation partition block can be the digital number used to identify the calculation partition block. The identification information of each calculation partition block is different.
[0067] Optionally, according to the arrangement order of the identification information of the N computing partition blocks, each subnet segment slice is circularly and parallelly sent to the N computing partition blocks for calculation, including: determining the N subnet segment slices with the smallest digital serial numbers among the un-sent subnet segment slices as the N subnet segment slices to be parallelly sent to the N computing partition blocks for calculation; parallelly sending the slice storage addresses and slice result storage addresses of the determined N subnet segment slices to the N computing partition blocks, so as to parallelly send the determined N subnet segment slices to the N computing partition blocks for calculation; among them, the slice storage address and slice result storage address of the subnet segment slice with the digital serial number ranked first in the sorting sequence of the digital serial numbers of the determined N subnet segment slices from small to large are sent to the computing partition block with the identification information ranked first in the sorting sequence of the identification information of the N computing partition blocks from small to large, the slice storage address and slice result storage address of the subnet segment slice with the digital serial number ranked second in the sorting sequence of the digital serial numbers of the determined N subnet segment slices from small to large are sent to the computing partition block with the identification information ranked second in the sorting sequence of the identification information of the N computing partition blocks from small to large, and so on, the slice storage address and slice result storage address of the subnet segment slice with the digital serial number ranked Nth in the sorting sequence of the digital serial numbers of the determined N subnet segment slices from small to large are sent to the computing partition block with the identification information ranked Nth in the sorting sequence of the identification information of the N computing partition blocks from small to large; after determining that the N computing partition blocks have completed the calculation process of the sent subnet segment slices, return to execute the operation of determining the N subnet segment slices with the smallest digital serial numbers among the un-sent subnet segment slices as the N subnet segment slices to be parallelly sent to the N computing partition blocks for calculation, until all subnet segment slices have been sent to the computing partition blocks for calculation or the number of the last remaining un-sent subnet segment slices is less than N.
[0068] Optionally, if all subnet segment slices have been sent to the computing partition blocks for calculation, it is determined that the current sending process ends.
[0069] Optionally, if the number of remaining subnet fragment slices that have not been distributed is M, where M is a positive integer less than N, then the slice storage addresses and slice result storage addresses of the last remaining M subnet fragment slices that have not been distributed are sent in parallel to the first M computing partition blocks among the N computing partition blocks, so as to distribute the last remaining M subnet fragment slices that have not been distributed to the M computing partition blocks in parallel for calculation, and determine the end of the current distribution process. Among them, the slice storage address and slice result storage address of the subnet fragment slice whose serial number among the last remaining M subnet fragment slices that have not been distributed ranks first in the sorting sequence of the serial numbers of the last remaining M subnet fragment slices that have not been distributed from small to large are sent to the computing partition block whose identification information ranks first in the sorting sequence of the identification information of the N computing partition blocks from small to large. The slice storage address and slice result storage address of the subnet fragment slice whose serial number among the last remaining M subnet fragment slices that have not been distributed ranks second in the sorting sequence of the serial numbers of the last remaining M subnet fragment slices that have not been distributed from small to large are sent to the computing partition block whose identification information ranks second in the sorting sequence of the identification information of the N computing partition blocks from small to large, and so on. The slice storage address and slice result storage address of the subnet fragment slice whose serial number among the last remaining M subnet fragment slices that have not been distributed ranks Mth in the sorting sequence of the serial numbers of the last remaining M subnet fragment slices that have not been distributed from small to large are sent to the computing partition block whose identification information ranks Mth in the sorting sequence of the identification information of the N computing partition blocks from small to large.
[0070] Step 103: For each subnet fragment that does not undergo graph fusion optimization processing, send the operators of the subnet fragment to each computing partition block of the electronic device for calculation.
[0071] Among them, each computing partition block is a hardware module divided according to hardware resources. Each computing partition block includes multiple computing units and a shared storage for the multiple computing units to use. Data is transferred between the shared storages of each computing partition block through the global storage of the electronic device.
[0072] Optionally, for each subnet fragment that does not undergo graph fusion optimization processing, sending the operators of the subnet fragment to each computing partition block of the electronic device for calculation includes: performing the following operations for each subnet fragment that does not undergo graph fusion optimization processing: in the order of arrangement of the operators in the subnet fragment, starting from the first operator in the subnet fragment, send the operator to a computing partition block in the electronic device, so that the computing partition block completes the calculation operation required by the operator through the computing units in the computing partition block and returns the calculation result. After obtaining the calculation result, continue to execute the next operator until the calculation result of the last operator is obtained.
[0073] In the technical solution of the embodiment of the present invention, by dividing the to-be-processed calculation request into multiple subnet segments, it is determined whether each subnet segment is subjected to graph fusion optimization processing; then, for each subnet segment subjected to graph fusion optimization processing, after slicing the subnet segment, each subnet segment slice is cyclically and parallelly sent to a single calculation partition block of the electronic device for calculation; for each subnet segment not subjected to graph fusion optimization processing, the operators of the subnet segment are sent to each calculation partition block of the electronic device for calculation; wherein, each calculation partition block is a hardware module divided according to hardware resources, each calculation partition block includes multiple calculation units and a shared memory for the multiple calculation units to use, and data transfer between the shared memories of each calculation partition block is performed through the global memory of the electronic device, which solves the problem that the calculation request processing solution in the related art fixedly uses the method of sequentially executing the operators in the subnet segment through the calculation units to execute each subnet segment, with poor flexibility, consuming a large amount of operation time and having a large time overhead. After dividing the calculation request into multiple subnet segments, for each subnet segment, it can be determined whether the time overhead of executing the subnet segment can be reduced by slicing the subnet segment and cyclically and parallelly sending each subnet segment slice to a single calculation partition block of the electronic device, and then for each determined subnet segment that can reduce the time overhead, after slicing the subnet segment, each subnet segment slice is cyclically and parallelly sent to a single calculation partition block of the electronic device for calculation, thereby improving the flexibility of the calculation request processing process, avoiding consuming a large amount of operation time, and reducing the time overhead of the calculation request processing process.
[0074] Embodiment 2
[0075] Figure 2 It is a flowchart of a method for parallel processing of calculation requests provided by Embodiment 2 of the present invention. The embodiments of the present invention can be combined with various optional solutions in one or more of the above embodiments. As Figure 2 shown, the method includes:
[0076] Step 201: Divide the to-be-processed calculation request into multiple subnet segments, and determine whether each subnet segment is subjected to graph fusion optimization processing.
[0077] Step 202: For each subnet segment subjected to graph fusion optimization processing, slice the subnet segment to obtain multiple subnet segment slices, and cyclically and parallelly send each subnet segment slice to N calculation partition blocks for calculation according to a preset polling method.
[0078] Step 203: For each subnet segment not subjected to graph fusion optimization processing, send the operators of the subnet segment to each calculation partition block of the electronic device for calculation.
[0079] Among them, each computing partition block is a hardware module divided according to hardware resources. Each computing partition block includes multiple computing units and a shared memory for the multiple computing units to use. Data transfer between the shared memories of the respective computing partition blocks is performed through the global memory of the electronic device.
[0080] The technical solution of the embodiment of the present invention can, after dividing a computing request into multiple subnet segments, for each subnet segment, determine whether it is possible to slice the subnet segment and then circularly and parallelly distribute the sliced subnet segments to a single computing partition block of the electronic device for computing, reduce the time overhead of executing the subnet segment, and then for each determined subnet segment that can reduce the time overhead, slice the subnet segment and then circularly and parallelly distribute the sliced subnet segments to a single computing partition block of the electronic device for computing, thereby improving the flexibility of the computing request processing process, avoiding consuming a large amount of operation time, and reducing the time overhead of the computing request processing process.
[0081] Embodiment III
[0082] Figure 3 It is a schematic structural diagram of a computing request parallel processing device provided in Embodiment III of the present invention. The device can be configured in an electronic device. As Figure 3 shown, the device includes: a processing evaluation module 301, a first processing module 302, and a second processing module 303.
[0083] Among them, the processing evaluation module 301 is configured to divide a computing request to be processed into multiple subnet segments and determine whether each subnet segment undergoes graph fusion optimization processing; the first processing module 302 is configured to, for each subnet segment that undergoes graph fusion optimization processing, slice the subnet segment and then circularly and parallelly distribute the sliced subnet segments to a single computing partition block of the electronic device for computing; the second processing module 303 is configured to, for each subnet segment that does not undergo graph fusion optimization processing, distribute the operators of the subnet segment to each computing partition block of the electronic device for computing; among them, each computing partition block is a hardware module divided according to hardware resources. Each computing partition block includes multiple computing units and a shared memory for the multiple computing units to use. Data transfer between the shared memories of the respective computing partition blocks is performed through the global memory of the electronic device.
[0084] The technical solution of the embodiment of the present invention divides the to-be-processed computing request into multiple subnet segments, and determines whether each subnet segment is to be subjected to graph fusion optimization processing; then, for each subnet segment to be subjected to graph fusion optimization processing, after slicing the subnet segment, each slice of the subnet segment is circularly and parallelly sent to a single computing partition block of the electronic device for calculation; for each subnet segment that is not to be subjected to graph fusion optimization processing, the operators of the subnet segment are sent to each computing partition block of the electronic device for calculation; wherein, each computing partition block is a hardware module divided according to hardware resources, each computing partition block includes multiple computing units and a shared memory for the multiple computing units to use, and data is transferred between the shared memories of each computing partition block through the global memory of the electronic device, which solves the problem that the computing request processing solution in the related art fixedly uses the method of sequentially executing the operators in the subnet segment through the computing units to execute each subnet segment, with poor flexibility, consuming a large amount of operation time and having a large time overhead. After dividing the computing request into multiple subnet segments, for each subnet segment, it can be determined whether it is possible to circularly and parallelly send each slice of the subnet segment to a single computing partition block of the electronic device for calculation after slicing the subnet segment, so as to reduce the time overhead of executing the subnet segment. Furthermore, for each subnet segment determined to be able to reduce the time overhead, after slicing the subnet segment, each slice of the subnet segment is circularly and parallelly sent to a single computing partition block of the electronic device for calculation, thereby improving the flexibility of the computing request processing process, avoiding consuming a large amount of operation time, and reducing the time overhead of the computing request processing process.
[0085] In an alternative embodiment of the embodiment of the present invention, optionally, when the processing and evaluation module 301 executes the operation of determining whether each subnet segment is to be subjected to graph fusion optimization processing, it is specifically configured to: perform the following operations for each subnet segment: determine the overhead before optimization and the overhead after optimization of the subnet segment according to the operator overhead parameter of the subnet segment; determine whether the overhead after optimization is less than the overhead before optimization; if the overhead after optimization is less than the overhead before optimization, determine that the subnet segment is to be subjected to graph fusion optimization processing.
[0086] In an alternative embodiment of the embodiment of the present invention, optionally, the processing and evaluation module 301 is further configured to: if the overhead after optimization is greater than or equal to the overhead before optimization, determine that the subnet segment is not to be subjected to graph fusion optimization processing.
[0087] In an alternative embodiment of the embodiment of the present invention, optionally, the electronic device includes N computing segmentation blocks; specifically, the first processing module 302 is configured to perform the following operations for each subnet segment undergoing graph fusion optimization processing: slice the subnet segment to obtain a plurality of subnet segment slices; and cyclically and parallelly distribute the respective subnet segment slices to the N computing segmentation blocks for calculation according to a preset polling manner.
[0088] In an alternative embodiment of the embodiment of the present invention, optionally, when the first processing module 302 performs the operation of cyclically and parallelly distributing the respective subnet segment slices to the N computing segmentation blocks for calculation according to a preset polling manner, it is specifically configured to: cyclically and parallelly distribute the respective subnet segment slices to the N computing segmentation blocks for calculation according to the arrangement order of the N computing segmentation blocks.
[0089] In an alternative embodiment of the embodiment of the present invention, optionally, when the first processing module 302 performs the operation of cyclically and parallelly distributing the respective subnet segment slices to the N computing segmentation blocks for calculation according to a preset polling manner, it is specifically configured to: cyclically and parallelly distribute the respective subnet segment slices to the N computing segmentation blocks for calculation according to the arrangement order of the identification information of the N computing segmentation blocks.
[0090] In an alternative embodiment of the embodiment of the present invention, optionally, each computing segmentation block reads the to-be-processed subnet segment slice from the global storage unit corresponding to the to-be-processed subnet segment slice, stores the read to-be-processed subnet segment slice in the shared storage of each computing segmentation block, calculates the to-be-processed subnet segment slice, and writes the calculation result back to the global storage unit corresponding to the calculation result of the to-be-processed subnet segment slice.
[0091] The computing request parallel processing device provided by the embodiment of the present invention can execute the computing request parallel processing method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.
[0092] Embodiment IV
[0093] Figure 4FIG. 0 shows a schematic structural diagram of an electronic device 10 that can be used to implement the method for parallel processing of computing requests according to an embodiment of the present invention. The electronic device is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, electronic devices, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, personal digital assistants, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0094] As Figure 4 shown, the electronic device 10 includes at least one processor 11, and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. The memory stores a computer program executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0095] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0096] The processor 11 can be various general-purpose and / or special-purpose processing components having processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the method for parallel processing of computing requests.
[0097] In some embodiments, the method for parallel processing of computing requests may be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as a storage unit. In some embodiments, part or all of the computer program may be loaded and / or installed onto the heterogeneous hardware accelerator via the ROM and / or the communication unit. When the computer program is loaded into the RAM and executed by the processor, one or more steps of the method for parallel processing of computing requests described above may be performed. Alternatively, in other embodiments, the processor may be configured to execute the method for parallel processing of computing requests by any other suitable means (e.g., by means of firmware).
[0098] The various embodiments of the systems and techniques described above in this document may be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: being implemented in one or more computer programs that may be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0099] The computer program for implementing the method of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer programs are executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer programs may be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or electronic device.
[0100] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0101] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a heterogeneous hardware accelerator that has: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the heterogeneous hardware accelerator. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).
[0102] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data electronic device), or a computing system that includes middleware components (e.g., an application electronic device), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by any form or medium of digital data communication (e.g., a communication network). Examples of the communication network include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0103] A computing system may include a client and an electronic device. The client and the electronic device are generally far from each other and usually interact via a communication network. The relationship between the client and the electronic device is generated by computer programs running on corresponding computers and having a client-electronic device relationship with each other. The electronic device may be a cloud electronic device, also known as a cloud computing electronic device or a cloud host, which is a host product in the cloud computing service system, and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.
[0104] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is imposed herein.
[0105] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for parallel processing of computing requests, characterized in that, Including: Dividing a to-be-processed computing request into multiple subnet segments, and determining whether each subnet segment is to be processed with graph fusion optimization; For each subnet segment to be processed with graph fusion optimization, after slicing the subnet segment, circulating and parallelly distributing each subnet segment slice to a single computing partition block of an electronic device for computing; For each subnet segment not to be processed with graph fusion optimization, distributing the operators of the subnet segment to each computing partition block of the electronic device for computing; Wherein, each computing partition block is a hardware module divided according to hardware resources, each computing partition block includes multiple computing units and a shared memory for the multiple computing units to use, and data is transferred between the shared memories of each computing partition block through the global memory of the electronic device.
2. The parallel processing method for computing requests according to claim 1, wherein Determining whether each subnet segment is to be processed with graph fusion optimization includes: Performing the following operations for each subnet segment: According to the operator overhead parameter of the subnet segment, determining the overhead before optimization and the overhead after optimization of the subnet segment; Judging whether the overhead after optimization is less than the overhead before optimization; If the overhead after optimization is less than the overhead before optimization, determining that the subnet segment is to be processed with graph fusion optimization.
3. The parallel processing method for computing requests according to claim 2, wherein After judging whether the overhead after optimization is less than the overhead before optimization, it further includes: If the overhead after optimization is greater than or equal to the overhead before optimization, determining that the subnet segment is not to be processed with graph fusion optimization.
4. The calculation request parallel processing method according to claim 1, wherein The electronic device includes N computing partition blocks; For each subnet segment to be processed with graph fusion optimization, after slicing the subnet segment, circulating and parallelly distributing each subnet segment slice to a single computing partition block of an electronic device for computing includes: Performing the following operations for each subnet segment to be processed with graph fusion optimization: Slicing the subnet segment to obtain multiple subnet segment slices; According to a preset round-robin manner, circulating and parallelly distributing each subnet segment slice to N computing partition blocks for computing.
5. The parallel processing method for computing requests according to claim 4, wherein According to a preset round-robin manner, circulating and parallelly distributing each subnet segment slice to N computing partition blocks for computing includes: According to the arrangement order of the N computing partition blocks, circulating and parallelly distributing each subnet segment slice to N computing partition blocks for computing.
6. The parallel processing method for computing requests according to claim 4, wherein According to a preset round-robin manner, circulating and parallelly distributing each subnet segment slice to N computing partition blocks for computing includes: According to the arrangement order of the identification information of the N computing partition blocks, circulating and parallelly distributing each subnet segment slice to N computing partition blocks for computing.
7. The parallel processing method for computing requests according to claim 4, characterized in that, Each computing partition block reads the to-be-processed subnet segment slice from the global memory unit corresponding to the to-be-processed subnet segment slice, stores the read to-be-processed subnet segment slice into the shared memory of each computing partition block, performs computing on the to-be-processed subnet segment slice, and writes the computing result back to the global memory unit corresponding to the computing result of the to-be-processed subnet segment slice.
8. A computing request parallel processing device, characterized in that, Including: A processing evaluation module, configured to divide a to-be-processed computing request into multiple subnet segments, and determine whether each subnet segment is to be processed with graph fusion optimization; The first processing module is configured to, for each subnet segment subjected to graph fusion optimization processing, slice the subnet segment and then circularly and parallelly distribute each sliced subnet segment to a single computing partition block of the electronic device for calculation; The second processing module is configured to, for each subnet segment not subjected to graph fusion optimization processing, distribute the operators of the subnet segment to each computing partition block of the electronic device for calculation; Wherein, each computing partition block is a hardware module divided according to hardware resources, each computing partition block includes multiple computing units and a shared memory for the multiple computing units to use, and data is transmitted between the shared memories of each computing partition block through the global memory of the electronic device.
9. An electronic device, characterized in that, The electronic device includes: At least one processor; And a memory communicatively connected to the at least one processor; Wherein, the memory stores a computer program executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the calculation request parallel processing method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and the computer instructions are used to implement the calculation request parallel processing method according to any one of claims 1-7 when executed by a processor.