Broadcast operation method, device, equipment and medium in artificial intelligence operator
By using the step size parameters and total dimensions provided by the main control processor in the artificial intelligence chip for offset calculation, the problems of poor universality and low efficiency of broadcast calculation methods are solved, and efficient broadcast calculation is achieved.
Patent Information
- Application Number
- CN202411081210.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-07
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2044-08-07
AI Technical Summary
The existing broadcast operation methods are poor in generality and inefficient, and it is necessary to redevelop broadcast operation programs for array structures of different dimensions, resulting in inefficiency.
The front-end master processor provides step size parameters and total dimensions of different tensors in broadcast operations, and performs offset operations in artificial intelligence chips, combining the strong parallelism of the front-end + back-end heterogeneous hardware framework to achieve efficient broadcast operations.
It improves the universality and efficiency of broadcast operations, and does not need to develop different broadcast operations programs for array structures of different dimensions, so that broadcast operations can be completed quickly and accurately.
Smart Images

Figure CN118860693B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a broadcast operation method, device, equipment and medium in an artificial intelligence operator. Background Art
[0002] During the process of an artificial intelligence model executing a task, operators in an operator library are often called for relevant operations. The broadcast operator (or broadcast calculation) is one of the numerous operators in the artificial intelligence operator library. Broadcast operations allow mathematical operations between arrays of different sizes. For example, when performing operations such as matrix multiplication and convolution in data processing and machine learning, broadcast operations are often required to ensure the dimensional compatibility of arrays. Therefore, how to perform efficient broadcast operations between different arrays is particularly important.
[0003] In related technologies, when performing broadcast operations between arrays of the same dimension, it is often necessary to develop a broadcast operation program corresponding to that dimension. When the dimension of the array that needs to perform broadcast operations changes, it is necessary to re-develop a broadcast operation program corresponding to the changed dimension. During the entire broadcast operation process, both the program development and the broadcast operation process are sequentially executed in the same hardware chip, resulting in poor generality and very low efficiency of the broadcast operation method. Summary of the Invention
[0004] The present invention provides a broadcast operation method, device, equipment and medium in an artificial intelligence operator, which is used to solve the defects of poor generality and very low efficiency of the existing broadcast operation method, and greatly improves the generality and broadcast efficiency of the broadcast operation method by combining the strong parallelism of the heterogeneous hardware framework of the front end + back end.
[0005] The present invention provides a broadcast operation method in an artificial intelligence operator, which is applied to an artificial intelligence chip, and the method includes the following steps.
[0006] In response to a broadcast operation instruction, obtain the stride parameters and the total dimension of different tensors in the broadcast operator in each dimension; each of the stride parameters and the total dimension is provided by a main control processor;
[0007] Based on each of the stride parameters and the total dimension, perform offset calculations respectively to obtain the target offsets corresponding to the first tensor and the second tensor in the different tensors;
[0008] Based on each of the target offsets, obtain the first element in the first tensor and the second element in the second tensor respectively, and perform a preset operation on the first element and the second element, and save the operation result to a third tensor in the different tensors.
[0009] A broadcast operation method in an artificial intelligence operator provided according to the present invention, wherein the target offsets corresponding to the first tensor and the second tensor in the different tensors are respectively obtained by performing offset calculations based on each of the step parameters and the total dimension, including: for each dimension in the first target tensor, when the dimension does not reach the total dimension, calculating the offset corresponding to the dimension based on the historical offset of the historical dimension of the dimension and the step parameter of the dimension; the first target tensor is any one of the first tensor and the second tensor; incrementing the dimension by 1, and repeating the above steps until the target offset corresponding to the dimension when the dimension reaches the total dimension is calculated.
[0010] A broadcast operation method in an artificial intelligence operator provided according to the present invention, wherein calculating the offset corresponding to the dimension based on the historical offset of the historical dimension of the dimension and the step parameter of the dimension includes: when the dimension matches a preset parameter representing a broadcast operation, calculating the offset corresponding to the dimension based on the historical offset, the step parameter of the dimension, and the index value of the dimension; when the dimension does not match the preset parameter, calculating the offset corresponding to the dimension based on the historical offset, the step parameter of the dimension, and the coordinates of the element corresponding to the dimension in the third tensor.
[0011] A broadcast operation method in an artificial intelligence operator provided according to the present invention, the method further includes: for each dimension in the third tensor, when the dimension is not less than 1, using the data processing thread corresponding to the dimension to perform dimension parameter calculation on the linear index value of the dimension and the step parameter of the dimension, to obtain the coordinates of the element corresponding to the dimension and the historical linear index value of the historical dimension of the dimension; decrementing the dimension by 1, and repeating the above steps until the coordinates of the elements corresponding to all dimensions in the third tensor are obtained when the dimension is less than 1.
[0012] A broadcast operation method in an artificial intelligence operator provided according to the present invention, wherein in response to a broadcast operation instruction, obtaining the step parameters and the total dimension of different tensors in a broadcast operator in each dimension respectively includes: based on the first tensor and the second tensor carried in the broadcast operation instruction, copying the step parameters and the total dimension of the first tensor, the second tensor, and the third tensor respectively after preprocessing in each dimension from the main control processor.
[0013] According to a broadcast operation method in an artificial intelligence operator provided by the present invention, the main control processor is configured to perform the following steps for each dimension in the second target tensor: when the dimension has not reached the total dimension, calculate the step parameter for the dimension based on the historical step parameter of the historical dimension of the dimension and the dimension value of the historical dimension; increment the dimension by 1, and repeat the above steps until the step parameter for each dimension in the second target tensor is calculated; wherein the second target tensor is any one of the first tensor, the second tensor, and the third tensor.
[0014] According to a broadcast operation method in an artificial intelligence operator provided by the present invention, the obtaining of the first element in the first tensor and the second element in the second tensor respectively based on the respective target offsets includes: performing memory access based on the target offset corresponding to the first tensor and the start address of the first tensor in the memory to obtain the first element in the first tensor; performing memory access based on the target offset corresponding to the second tensor and the start address of the second tensor in the memory to obtain the second element in the second tensor.
[0015] The present invention also provides a broadcast operation device in an artificial intelligence operator, which is applied to an artificial intelligence chip and includes the following units.
[0016] A parameter acquisition unit, configured to obtain the step parameters for each dimension and the total dimension of different tensors in a broadcast operator in response to a broadcast operation instruction; each of the step parameters and the total dimension is provided by the main control processor.
[0017] A broadcast operation unit, configured to perform offset amount operations respectively based on each of the step parameters and the total dimension to obtain the respective target offsets corresponding to the first tensor and the second tensor in the different tensors; obtain the first element in the first tensor and the second element in the second tensor respectively based on the respective target offsets, and perform a preset operation on the first element and the second element, and save the operation result to the third tensor in the different tensors.
[0018] The present invention also provides an electronic device, including a main control processor and an artificial intelligence chip connected in the form of a heterogeneous framework. The artificial intelligence chip is configured to, in response to a broadcast operation instruction, obtain the stride parameters and the total dimension of different tensors in each dimension of a broadcast operator; perform offset calculations respectively based on each of the stride parameters and the total dimension to obtain the target offsets corresponding to a first tensor and a second tensor in the different tensors; obtain a first element in the first tensor and a second element in the second tensor respectively based on each of the target offsets, perform a preset calculation on the first element and the second element, and save the calculation result to a third tensor in the different tensors; and the main control processor is configured to provide each of the stride parameters and the total dimension to the artificial intelligence chip.
[0019] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the broadcast operation method in any one of the above artificial intelligence operators.
[0020] The broadcast operation method, device, equipment and medium in the artificial intelligence operator provided by the present invention. In the broadcast operation method in the artificial intelligence operator, when the artificial intelligence chip responds to a broadcast operation instruction, it first obtains the stride parameters and the total dimension of different tensors in each dimension of the broadcast operator, then further performs offset calculations respectively based on each of the stride parameters and the total dimension to obtain the target offsets corresponding to the first tensor and the second tensor in the different tensors, and then obtains the first element in the first tensor and the second element in the second tensor respectively based on each of the target offsets, performs a preset calculation on the first element and the second element, and saves the calculation result to the third tensor in the different tensors. Since the multiple stride parameters and the total dimension of the different tensors involved in the broadcast operation are all provided by the main control processor, therefore, by means of the main control processor at the front end providing the stride parameters and the total dimension of the different tensors in the broadcast operation and performing the broadcast operation based on all the stride parameters and the total dimension in the artificial intelligence chip, the broadcast operation result is determined efficiently and accurately. The entire broadcast operation process is simple and easy to implement. It is neither necessary to develop corresponding broadcast operation programs for different-dimensional array structures, nor is it necessary to use the main control processor provided at the front end to assist the artificial intelligence operator at the back end to quickly and accurately complete the broadcast operation, rather than being limited to the hardware chip of the artificial intelligence chip. Combining the strong parallelism of the front-end + back-end heterogeneous hardware framework, the generality and broadcast efficiency of the broadcast operation method are greatly improved. Description of the Drawings
[0021] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0022] Figure 1 It is one of the schematic flowcharts of the broadcast operation method in the artificial intelligence operator provided by the present invention.
[0023] Figure 2 It is the schematic diagram of the broadcast operation provided by the present invention.
[0024] Figure 3 It is the flowchart of the offset calculation provided by the present invention.
[0025] Figure 4 It is the flowchart of the coordinate calculation provided by the present invention.
[0026] Figure 5 It is the flowchart of the step parameter calculation provided by the present invention.
[0027] Figure 6 It is the second schematic flowchart of the broadcast operation method in the artificial intelligence operator provided by the present invention.
[0028] Figure 7 It is the schematic diagram of the relevant information of tensor A provided by the present invention.
[0029] Figure 8 It is the schematic diagram of the relevant information of tensor B provided by the present invention.
[0030] Figure 9 It is the schematic diagram of the relevant information of tensor C provided by the present invention.
[0031] Figure 10 It is the schematic diagram of the structure of the broadcast operation device in the artificial intelligence operator provided by the present invention.
[0032] Figure 11 It is the schematic diagram of the structure of the electronic device provided by the present invention. Detailed implementation manners
[0033] To make the objectives, technical solutions and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention in conjunction with the drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0034] In an embodiment of the present invention, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects and indicates that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone, where A and B can be singular or plural. In the written description of the present invention, the character " / " generally represents an "or" relationship between the associated objects before and after. In addition, it should be noted that the serial numbers assigned to the objects described in the present invention, such as "first", "second", etc., are only used to distinguish the described objects and do not have any sequential or technical meaning.
[0035] In the field of artificial intelligence technology, when an artificial intelligence model is used to execute a task, operators in the artificial intelligence operator library are often called for relevant operations. The broadcast operator (or broadcast calculation) is one of the many operators in the operator library. Broadcast operations allow mathematical operations between arrays of different sizes. For example, when performing operations such as matrix multiplication and convolution in data processing and machine learning, broadcast operations are often required to ensure the dimensional compatibility of the arrays. Therefore, how to perform efficient broadcast operations between different arrays is particularly important.
[0036] In the related art, when performing broadcast operations between arrays of the same dimension, it is often necessary to develop a broadcast operation program corresponding to that dimension. When the dimension of the array that requires broadcast operations changes, it is necessary to re-develop a broadcast operation program corresponding to the changed dimension. Both the program development and the broadcast operation process in the entire broadcast operation process are sequentially executed in the same hardware chip, resulting in poor generality and very low efficiency of the broadcast operation method.
[0037] To solve the above technical problems, the present invention provides a broadcast operation method, device, equipment, and medium in an artificial intelligence operator. By providing the step parameters and total dimensions of different tensors in the broadcast operation by the main control processor at the front end, and the method of performing broadcast operations based on all step parameters and total dimensions in the artificial intelligence chip, the broadcast operation result is determined efficiently and accurately. The entire broadcast operation process is simple and easy to implement. It is neither necessary to develop corresponding broadcast operation programs for different array structures of different dimensions, nor to use the main control processor provided by the front end to assist the artificial intelligence operator at the back end to quickly and accurately complete the broadcast operation, rather than being limited to the hardware chip of the artificial intelligence chip. Combining the strong parallelism of the heterogeneous hardware framework of the front end + back end, the generality and broadcast efficiency of the broadcast operation method are greatly improved.
[0038] The following is combined with Figures 1 - 11Describe the broadcast operation method, device, equipment and medium in the artificial intelligence operator of the present invention. The execution subject of the broadcast operation method in the artificial intelligence operator is an artificial intelligence chip, which is also called an artificial intelligence accelerator or a computing card, that is, a module specifically used to process a large number of computing tasks in artificial intelligence applications, such as GPU (Graphics Processing Unit), TPU (Tensor Processing Unit), NPU (Neural network Processing Unit), DPU (Deep learning Processing Unit), APU (Accelerated Processing Unit), and GPGPU (General-Purpose computing on Graphics Processing Unit), etc. Further, the execution subject of the broadcast operation method in the artificial intelligence operator can also be applied to the broadcast operation device in the artificial intelligence operator provided in the artificial intelligence chip. The broadcast operation device in the artificial intelligence operator can be implemented by software, hardware or a combination of both. Here, taking the execution subject of the broadcast operation method in the artificial intelligence operator as an artificial intelligence chip as an example, the broadcast operation method in the artificial intelligence operator will be described.
[0039] Referring to Figure 1 , which is one of the flow diagrams of the broadcast operation method in the artificial intelligence operator provided by the present invention. As Figure 1 shown, the broadcast operation method in the artificial intelligence operator includes steps 110 to 130.
[0040] Step 110: In response to a broadcast operation instruction, obtain the stride parameters and the total dimension of different tensors in each dimension in the broadcast operator; each stride parameter and the total dimension are provided by the main control processor.
[0041] Among them, the broadcast operation instruction can be an instruction automatically generated when the artificial intelligence operator needs to perform a broadcast operation on a tensor during the implementation of its algorithm logic.
[0042] The total dimension can be the dimension of each tensor among different tensors, such as all 3D, all 4D, all 5D or all higher dimensions, etc.
[0043] The main control processor can be a processor that at least has functions such as obtaining tensor dimension information, multiplication operation, tensor dimension calculation, and loop iteration operation. For example, it can be a main control chip, a main control module, a Central Processing Unit (CPU), an Advanced RISC Machine (ARM) processor, an X86 processor, etc.
[0044] The different tensors involved in the broadcast operation usually include three tensors, among which two tensors are from the broadcast sender and one tensor is from the broadcast receiver. The dimension information and element information of the two tensors of the broadcast sender are both known. Each dimension information of the one tensor of the broadcast receiver is determined based on the dimension information of the two tensors of the broadcast sender, and the dimension information of the one tensor of the broadcast receiver is determined by the respective dimensions of the two tensors of the broadcast sender.
[0045] Exemplarily, when the dimension information of the two tensors of the broadcast sender is m×1 dimension and 1×n respectively, the dimension information of the one tensor of the broadcast receiver can be deduced as m×n.
[0046] For the two tensors of the broadcast sender and the one tensor of the broadcast receiver, multiple stride parameters of each of them need to be determined in the main control processor at the front end. Each stride parameter can be recorded as a stride parameter and represents the number of elements inserted in the memory between two consecutive elements (such as the two consecutive elements 0 and 1) in a certain dimension of a certain tensor of the broadcast sender or the one tensor of the broadcast receiver.
[0047] The process of determining the stride parameters for different tensors in each dimension is executed in the main control processor at the front end. The purpose is to assist the artificial intelligence chip at the back end to perform the broadcast operation quickly and accurately without the need to re-develop the broadcast algorithm program.
[0048] Specifically, when the artificial intelligence chip responds to the broadcast operation instruction, it first obtains the stride parameters and the total dimension of different tensors in each dimension in the broadcast operator. The obtaining method can be directly reading from its memory when the above information is pre-stored in the memory of the artificial intelligence chip; or, it can also be called from the cloud server when the above information is pre-stored in the cloud server; or, the artificial intelligence chip can also transmit the parameter dimension obtaining instruction carrying the two tensors of the broadcast sender to the main control processor at the front end through communication transmission or copying, so that the main control processor calculates the stride parameter information for each dimension of the two tensors of the broadcast sender, and calculates the dimension deduction and the stride parameter information for each dimension of the data of the one tensor of the broadcast receiver, and feeds back the calculation results and the total dimension to the artificial intelligence chip. The present invention does not make specific limitations on this.
[0049] Step 120: Perform offset calculations respectively based on each step size parameter and the total dimension to obtain the target offsets corresponding to the first tensor and the second tensor in different tensors respectively.
[0050] Among them, both the first tensor and the second tensor belong to the broadcast sender.
[0051] Specifically, for each tensor in different tensors, the number of step size parameters it contains is the same as the dimension of the corresponding tensor. For example, when the dimension of the tensor is 3D, it correspondingly contains 3 step size parameters; on this basis, the offset corresponding to the next dimension can be calculated according to the step size parameter of the current dimension in the first tensor or the second tensor and the preset initial offset value, and the offset corresponding to the next dimension is used as the new initial offset value, and the step of calculating the offset corresponding to the next dimension is repeated; until the respective target offsets that meet the preset conditions are obtained.
[0052] Step 130: Obtain the first element in the first tensor and the second element in the second tensor respectively based on each target offset, and perform a preset operation on the first element and the second element, and save the operation result to the third tensor in different tensors.
[0053] Among them, the preset operation may include at least one of other operations such as addition operation, subtraction operation, multiplication operation, division operation, maximum value operation, modulo operation, and integer operation.
[0054] Specifically, in the case of obtaining the target offset corresponding to the first tensor and the target offset of the second tensor, the first element can be obtained from the first tensor based on the target offset corresponding to the first tensor, and the second element can be obtained from the second tensor based on the target offset corresponding to the second tensor. After performing the preset operation on the first element and the second element, the operation result is saved to the third tensor of the broadcast receiver.
[0055] Exemplarily, referring to Figure 2 the schematic diagram of the broadcast operation shown, as Figure 2 shown, if it is necessary to perform a broadcast operation on an N-dimensional tensor A and a tensor B to obtain an N-dimensional tensor C, and the dimension information of tensor A is [a1, a2, a3, …, a N , the dimension information of tensor B is [b1, b2, b3, …, b N , and the dimension information of tensor C is [c1, c2, c3, …, c N ; if the x-th dimension of tensor B needs to be broadcast to the corresponding dimension of tensor A, then b x is equal to 1, and the dimension of this dimension of tensor C is equal to the dimension of this dimension of tensor A, that is, c x is equal to a x; By broadcasting the elements [5] in this dimension of tensor B to the elements [6 2 10 4] in this dimension of tensor A, the result of the broadcast operation shown as Figure 2 is [11 7 15 9].
[0056] It should be noted that for the acquisition methods of the first element and the second element, it can be read in combination with the corresponding target offsets when the first tensor and the second tensor are pre-stored in the artificial intelligence chip; or, when the first tensor and the second tensor are stored in the main control processor or other third-party external storage devices, the first element and the second element can be obtained by sending element access requests carrying different target offsets to them. The present invention does not make specific limitations on this.
[0057] In the broadcast operation method of the artificial intelligence operator provided by the present invention, when the artificial intelligence chip responds to the broadcast operation instruction, it first obtains the step length parameters and the total dimension of different tensors in each dimension of the broadcast operator, and then further performs offset operations based on each step length parameter and the total dimension to obtain the respective target offsets of the first tensor and the second tensor in different tensors. Then, based on each target offset, it obtains the first element in the first tensor and the second element in the second tensor, and performs a preset operation on the first element and the second element, and saves the operation result to the third tensor in different tensors. Since the multiple step length parameters and the total dimension of different tensors involved in the broadcast operation are provided by the main control processor, therefore, by the main control processor at the front end providing the step length parameters and the total dimension of different tensors in the broadcast operation, and the way of performing the broadcast operation based on all step length parameters and the total dimension in the artificial intelligence chip, the broadcast operation result is determined efficiently and accurately. The entire broadcast operation process is simple and easy to implement. It is neither necessary to develop corresponding different broadcast operation programs for different-dimensional array structures, and at the same time, the main control processor provided by the front end can be used to assist the artificial intelligence operator at the back end to quickly and accurately complete the broadcast operation, rather than being limited to the artificial intelligence chip as a hardware chip. Combining the strong parallelism of the front-end + back-end heterogeneous hardware framework, the generality and broadcast efficiency of the broadcast operation method are greatly improved.
[0058] Based on the above Figure 1 shown broadcast operation method of the artificial intelligence operator, in an exemplary embodiment, in step 120, the artificial intelligence chip performs offset operations based on each step length parameter and the total dimension to obtain the respective target offsets of the first tensor and the second tensor in different tensors, which is specifically implemented through the following steps.
[0059] For each dimension in the first target tensor, when the dimension has not reached the total dimension, calculate the offset corresponding to the dimension based on the historical offset of the historical dimension of the dimension and the stride parameter under the dimension; the first target tensor is either the first tensor or the second tensor; increment the dimension by 1 and repeat the above steps until the target offset corresponding to when the dimension reaches the total dimension is calculated.
[0060] Specifically, referring to Figure 3 the offset calculation flow chart shown, as Figure 3 shown, set the dimension i in the first target tensor to start from 1, the maximum value of dimension i is the total dimension N, and set the historical offset of the historical dimension when dimension i is 1 as the initial offset offset0, and offset0 = 0; based on this, start the offset calculation process, and based on the historical offset offset i-1 of the historical dimension i - 1 of dimension i and the stride parameter stride i under dimension i, calculate the offset offset i corresponding to dimension i; increment the value of i by 1 and repeat the above process until the offset offset N corresponding to dimension N is calculated, and use the offset offset N corresponding to dimension N as the target offset corresponding to the first target tensor, that is, obtain the target offsets corresponding to the first tensor and the second tensor respectively.
[0061] Based on the above Figure 1 shown broadcast operation method in the artificial intelligence operator, in an exemplary embodiment, the artificial intelligence chip calculates the offset corresponding to the dimension based on the historical offset of the historical dimension of the dimension and the stride parameter under the dimension, and its specific implementation process is achieved through the following steps.
[0062] When the dimension matches the preset parameter representing the broadcast operation, calculate the offset corresponding to the dimension based on the historical offset, the stride parameter under the dimension, and the index value of the dimension; when the dimension does not match the preset parameter, calculate the offset corresponding to the dimension based on the historical offset, the stride parameter under the dimension, and the coordinates of the element corresponding to the dimension in the third tensor.
[0063] Specifically, continue to refer to Figure 3 the offset calculation flow chart shown. First, determine whether dimension i matches the preset parameter representing the broadcast operation, that is, whether dimension i is equal to the preset parameter. If dimension i is equal to the preset parameter, it is determined that the broadcast operation is performed on dimension i in the first target tensor. At this time, based on the historical offset offset i-1 of the historical dimension i - 1 of dimension i, the stride parameter stride iand the index value index of dimension i in the first target tensor i , calculate the offset offset corresponding to dimension i i , and its specific calculation process is implemented through Equation (1).
[0064] offset i = offset i-1 (1).
[0065] Conversely, if dimension i is not equal to the preset parameter, it is determined that the broadcast operation is not performed on dimension i in the first target tensor. At this time, based on the historical offset offset of historical dimension i-1 of dimension i i-1 , the stride parameter stride under dimension i i and the coordinate coord of dimension i in the third tensor i , calculate the offset offset corresponding to dimension i i , and its specific calculation process is implemented through Equation (2).
[0066] offset i = offset i-1 + stride i × coord i (2).
[0067] Exemplarily, the preset parameter characterizing the broadcast operation can be set in advance, for example, it can take the value of 1.
[0068] Based on the above Figure 1 shown broadcast operation method in the artificial intelligence operator, in an exemplary embodiment, the artificial intelligence chip can calculate the coordinates of each element corresponding to each dimension in the third tensor through the following steps.
[0069] For each dimension in the third tensor, when the dimension is not less than 1, use the data processing thread corresponding to the dimension to perform dimension parameter calculation on the linear index value of the dimension and the stride parameter under the dimension, and obtain the coordinates of the element corresponding to the dimension and the historical linear index value of the historical dimension of the dimension; let the dimension be reduced by 1, and repeat the above steps until all dimensions in the third tensor each obtain the coordinates of the corresponding elements when the dimension is less than 1.
[0070] Specifically, referring to Figure 4 shown coordinate calculation flow chart, as Figure 4 shown, set dimension j in the third tensor to start from the total dimension N, the minimum value of dimension j is 1, and set the maximum linear index value linear_index N corresponding to when dimension j is N as a known quantity set in advance; based on this, start the coordinate calculation process, based on the linear index value linear_index of dimension j jand the stride parameter stride under dimension j j , calculate the coordinate coord of dimension j j and the historical linear index value linear_index of historical dimension j-1 of dimension j j-1 , and its calculation process can be implemented through equations (3) and (4); let the value of j be decreased by 1, and repeat the above process until all coordinates corresponding to each dimension in the third tensor are obtained when the dimension is less than 1, that is, the coordinate coord1 corresponding to element of dimension 1 in the third tensor, the coordinate coord2 corresponding to element of dimension 2, ……, the coordinate coord corresponding to element of dimension N N .
[0071] coord j = linear_index j / stride j (3).
[0072] linear_index j-1 = linear_index j % stride j (4).
[0073] In equations (3) and (4), / represents integer division, and % represents remainder; and the linear index value linear_index of dimension j j can be specifically calculated through equation (5).
[0074] linear_index j = linear_index j-1 + coord j × stride j (5).
[0075] It should be noted that the artificial intelligence chip initiates corresponding data processing threads for data processing according to the total number of elements in the third tensor, and each data processing thread is used to process the coordinates of one element in the third tensor, such as each data processing thread is used to calculate the coordinates of the corresponding element in the corresponding dimension of the third tensor.
[0076] Based on the above Figure 1 broadcast operation method in the artificial intelligence operator shown, in an exemplary embodiment, step 110 is specifically implemented through the following steps.
[0077] Based on the first tensor and the second tensor carried in the broadcast operation instruction, copy the first tensor, the second tensor, the stride parameter and the total dimension after preprocessing in each dimension from the main control processor respectively.
[0078] Specifically, when an AI operator needs to perform a broadcast operation on tensors during the implementation of its algorithm logic, it can automatically generate a broadcast operation instruction. The broadcast operation instruction carries two tensors of the broadcast sender, such as the first tensor and the second tensor, and each dimension information of the third tensor generated by broadcasting the data of a certain dimension in the first tensor to the second tensor and / or broadcasting the data of a certain dimension in the second tensor to the first tensor. Each dimension information of the third tensor can be deduced from the first tensor and the second tensor.
[0079] Exemplarily, when the dimension information of the first tensor (such as tensor A) is m×1 and the dimension information of the second tensor (such as tensor B) is 1×n, the second dimension of tensor A can be broadcast to tensor B, and the first dimension of tensor B can be broadcast to tensor A. In this way, the dimension information of the third tensor (such as tensor C) is deduced to be m×n.
[0080] In addition, the process of deducing the dimension information of the third tensor from the first tensor and the second tensor, and the preprocessing process of preprocessing the first tensor, the second tensor, and the third tensor respectively to obtain the stride parameters in each dimension are all executed on the main control processor.
[0081] The main control processor can calculate the stride parameters of the above three tensors in each dimension based on the stride parameter calculation requirements of the AI chip, and give the first tensor, the second tensor, and the third tensor their respective multiple stride parameters and the total dimension N to the AI chip by means of copying.
[0082] Based on the above Figure 1 In an example embodiment of the broadcast operation method in the AI operator shown, it should be noted that for the main control processor, the main control processor is used to perform the following steps for each dimension in the second target tensor.
[0083] When the dimension has not reached the total dimension, calculate the stride parameter of the dimension based on the historical stride parameter of the historical dimension of the dimension and the dimension value of the historical dimension; increment the dimension by 1, and repeat the above steps until the stride parameters of each dimension in the second target tensor are calculated; where the second target tensor is any one of the first tensor, the second tensor, and the third tensor.
[0084] Specifically, referring to Figure 5 , the flowchart of the stride parameter calculation provided by the present invention, as Figure 5As shown, set the dimension k in the second target tensor to start from 2, the maximum value of dimension k is the total dimension N, and set the historical step parameter of the historical dimension corresponding to when dimension k is 2 to the initial value of the step parameter stride1, and set the initial value of the step parameter as a known quantity, such as stride1 = 1; Based on this, start the step parameter calculation process, based on the historical step parameter stride of the historical dimension k - 1 of dimension k k-1 and the dimension value Dimension of the historical dimension k - 1 k-1 , calculate the step parameter stride under dimension k k , which can be specifically implemented by Equation (6).
[0085] stride k = stride k-1 × Dimension k-1 (6).
[0086] Let the value of k increase by 1, and repeat the above process until the step parameters under each dimension in the second target tensor are calculated, that is, obtain the step parameter stride1 under dimension 1 in the second target tensor, the step parameter stride2 under dimension 2, ……, the step parameter stride under dimension N N .
[0087] Based on the above Figure 1 shown broadcast operation method in the artificial intelligence operator, in an exemplary embodiment, in step 130, the artificial intelligence chip obtains the first element in the first tensor and the second element in the second tensor based on each target offset respectively, which is specifically implemented through the following steps.
[0088] Perform memory access based on the target offset corresponding to the first tensor and the starting address of the first tensor in memory to obtain the first element in the first tensor; perform memory access based on the target offset corresponding to the second tensor and the starting address of the second tensor in memory to obtain the second element in the second tensor.
[0089] Specifically, the storage method of the artificial intelligence chip for multi-dimensional data structures is usually linear. When indexing an element in a multi-dimensional data structure in memory, first calculate the offset of the element in memory, and then combine the memory starting address for memory access to facilitate reading the element. Based on this, when both the first tensor and the second tensor are stored linearly in the memory of the artificial intelligence chip, memory access can be performed respectively according to the target offset corresponding to the first tensor and the starting address of the first tensor in memory, and according to the target offset corresponding to the second tensor and the starting address of the second tensor in memory, to obtain the first element in the first tensor and the second element in the second tensor.
[0090] Exemplarily, referring toFigure 6 , which is the second flowchart of the broadcast operation method in the artificial intelligence operator provided by the present invention. As Figure 6 shown, the main control processor can specifically be a central processing unit. The central processing unit calculates the stride of tensor A (i.e., the first tensor), the stride of tensor B (i.e., the second tensor), and the stride of tensor C (i.e., the third tensor), that is, calculates the step length parameters of the first tensor, the second tensor, and the third tensor in each dimension respectively, and then copies the respective strides and the total dimension of the above three tensors to the artificial intelligence chip. The artificial intelligence chip calculates the coordinates of the elements in tensor C corresponding to each data processing thread accordingly, then calculates the respective target offsets of tensor A and tensor B corresponding to each element in tensor C according to the coordinates of each element in tensor C, and reads the first element in tensor A and the second element in tensor B from the memory according to the respective target offsets for calculation and data writing back, that is, performs a preset operation according to the first element and the second element, and saves the operation result to tensor C. The specific process involved can refer to the foregoing embodiments. It will not be elaborated here.
[0091] Specifically, when the total dimensions of tensor A, tensor B, and tensor C are all 3, and the dimension of tensor A is [2, 1, 4], the dimension of tensor B is [2, 3, 1], and the dimension of tensor C is [2, 3, 4], the relevant information of tensor A, tensor B, and tensor C respectively can refer to Figure 7 , Figure 8 and Figure 9 shown. The relevant information here can specifically include each linear index value (linear idx), one-dimensional arrangement, and multi-dimensional coordinates of each element of tensor A, tensor B, or tensor C. It should be noted here that the dimension [2, 3, 4] of tensor C is deduced from the dimensions of tensor A and tensor B.
[0092] At this time, in combination with the foregoing step length parameter calculation method (such as the step length parameter calculation flowchart shown in Figure 5 ), the step length parameters of tensor A in 3 dimensions are calculated as [1, 2, 2], the step length parameters of tensor B in 3 dimensions are [1, 2, 6], and the step length parameters of tensor C in 3 dimensions are [1, 2, 6]. Further according to the above step length parameters, the multi-dimensional coordinates of the corresponding elements can be calculated from the corresponding linear index values, or the corresponding linear index values can be calculated according to the multi-dimensional coordinates of the corresponding elements.
[0093] For example, taking the element with the linear index value linear_index of 16 in tensor C as an example:
[0094] The coordinate of the third dimension coord3 = linear_index3 / stride3 = 16 / 6 = 2;
[0095] linear_index2 = linear_index3 % stride3 = 16 % 6 = 4;
[0096] The coordinate coord2 of the second dimension = linear_index2 / stride2 = 4 / 2 = 2;
[0097] linear_index1 = linear_index2 % stride2 = 4 % 2 = 2;
[0098] The coordinate coord1 of the first dimension = linear_index1 / stride1 = 0 / 1 = 0.
[0099] For another example, taking the calculation of the corresponding linear index value according to multi - dimensional coordinates as an example:
[0100] The initial linear index value linear_index0 = 0;
[0101] The first linear index value linear_index1 = linear_index0 + coord1 × stride1 = 0 + 0 × 1 = 0;
[0102] The second linear index value linear_index2 = linear_index1 + coord2 × stride2 = 0 + 2 × 2 = 4;
[0103] The next linear index value linear_index3 = linear_index2 + coord3 × stride3 = 4 + 2 × 6 = 16.
[0104] Therefore, the linear index value of the coordinate (0, 2, 2) in tensor C is 16.
[0105] Furthermore, since tensor C contains 24 elements, a total of 24 data - processing threads are initiated. Each data - processing thread is used to calculate the coordinates of one element in tensor C; each data - processing thread has a unique number, and each number belongs to the range greater than or equal to 0 and less than or equal to 23.
[0106] Taking the data - processing thread numbered 16 as an example, the data - processing thread can calculate the coordinates of the corresponding element in tensor C as coord(0, 2, 2) based on the stride parameters of the 3 dimensions of tensor C.
[0107] Then, setting the initial offset offset0 = 0, the process of calculating the target offset of tensor A is as follows:
[0108] For the first dimension, no broadcasting operation is performed: offset1 = offset0 + coord1 × stride1 = 0 + 0 × 1 = 0;
[0109] For the second dimension, broadcasting operation is performed: offset2 = offset1 + index2 × stride2 = 0 + 0 × 2 = 0;
[0110] For the third dimension, no broadcasting operation is performed: offset3 = offset2 + coord3 × stride3 = 0 + 2 × 2 = 4.
[0111] Therefore, the target offset of tensor A is 4, and the first element read from tensor A is 5.
[0112] Similarly, setting the initial offset offset0 = 0, the process of calculating the target offset of tensor B is as follows:
[0113] For the first dimension, no broadcasting operation is performed: offset1 = offset0 + coord1 × stride1 = 0 + 0 × 1 = 0;
[0114] For the second dimension, no broadcasting operation is performed: offset2 = offset1 + coord2 × stride2 = 0 + 2 × 2 = 4;
[0115] For the third dimension, broadcasting operation is performed: offset3 = offset2 + index3 × stride3 = 4 + 0 × 6 = 4.
[0116] Therefore, the target offset of tensor B is 4, and the second element read from tensor B is 5.
[0117] Through the broadcasting operation method in the artificial intelligence operator provided by the present invention, the same broadcasting operation program can be used to implement tensor operations in any dimension, support broadcasting of any multi-dimensional array, and combined with the strong parallelism of the front-end + back-end heterogeneous hardware framework, it can also greatly improve the generality and broadcasting efficiency of the broadcasting operation method.
[0118] Next, the broadcasting operation device in the artificial intelligence operator provided by the present invention will be described. The broadcasting operation device in the artificial intelligence operator described below can be mutually corresponding and referred to the broadcasting operation method in the artificial intelligence operator described above.
[0119] Referring to Figure 10 , which is a schematic structural diagram of the broadcasting operation device in the artificial intelligence operator provided by the present invention. As Figure 10 shown, the broadcasting operation device 1000 in the artificial intelligence operator is applied to an artificial intelligence chip and includes: a parameter acquisition unit 1010 and a broadcasting operation unit 1020.
[0120] A parameter acquisition unit 1010, configured to obtain the stride parameters and the total dimension of different tensors in each dimension of a broadcast operator in response to a broadcast operation instruction; the stride parameters and the total dimension are provided by a main control processor.
[0121] A broadcast operation unit 1020, configured to perform offset calculations respectively based on the stride parameters and the total dimension to obtain the target offsets corresponding to a first tensor and a second tensor in different tensors; obtain a first element in the first tensor and a second element in the second tensor respectively based on the target offsets, and perform a preset operation on the first element and the second element, and save the operation result to a third tensor in the different tensors.
[0122] Optionally, the broadcast operation unit 1020 is specifically configured to, for each dimension in a first target tensor, calculate the offset corresponding to the dimension based on the historical offset of the historical dimension of the dimension and the stride parameter of the dimension when the dimension does not reach the total dimension; the first target tensor is any one of the first tensor and the second tensor; increment the dimension by 1, and repeat the above steps until the target offset corresponding to when the dimension reaches the total dimension is calculated.
[0123] Optionally, the broadcast operation unit 1020 is specifically configured to calculate the offset corresponding to the dimension based on the historical offset, the stride parameter of the dimension, and the index value of the dimension when the dimension matches a preset parameter representing a broadcast operation; calculate the offset corresponding to the dimension based on the historical offset, the stride parameter of the dimension, and the coordinates of the element corresponding to the dimension in the third tensor when the dimension does not match the preset parameter.
[0124] Optionally, the broadcast operation unit 1020 is specifically configured to, for each dimension in the third tensor, when the dimension is not less than 1, use a data processing thread corresponding to the dimension to perform dimension parameter calculations on the linear index value of the dimension and the stride parameter of the dimension to obtain the coordinates of the element corresponding to the dimension and the historical linear index value of the historical dimension of the dimension; decrement the dimension by 1, and repeat the above steps until the coordinates of the elements corresponding to all dimensions in the third tensor are obtained when the dimension is less than 1.
[0125] Optionally, the parameter acquisition unit 1010 is specifically configured to copy the stride parameters and the total dimension of the first tensor, the second tensor, and the third tensor respectively after preprocessing in each dimension from the main control processor based on the first tensor and the second tensor carried in the broadcast operation instruction.
[0126] Optionally, the main control processor in the parameter acquisition unit 1010 is configured to perform the following steps for each dimension in the second target tensor: when the dimension has not reached the total dimension, calculate the step parameter for the dimension based on the historical step parameter and the dimension value of the historical dimension of the dimension; increment the dimension by 1, and repeat the above steps until the step parameters for each dimension in the second target tensor are calculated; where the second target tensor is any one of the first tensor, the second tensor, and the third tensor.
[0127] Optionally, the broadcast operation unit 1020 is specifically configured to perform memory access based on the target offset corresponding to the first tensor and the starting address of the first tensor in the memory to obtain the first element in the first tensor; perform memory access based on the target offset corresponding to the second tensor and the starting address of the second tensor in the memory to obtain the second element in the second tensor.
[0128] Figure 11 An example of the physical structure diagram of an electronic device is shown in Figure 11 As shown, the electronic device may include: an Artificial Intelligence (AI) chip 1110, a communication interface 1120, a memory 1130, a communication bus 1140, and a main control processor connected to the AI chip 1110 in a heterogeneous framework form. Among them, the AI chip 1110, the communication interface 1120, and the memory 1130 complete mutual communication through the communication bus 1140. The AI chip 1110 can call the logical instructions in the memory 1130 to execute the broadcast operation method in the artificial intelligence operator. The method includes: in response to a broadcast operation instruction, obtaining the step parameters and the total dimension for each dimension of different tensors in the broadcast operator; the step parameters and the total dimension are provided by the main control processor; performing offset calculation based on each step parameter and the total dimension to obtain the target offsets corresponding to the first tensor and the second tensor in different tensors respectively; obtaining the first element in the first tensor and the second element in the second tensor based on each target offset respectively, and performing a preset operation on the first element and the second element, and saving the operation result to the third tensor in different tensors; the main control processor is used to provide each step parameter and the total dimension to the AI chip 1110.
[0129] In addition, when the logical instructions in the above-mentioned memory 1130 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs, Read-Only Memories), random access memories (RAMs, Random Access Memories), magnetic disks, or optical discs that can store program codes.
[0130] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the broadcast operation method in the artificial intelligence operators provided by the above-mentioned various methods. The method includes: in response to a broadcast operation instruction, obtaining the stride parameters and the total number of dimensions of different tensors in the broadcast operator in each dimension; the stride parameters and the total number of dimensions are provided by the main control processor; performing offset calculations respectively based on the stride parameters and the total number of dimensions to obtain the target offsets corresponding to the first tensor and the second tensor in different tensors; respectively obtaining the first element in the first tensor and the second element in the second tensor based on the target offsets, and performing a preset operation on the first element and the second element, and saving the operation result to the third tensor in different tensors.
[0131] In yet another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the broadcast operation method in the artificial intelligence operators provided by the above-mentioned various methods. The method includes: in response to a broadcast operation instruction, obtaining the stride parameters and the total number of dimensions of different tensors in the broadcast operator in each dimension; the stride parameters and the total number of dimensions are provided by the main control processor; performing offset calculations respectively based on the stride parameters and the total number of dimensions to obtain the target offsets corresponding to the first tensor and the second tensor in different tensors; respectively obtaining the first element in the first tensor and the second element in the second tensor based on the target offsets, and performing a preset operation on the first element and the second element, and saving the operation result to the third tensor in different tensors.
[0132] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0133] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0134] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. However, these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A broadcast operation method in an artificial intelligence operator, characterized in that, Applied to an artificial intelligence chip, the method includes: In response to a broadcast operation instruction, obtaining the stride parameters and the total dimension of different tensors in each dimension in a broadcast operator; each of the stride parameters and the total dimension is provided by a main control processor; the main control processor is connected to the artificial intelligence chip in the form of a heterogeneous framework; the broadcast operation instruction is an instruction automatically generated when an artificial intelligence operator needs to perform a broadcast operation on a tensor during the implementation of its algorithm logic; the total dimension is the dimension of each tensor in different tensors; each of the stride parameters represents the number of elements inserted in memory between two consecutive elements in a certain dimension in any tensor of a broadcast sender or one tensor of a broadcast receiver; Performing offset calculations respectively based on each of the stride parameters and the total dimension to obtain the target offsets corresponding to the first tensor and the second tensor in the different tensors; Accessing memory based on each of the target offsets, respectively obtaining the first element in the first tensor and the second element in the second tensor, and performing a preset operation on the first element and the second element, and writing the operation result correspondingly into the memory corresponding to the third tensor in the different tensors; the first tensor and the second tensor both belong to the broadcast sender; the third tensor belongs to the broadcast receiver; The accessing memory based on each of the target offsets to respectively obtain the first element in the first tensor and the second element in the second tensor includes: Performing memory access based on the target offset corresponding to the first tensor and the starting address of the first tensor in memory to obtain the first element in the first tensor; Performing memory access based on the target offset corresponding to the second tensor and the starting address of the second tensor in the memory to obtain the second element in the second tensor.
2. The broadcast operation method in the artificial intelligence operator according to claim 1, characterized in that, The performing offset calculations respectively based on each of the stride parameters and the total dimension to obtain the target offsets corresponding to the first tensor and the second tensor in the different tensors includes: For each dimension in the first target tensor, when the dimension does not reach the total dimension, calculating the offset corresponding to the dimension based on the historical offset of the historical dimension of the dimension and the stride parameter in the dimension; the first target tensor is any one of the first tensor and the second tensor; Incrementing the dimension by 1, and repeating the above steps until the target offset corresponding to when the dimension reaches the total dimension is calculated.
3. The broadcast operation method in the artificial intelligence operator according to claim 2, wherein The calculating the offset corresponding to the dimension based on the historical offset of the historical dimension of the dimension and the stride parameter in the dimension includes: When the dimension matches a preset parameter characterizing the broadcast operation, calculating the offset corresponding to the dimension based on the historical offset, the stride parameter in the dimension, and the index value of the dimension; the dimension matching the preset parameter characterizing the broadcast operation is used to characterize that the dimension is equal to the preset parameter and the dimension performs a broadcast operation; In the case where the dimension does not match the preset parameter, calculate the offset corresponding to the dimension based on the historical offset, the step parameter in the dimension, and the coordinates of the elements corresponding to the dimension in the third tensor.
4. The broadcast operation method in the artificial intelligence operator according to claim 3, wherein The method further includes: For each dimension in the third tensor, when the dimension is not less than 1, use the data processing thread corresponding to the dimension to perform dimension parameter calculation on the linear index value of the dimension and the step parameter in the dimension, to obtain the coordinates of the elements corresponding to the dimension and the historical linear index value of the historical dimension of the dimension; Decrease the dimension by 1, and repeat the above steps until the coordinates of the elements corresponding to all dimensions in the third tensor are obtained when the dimension is less than 1.
5. The broadcast operation method in the artificial intelligence operator according to any one of claims 1 to 4, characterized in that, The obtaining, in response to a broadcast operation instruction, of the step parameters and the total dimension of different tensors in each dimension in the broadcast operator includes: Based on the first tensor and the second tensor carried in the broadcast operation instruction, copy from the main control processor the step parameters and the total dimension of the first tensor, the second tensor, and the third tensor respectively after preprocessing in each dimension.
6. The broadcast operation method in the artificial intelligence operator according to claim 5, characterized in that, The main control processor is configured to perform the following steps for each dimension in the second target tensor: When the dimension has not reached the total dimension, calculate the step parameter in the dimension based on the historical step parameter of the historical dimension of the dimension and the dimension value of the historical dimension; Increase the dimension by 1, and repeat the above steps until the step parameters in each dimension of the second target tensor are calculated; where the second target tensor is any one of the first tensor, the second tensor, and the third tensor.
7. A broadcast operation device in an artificial intelligence operator, characterized in that, Applied to an artificial intelligence chip, the device includes: A parameter acquisition unit, configured to obtain the step parameters and the total dimension of different tensors in each dimension in the broadcast operator in response to a broadcast operation instruction; each of the step parameters and the total dimension is provided by a main control processor; the main control processor is connected to the artificial intelligence chip in the form of a heterogeneous framework; the broadcast operation instruction is an instruction automatically generated when an artificial intelligence operator needs to perform a broadcast operation on a tensor during the implementation of its algorithm logic; the total dimension is the dimension of each tensor among different tensors; each of the step parameters represents the number of elements inserted in the memory between two consecutive elements in a certain dimension of any tensor of the broadcast sender or one tensor of the broadcast receiver; A broadcast operation unit, configured to perform offset calculation respectively based on each of the step parameters and the total dimension to obtain the target offsets corresponding to the first tensor and the second tensor in the different tensors; access the memory based on each of the target offsets to respectively obtain a first element in the first tensor and a second element in the second tensor, and perform a preset operation on the first element and the second element, and write the operation result correspondingly into the memory corresponding to the third tensor in the different tensors; the first tensor and the second tensor both belong to the broadcast sender; the third tensor belongs to the broadcast receiver; The broadcast operation unit is specifically configured to: perform memory access based on the target offset corresponding to the first tensor and the starting address of the first tensor in the memory to obtain the first element in the first tensor; perform memory access based on the target offset corresponding to the second tensor and the starting address of the second tensor in the memory to obtain the second element in the second tensor.
8. An electronic device, characterized in that, It includes a main control processor and an artificial intelligence chip connected in the form of a heterogeneous framework. The artificial intelligence chip is used to, in response to a broadcast operation instruction, obtain the stride parameters and the total dimension of different tensors in each dimension in the broadcast operator. The broadcast operation instruction is an instruction automatically generated when an artificial intelligence operator needs to perform a broadcast operation on a tensor during the implementation of its algorithm logic. The total dimension is the dimension of each tensor among different tensors. Each of the stride parameters represents the number of elements inserted in the memory between two consecutive elements in a certain dimension in any tensor of the broadcast sender or one tensor of the broadcast receiver. Offset calculations are respectively performed based on each of the stride parameters and the total dimension to obtain the target offsets corresponding to the first tensor and the second tensor among the different tensors. Access the memory based on each of the target offsets to respectively obtain the first element in the first tensor and the second element in the second tensor, perform a preset operation on the first element and the second element, and write the operation result correspondingly into the memory corresponding to the third tensor among the different tensors. The first tensor and the second tensor both belong to the broadcast sender. The third tensor belongs to the broadcast receiver and is further configured to: perform memory access based on the target offset corresponding to the first tensor and the starting address of the first tensor in the memory to obtain the first element in the first tensor. Perform memory access based on the target offset corresponding to the second tensor and the starting address of the second tensor in the memory to obtain the second element in the second tensor. The main control processor is used to provide each of the stride parameters and the total dimension to the artificial intelligence chip.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the broadcast operation method in the artificial intelligence operator according to any one of claims 1 to 6.
Citation Information
Patent Citations
Data processing method and device, terminal equipment and computer readable storage medium
CN114491399A
Data processing method and device, electronic equipment and storage medium
CN116822612A