Mean operator execution method, chip, device, medium and program product

By designing parallel execution units on artificial intelligence chips and utilizing on-chip caches, the problem of low execution efficiency of mean operators in the prior art is solved, and more efficient computing performance is achieved.

CN118505486BActive Publication Date: 2025-05-23SHANGHAI BIREN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410579428.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-10
Publication Date
2025-05-23
Estimated Expiration
2044-05-10

AI Technical Summary

Technical Problem

When using artificial intelligence chips to execute mean operators, the prior art cannot effectively exert the computing power performance of the chip, resulting in low execution efficiency of the operator.

Method used

By designing an architecture of N parallel execution units on an artificial intelligence chip, data exchange and accumulation are used to use on-chip cache to achieve parallel processing of multiple pending data.

Benefits of technology

The computing efficiency of the mean operator is improved and the parallel computing power of artificial intelligence chips is fully utilized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118505486B_ABST
    Figure CN118505486B_ABST
Patent Text Reader

Abstract

The embodiment of the present invention provides an execution method, chip, device, medium and program product of a mean operator, so as to improve the execution efficiency of the mean operator. The method includes: N execution units execute a first operation in parallel; the first operation executed by one execution unit includes: loading multiple data to be processed in an input image from a video memory, executing a first operation on the multiple data to be processed, and writing the first operation result to an on-chip cache, and accumulating the first operation result with the execution results of the first operation written to the on-chip cache by other N-1 execution units among the N execution units to obtain a first accumulated result; the reference execution unit among the N execution units loads the first accumulated result from the on-chip cache, determines a first mean result according to the first accumulated result, and writes the first mean result to the video memory, so as to effectively exert the parallel computing capability of the chip and improve the execution efficiency of the mean operator.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to the field of artificial intelligence technology, and in particular to an execution method, chip, device, medium, and program product of a mean operator. Background Art

[0002] Artificial intelligence models usually refer to neural network models that are trained to perform inference and prediction, such as image inference models, speech inference models, etc. In practical applications, artificial intelligence chips are used to implement the operations of artificial intelligence models. The operations of artificial intelligence models can be implemented by operators in the computation graph. The computation graph is a multi-graph structure used to represent the computational tasks and data flow processes of artificial intelligence models. Operators refer to various operations performed on tensors at each layer in the artificial intelligence model.

[0003] For example, the mean operator used in the field of image processing calculates the mean of the entire input image, and outputs an average value for each image channel in the input image. In related technologies, the mean operation of image data of a channel in the input image is performed at the granularity of threads or thread blocks, which cannot effectively exert the computing power of artificial intelligence chips, resulting in low execution efficiency of the operator. Summary of the invention

[0004] The embodiments of the present invention provide an execution method of a mean operator, an artificial intelligence chip, and an electronic device to effectively exert the computing power performance of the artificial intelligence chip and improve the execution efficiency of the mean operator.

[0005] In a first aspect, the present application provides an execution method of a mean operator, which is applied to an artificial intelligence chip, wherein the artificial intelligence chip includes N execution units, a video memory, and an on-chip cache, where N is a positive integer; the method includes: N execution units execute a first operation in parallel; wherein the first operation executed by one execution unit includes: loading multiple to-be-processed data in an input image from the video memory, performing a first operation on the multiple to-be-processed data to obtain a first operation result, and the multiple to-be-processed data loaded by different execution units in the N execution units are different; writing the first operation result into the on-chip cache, and accumulating the first operation result with the execution results of the first operation written into the on-chip cache by other N-1 execution units in the N execution units; a reference execution unit in the N execution units loads a first accumulated result from the on-chip cache, the first accumulated result includes an accumulated result of the execution results of the first operation corresponding to the N execution units respectively, and the reference execution unit is an execution unit as a core among the N execution units; the reference execution unit determines a first mean result according to the first accumulated result, and writes the first mean result out to the video memory.

[0006] Through the above scheme, N execution units execute the first operation in parallel, wherein the first operation executed by each execution unit includes: executing a first operation on the loaded multiple data to be processed, and writing the execution result of the first operation into the on-chip cache, and the execution results of the first operation corresponding to the N execution units can be accumulated in the on-chip cache, so that data exchange between execution units can be achieved through the on-chip cache, and the N execution units process the data in the input image in parallel, which can effectively exert the parallel computing capability of the artificial intelligence chip and improve the computing efficiency of the mean operator.

[0007] In a possible implementation, a plurality of data to be processed include first channel values ​​respectively corresponding to a plurality of pixel points in an input image, each execution unit in the N execution units includes K threads, and each thread corresponds to a register; the plurality of data to be processed in the input image are loaded from a video memory, a first operation is performed on the plurality of data to be processed, and a first operation result is obtained, including: the K threads in the one execution unit perform a second operation in parallel; wherein the second operation performed by one thread includes: loading part of the first channel values ​​in the plurality of data to be processed from the video memory into a register corresponding to the one thread; performing a first operation on part of the first channel values ​​in a register corresponding to the one thread, and loading the result of the first operation performed by the one thread from the register corresponding to the one thread to the register corresponding to the reference thread; the reference thread is a core thread among the K threads in the one execution unit, and the execution results of the first operation respectively corresponding to the K threads are accumulated in the register corresponding to the reference thread to obtain the first operation result.

[0008] In one possible implementation, N execution units are divided into M execution groups, and the on-chip cache includes group shared memories corresponding to the M execution groups respectively; a first operation result is written into the on-chip cache, and the first operation result is accumulated with the execution results of the first operation written into the on-chip cache by other N-1 execution units among the N execution units, including: writing the first operation result into the target group shared memory corresponding to the target execution group to which an execution unit belongs; accumulating the first operation result with the execution results of the first operation written into the target group shared memory by other execution units except one execution unit in the target execution group to obtain a second accumulated result.

[0009] In a possible implementation, the on-chip cache also includes a secondary cache; after accumulating the first operation result with the execution result of the first operation written into the target group shared memory by other execution units except one execution unit in the target execution group to obtain a second accumulated result, it also includes: the first execution unit in the target execution group loads the second accumulated result from the target group shared memory, the first execution unit being the core execution unit in the target execution group; the first execution unit writes the loaded second accumulated result into the secondary cache, and accumulates the second accumulated result with the second accumulated results corresponding to the other M-1 execution groups in the M execution groups written into the secondary cache; the result obtained by accumulating the second accumulated results corresponding to the M execution groups is used as the first accumulated result; the reference execution unit among the N execution units loads the first accumulated result from the on-chip cache, including: the reference execution unit among the N execution units loads the first accumulated result from the secondary cache.

[0010] In a possible implementation, a plurality of data to be processed include first channel values ​​corresponding to a plurality of pixel points in an input image, respectively, and a first operation is a sum operation; performing the first operation on the plurality of data to be processed to obtain a first operation result comprises: performing a sum operation on all first channel values ​​included in the plurality of data to be processed to obtain a first operation result; and determining a first mean result based on the first accumulation result by a reference execution unit comprises: the reference execution unit performs a mean operation on the first accumulation result based on the number of first channel values ​​loaded by N execution units to obtain a first mean result.

[0011] In a possible implementation, a plurality of data to be processed include first channel values ​​corresponding to a plurality of pixel points in an input image, and a first operation is an averaging operation; performing the first operation on the plurality of data to be processed to obtain a first operation result includes: performing an averaging operation on the sum of the first channel values ​​included in the plurality of data to be processed according to the number of the first channel values ​​included in the plurality of data to be processed to obtain a first operation result; a reference execution unit determines a first mean result according to a first accumulation result, including: the reference execution unit performs an averaging operation on the first accumulation result according to N to obtain a first mean result.

[0012] In a possible implementation, an input image includes first channel data, and N data blocks obtained after the first channel data is segmented respectively include multiple different data to be processed; multiple data to be processed in the input image are loaded from a video memory, and a first operation is performed on the multiple data to be processed to obtain a first operation result, including: loading a first data block in the first channel data from the video memory, performing a first operation on the first data block, and obtaining a first operation result; the first data block is any data block among the N data blocks obtained by segmenting the first channel data, and different execution units among the N execution units load different data blocks.

[0013] In a possible implementation, a plurality of data to be processed belong to first channel data in an input image, and the input image also includes second channel data; after the reference execution unit determines a first mean result according to a first accumulation result, and writes the first mean result to a video memory, it further includes: N execution units executing a third operation in parallel; wherein the third operation executed by one execution unit includes: loading a second data block in the second channel data from the video memory, executing a third operation on the second data block, and obtaining a third operation result; the second data block is any one of the N data blocks obtained after the second channel data is segmented, and different execution units in the N execution units load different data blocks; writing the third operation result into an on-chip cache, and accumulating the third operation result with the execution results of the third operation written into the on-chip cache by other N-1 execution units in the N execution units; the reference execution unit in the N execution units loads a third accumulation result from the on-chip cache, and the third accumulation result includes the accumulation result of the execution results of the third operation corresponding to the N execution units respectively; the reference execution unit determines a second mean result according to the third accumulation result, and writes the second mean result to the video memory.

[0014] In a possible implementation, a plurality of data to be processed belong to first channel data in an input image, and the input image also includes second channel data; the artificial intelligence chip also includes F execution units; the method also includes: the F execution units execute a fourth operation in parallel; wherein, the fourth operation executed by one of the F execution units includes: loading a third data block in the second channel data from a video memory, executing a fourth operation on the third data block, and obtaining a fourth operation result; the third data block is any one of the F data blocks obtained by segmenting the second channel data, and different execution units of the F execution units load different data blocks; writing the fourth operation result into an on-chip cache, and accumulating the fourth operation result with the execution results of the fourth operation written into the on-chip cache by the other F-1 execution units of the F execution units; the reference execution unit of the F execution units loads the fourth accumulated result from the on-chip cache, determines a third mean result based on the fourth accumulated result, and writes the third mean result out to the video memory; wherein, the fourth accumulated result includes the accumulated result of the execution results of the fourth operation corresponding to the F execution units respectively.

[0015] In a second aspect, the present application provides an artificial intelligence chip, comprising N execution units, a video memory, and an on-chip cache, where N is a positive integer;

[0016] The N execution units are used to execute a first operation in parallel; wherein the first operation executed by one execution unit includes: loading a plurality of to-be-processed data in the input image from the video memory, executing a first operation on the plurality of to-be-processed data, and obtaining a first operation result; the plurality of to-be-processed data loaded by different execution units in the N execution units are different; writing the first operation result into an on-chip cache, and accumulating the first operation result with the execution results of the first operation written into the on-chip cache by other N-1 execution units in the N execution units;

[0017] The reference execution unit among the N execution units is further used to load a first cumulative result from the on-chip cache, determine a first average result based on the first cumulative result, and write the first average result to the video memory; the first cumulative result includes the cumulative result of the execution results of the first operations respectively corresponding to the N execution units, and the reference execution unit is the core execution unit among the N execution units.

[0018] In a third aspect, the present application provides a computer device, comprising a memory, a processor chip, and a computer program stored in the memory and executable on the processor chip, wherein the processor chip implements the method in the above-mentioned first aspect or any possible implementation manner of the first aspect when executing the computer program.

[0019] In a fourth aspect, the present application provides a computer-readable storage medium, comprising computer-executable instructions. When the computer-executable instructions are executed on a computer, the computer executes the method in the above-mentioned first aspect or any possible implementation manner of the first aspect.

[0020] In a fifth aspect, the present application provides a computer program product, wherein the computer program product stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a computer, the computer executes the method in the above-mentioned first aspect or any possible implementation manner of the first aspect.

[0021] The technical effects that can be achieved in any of the second to fifth aspects mentioned above can refer to the description of the beneficial effects in the first aspect mentioned above or any possible implementation method of the first aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1A and Figure 1B This is a schematic diagram of the structure of an artificial intelligence chip applicable to the embodiments of the present application;

[0023] Figure 2 A schematic diagram of executing a mean operator in the related art;

[0024] Figure 3 A schematic diagram of executing a mean operator in the related art;

[0025] Figure 4 A flowchart of a method for executing a mean operator provided by an embodiment of the present invention;

[0026] Figure 5 A schematic diagram of the RGB image format provided by an embodiment of the present invention;

[0027] Figure 6 A schematic diagram of the NV12 image format provided by an embodiment of the present invention;

[0028] Figure 7 A schematic diagram of executing the mean operator provided in an embodiment of the present invention;

[0029] Figure 8 A flowchart of a method for executing a mean operator provided by an embodiment of the present invention;

[0030] Fig. 9 A flowchart of a method for executing a mean operator provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0031] In order to make the purpose, technical solution and beneficial effects of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0032] Figure 1A A schematic diagram of the structure of an artificial intelligence chip applicable to an embodiment of the present application.

[0033] like Figure 1A As shown, the artificial intelligence chip 100 includes at least a video memory 101, an on-chip cache 102, and a plurality of execution units (EUs) 103. Each execution unit 103 may include a plurality of threads (for example, 40 threads) and registers corresponding to each thread (Thread). The present application does not limit the number of threads included in each execution unit 103.

[0034] The video memory 101 may be a high bandwidth memory (HBM) or other types of memory. The on-chip cache 102 may be a shared cache of multiple execution units or a shared cache of some of the multiple execution units. For example, the on-chip cache 102 is a group shared memory (GSM). For another example, the on-chip cache 102 is a secondary cache (L2 Cache).

[0035] Above Figure 1A The access speeds of the several storage structures mentioned in the figure are from fast to slow: register, on-chip cache 102, and video memory 101.

[0036] Figure 1B This is a schematic diagram of the structure of another artificial intelligence chip applicable to the embodiments of the present application.

[0037] like Figure 1B As shown, the artificial intelligence chip 110 includes at least a video memory 101 and a plurality of programmable multiprocessors 104. The programmable multiprocessor 104 may be a streaming processing cluster (SPC for short), and each programmable multiprocessor 104 includes: a plurality of execution units 103. Each execution unit 103 may include a plurality of threads (for example, 40 threads) and registers corresponding to each thread.

[0038] The multiple execution units 103 in each programmable multiprocessor 104 may be divided into multiple execution groups 105, where the execution group may be a computing unit (CU). Each execution group 105 corresponds to a group shared memory (GSM) 106. Each execution unit 103 in the execution group 105 may use the same group shared memory (GSM) 106, and different execution groups 105 correspond to different group shared memories 106.

[0039] The artificial intelligence chip 110 also includes a second-level cache (L2 Cache) 107 , and multiple execution groups 105 can access the second-level cache 107 .

[0040] The video memory 101 may be a high bandwidth memory (HBM for short) or other types of memory.

[0041] Above Figure 1B The access speeds of the several storage structures mentioned in the figure are as follows from fast to slow: register, group shared memory 106 , secondary cache 107 , and video memory 101 .

[0042] It should be understood that the artificial intelligence chip in this application includes Figure 1A , Figure 1B In addition to the structures shown, other structures may also be included, and this application does not limit this. This application also does not limit the number of computing cores, execution groups, execution units, etc. included in the artificial intelligence chip.

[0043] In the embodiments of the present application, the artificial intelligence chip may be: a graphics processing unit (Graphics Processing Unit, referred to as GPU), a general-purpose graphics processing unit (General-Purpose Computing On Graphics Processing Units, referred to as GPGPU), etc.

[0044] In the field of image processing, the mean operator is used to calculate the mean of the entire input image, and each image channel in the input image outputs an average value. When the artificial intelligence chip executes the mean operator, the channel data of each channel in the image to be processed needs to be loaded from the video memory to the chip, and the sum and mean calculations are performed.

[0045] In one implementation, for example, the image to be processed includes L channels, and the artificial intelligence chip uses L threads to process the channel data corresponding to the L channels in the image to be processed, and one thread processes the channel data corresponding to one channel. Take the example of using Thread0 in EU0 in CU0 in SPC0 to process the first channel data (for example, image array) corresponding to the first channel, as shown in FIG. Figure 2 As shown, a Thread0 is used to load all the data in the first channel data (image array) corresponding to the first channel, and then a thread sum (Thread Sum) operation is performed on all the data in the image array. After that, Thread0 performs a mean (Mean) operation on the result of the sum operation to obtain the mean result of the first channel data corresponding to the first channel, and Thread0 writes the mean result of the first channel to the video memory.

[0046] In another implementation, the artificial intelligence chip uses a thread block to process channel data corresponding to a channel. Take the example of using the thread block in EU0 in CU0 in SPC0 to process the first channel data (imagearray) corresponding to the first channel. For example, the thread block in EU0 includes 40 threads, which are: Figure 3 As shown in Thread0, Thread1, ..., Thread40, each thread loads part of the data in the first channel, and then distributes the thread summation, and then performs the thread block summation, that is, the summation results corresponding to Thread0 to Thread31 are accumulated, and then the accumulated results are averaged to obtain the average result of the first channel data corresponding to the first channel, and Thread0 writes the average result of the first channel to the video memory. This method does not have the problem of communication and data transfer between thread blocks, but it will waste the performance of the unused thread blocks.

[0047] Both of the above two implementation methods cannot effectively bring into play the computing power of artificial intelligence chips, resulting in low execution efficiency of operators.

[0048] In view of this, the present application provides a method for executing a mean operator to effectively utilize the computing power performance of an artificial intelligence chip and improve the execution efficiency of the mean operator. The method for executing the mean operator provided in the present application can be implemented by an operator task execution device. The operator task execution device can be an artificial intelligence chip or a device including an artificial intelligence chip. The following method embodiments are introduced by taking an artificial intelligence chip as the execution subject as an example.

[0049] Figure 4 A flowchart of a method for executing a mean operator provided in an embodiment of the present application. The method for executing the mean operator can be applied to an artificial intelligence chip, which includes N execution units, a video memory, and an on-chip cache, where N is a positive integer, such as Figure 4 As shown, the execution method of the mean operator includes:

[0050] In step 401 , N execution units execute a first operation in parallel; wherein the first operation executed by one execution unit includes the following steps 401 - 1 and 401 - 2 .

[0051] Step 401 - 1 , loading a plurality of data to be processed in an input image from a video memory, and performing a first operation on the plurality of data to be processed to obtain a first operation result.

[0052] When each of the N execution units executes step 401 - 1 , different execution units among the N execution units load different multiple pieces of to-be-processed data.

[0053] The input image may include channel data of one channel, or may include channel data of multiple channels. Take N execution units processing the first channel data in the input image as an example, that is, the multiple data to be processed loaded when the N execution units respectively execute step 401-1 all belong to the first channel data in the input image. The first channel data can be divided into N data blocks, and the N data blocks obtained by dividing the first channel data respectively include different multiple data to be processed. The N execution units perform the first operation on the first channel data in parallel, wherein one execution unit executing step 401-1 can be understood as: one execution unit loads the first data block from the video memory, and performs the first operation on the first data block to obtain the first operation result. The first data block is any one of the N data blocks obtained by dividing the first channel data, and different execution units in the N execution units load different data blocks.

[0054] The first channel data includes first channel values ​​corresponding to all pixels in the input image.

[0055] The input image in the embodiment of the present application may support multiple image formats, including but not limited to GRAY, RGB, BGR, RGBP, BGRP, NV12, NV21, YU12, YV12, RGBA, BGRA, etc. shown in Table 1 below. Input images of different image formats have different numbers of channels. The following is a diagram showing the attribute information of different image formats: the number of channels and the number of planars.

[0056] Table 1

[0057] Image Format Number of channels Number of planes GRAY 1 1 RGB / BGR 3 1 RGBP / BGRP 3 3 NV12 / NV21 3 2 YU12 / YV12 3 3 RGBA / BGRA 4 1 The following describes the number of channels of the input image in step 401 - 1 and the first channel data by taking the image formats such as GRAY, RGB and NV12 in Table 1 as examples.

[0058] Taking the input image in GRAY format as an example, the number of channels is 1, and the first channel data in the input image is the image data of all pixels included in the entire input image.

[0059] Taking the input image in RGB format as an example, the number of channels is 3, namely R channel, G channel and B channel. The first channel data in the above step 401-1 can be the channel data corresponding to the R channel of all pixels included in the input image in RGB format, or the channel data corresponding to the G channel of all pixels included in the input image, or the channel data corresponding to the B channel of all pixels included in the input image. Figure 5 Taking the input image in RGB format as shown, the channel data corresponding to the R channel of all pixels includes 12 R channel values, the channel data corresponding to the G channel of all pixels includes 12 G channel values, and the channel data corresponding to the B channel of all pixels includes 12 B channel values.

[0060] Taking the input image in NV12 format as an example, the number of channels is 3, namely, Y channel, U channel and V channel. The first channel data in the above step 401-1 can be the channel data corresponding to the Y channel of all pixels included in the input image, or the channel data corresponding to the U channel of all pixels included in the input image, or the channel data corresponding to the V channel of all pixels included in the input image. Figure 6Take the input image in NV12 format as an example, where the channel data corresponding to the Y channel of all pixels includes 24 Y channel values, namely Y1 to Y24, the channel data corresponding to the U channel of all pixels includes 6 U channel values, namely U1 to U6, and the channel data corresponding to the V channel of all pixels includes 6 V channel values, namely V1 to V6. The channel data corresponding to the Y channel, U channel, and V channel of all pixels in YU12, YV12, and NV21 are similar to the NV12 format and will not be repeated here.

[0061] The following is the first channel data including Figure 6 Taking the 24 Y channel values ​​in as an example, the first data block is explained.

[0062] Taking N as 3 in step 401 as an example, the 24 Y channel values ​​included in the first channel data can be divided into 3 data blocks, each of which includes 8 Y channel values. In other words, each execution unit loads 8 Y channel values ​​from the video memory and performs the first operation on the 8 Y channel values. Among them, the 8 Y channel values ​​loaded by different execution units are not the same. For example, the three execution units are execution unit 1, execution unit 2, and execution unit 3. Execution unit 1 loads 8 Y channel values. Figure 6 Y1~Y8 in execution unit 2 loads Y9~Y16, and execution unit 3 loads Figure 6 Y17~Y24 in.

[0063] In the above step 401-1, a first operation is performed on multiple data to be processed to obtain a first operation result, which includes multiple implementation methods. In the following embodiments, taking the multiple data to be processed as a first data block as an example, possible implementation methods A1 and A2 are provided.

[0064] In implementation A1, the first data block includes first channel values ​​corresponding to a plurality of pixel points in the input image, the first operation is a sum operation, and the one execution unit performs a sum operation on all first channel values ​​included in the first data block to obtain a first operation result.

[0065] The first data block includes Figure 6 As an example, the execution unit is execution unit 1. Execution unit 1 performs a sum operation on Y1 to Y8 to obtain a first operation result y 1 Satisfies the following formula (1):

[0066] y 1 =Y 1 +Y 2 +Y 3 +Y 4 +Y 5 +Y 6 +Y 7 +Y8 Formula (1)

[0067] In implementation A2, the first data block includes first channel values ​​corresponding to multiple pixel points in the input image, and the first operation is an averaging operation; the one execution unit performs an averaging operation on the sum of the first channel values ​​included in the first data block according to the number of first channel values ​​included in the first data block to obtain a first operation result.

[0068] The first data block includes Figure 6 As an example, the execution unit is the execution unit 1. The number of first channel values ​​included in the first data block is 8. The execution unit 1 performs an average operation on the sum of Y1 to Y8 to obtain the first operation result y 2 Satisfies the following formula (2):

[0069]

[0070] Based on any of the above implementations, each of the N execution units may include K threads, and each thread corresponds to a register; the above step 401-1 may be implemented in the following manner: the K threads in one execution unit execute the second operation in parallel; wherein the second operation executed by one thread includes: loading a portion of the first channel values ​​in the first data block from the video memory into the register corresponding to the one thread; the one thread executes a first operation on the portion of the first channel values ​​in the register, and loads the result of the first operation executed by the one thread from the register corresponding to the one thread to the register corresponding to the reference thread among the K threads, wherein the reference thread is a thread serving as a core among the K threads in the one execution unit; the execution results of the first operation corresponding to the K threads in the one execution unit are accumulated in the reference thread in the one execution unit to obtain the first operation result.

[0071] Still including the first data block Figure 6Y1~Y8 in, K is 4, and the reference thread is thread 1 as an example. Thread 1 loads Y1 and Y2 from the video memory to the register 1 corresponding to thread 1, and performs a sum operation on Y1 and Y2 in register 1. The sum of Y1 and Y2 is stored in register 1; thread 2 loads Y3 and Y4 from the video memory to the register 2 corresponding to thread 2, and performs a sum operation on Y3 and Y4 in register 2, and then thread 2 loads the sum of Y3 and Y4 from register 2 to register 1; thread 3 loads Y5 and Y6 from the video memory to the register 3 corresponding to thread 3, and performs a sum operation on Y5 and Y6 in register 1. 3, Y5 and Y6 are summed, and then thread 3 loads the sum of Y5 and Y6 from register 3 to register 1; thread 4 loads Y7 and Y8 from the video memory to register 4 corresponding to thread 4, and sums Y7 and Y8 in register 4, and then thread 4 loads the sum of Y7 and Y8 from register 4 to register 1; the sum of Y1 and Y2 in register 1 is accumulated with the sum of Y1 and Y2 loaded into register 1, the sum of Y3 and Y4, the sum of Y5 and Y6, and the sum of Y7 and Y8, thereby realizing the summation operation at the thread warp level.

[0072] Step 401 - 2 , write the first operation result into the on-chip cache, and accumulate the first operation result with the execution results of the first operation written into the on-chip cache by other N−1 execution units among the N execution units.

[0073] Taking N as 3 and the first operation as a sum operation as an example, execution unit 1, execution unit 2 and execution unit 3 respectively execute the above steps 401-1 and 401-2. Figure 6 For example, when execution unit 1 executes step 401-1, execution unit 1 loads Figure 6 Y1~Y8 in the chip are obtained by performing a sum operation on Y1~Y8, and the sum of Y1~Y8 is written into the on-chip cache.

[0074] When execution unit 2 executes step 401-1, execution unit 2 loads Figure 6 Y9~Y17 in the chip are obtained by performing a sum operation on Y9~Y17, and the sum of Y9~Y17 is written into the on-chip cache.

[0075] When execution unit 3 executes step 401-1, execution unit 3 loads Figure 6 Y18~Y24 in the chip are obtained by performing a sum operation on Y18~Y24, and the sum of Y18~Y24 is written into the on-chip cache.

[0076] Execution unit 1 writes the sum of Y1-Y8 into the first storage address of the on-chip cache, execution unit 2 writes the sum of Y9-Y17 into the first storage address of the on-chip cache, and execution unit 3 writes the sum of Y18-Y24 into the first storage address of the on-chip cache. In this way, the sum of Y1-Y8, the sum of Y9-Y17, and the sum of Y18-Y24 are accumulated at the same storage address of the on-chip cache to obtain the first accumulation result in step 402.

[0077] Step 402: A reference execution unit among the N execution units loads a first accumulated result from an on-chip cache, where the first accumulated result includes an accumulated result of execution results of first operations respectively corresponding to the N execution units.

[0078] For N execution units, there is an execution unit that serves as a core and is called a reference execution unit.

[0079] In step 402, the first accumulated result is the accumulated result of the first operation result and the execution results of the first operations of the other N-1 execution units among the N execution units except the one execution unit.

[0080] Step 403: The reference execution unit among the N execution units determines a first average result according to the first accumulation result, and writes the first average result to the video memory.

[0081] The types of the first operation in the above step 401-1 are different, and accordingly, the implementation methods of the reference execution unit in the N execution units in step 403 determining the first mean value result according to the first accumulation result are also different. For details, please refer to the following implementation methods B1 and B2.

[0082] Implementation method B1 is based on the above implementation method A1, the first operation is a sum operation, and the reference execution unit among the N execution units performs an averaging operation on the first accumulation result according to the number of all first channel values ​​loaded by the N execution units to obtain a first average result.

[0083] Load with 3 execution units Figure 6 Taking Y1 to Y24 as an example, the number of first channel values ​​loaded by N execution units is 24. Combining the above implementation A1, it can be known that the first accumulated result Y in step 403 sum The following formula (3) is satisfied:

[0084] Y sum =Y 1 +Y 2 +…+Y 23 +Y 24 Formula (3)

[0085] According to the number 24 of the first channel values ​​loaded by the three execution units, the first accumulation result Y sum Perform the averaging operation to obtain the first average result Satisfies the following formula (4):

[0086]

[0087] In this example, the first mean result is the mean result corresponding to the first channel data.

[0088] Implementation method B2, based on the above-mentioned method A2, the first operation is an averaging operation, and the reference execution unit among the N execution units can perform an averaging operation on the first accumulated result according to N to obtain a first average result.

[0089] Load with 3 execution units Figure 6 Taking Y1 to Y24 as an example, then in combination with the above implementation A2, it can be known that the first accumulated result Y in step 403 sum The following formula (5) is satisfied:

[0090]

[0091] According to the number of execution units 3, the first accumulated result Y sum Perform the averaging operation to obtain the first average result Satisfies the following formula (6):

[0092]

[0093] In an embodiment of the present application, N execution units process the first channel data in parallel, each execution unit loads a data block obtained after segmentation of the first channel data to perform a first operation, and writes the execution result of the first operation into an on-chip cache. The execution results of the first operation corresponding to the N execution units can be accumulated in the on-chip cache, for example, accumulated in the same storage address in the on-chip cache, and finally the first accumulated result is stored in this storage address. Data transfer between execution units can be achieved through the on-chip cache, and N execution units process the first channel data in parallel. Compared with processing the channel data of a channel at the granularity of threads or thread blocks, the concurrent computing capability of the artificial intelligence chip can be effectively exerted, thereby improving the execution efficiency of the mean operator.

[0094] In some other embodiments, N execution units can be divided into M execution groups, each execution group includes multiple execution units, and the on-chip cache includes group shared memories corresponding to the M execution groups. For one execution group, the execution group corresponds to one group shared memory, and each execution unit in the execution group uses the group shared memory corresponding to the execution group. The above step 401-2 can be implemented in the following manner: writing the first operation result into the target group shared memory corresponding to the target execution group to which an execution unit belongs; accumulating the first operation result and the execution results of the first operation written into the target group shared memory by other execution units in the target execution group except one execution unit to obtain a second accumulated result.

[0095] Taking N as 4 and M as 2 as an example, the 4 execution units are divided into 2 execution groups, namely execution group 1 and execution group 2. Each execution group includes 2 execution units. Execution unit 1 and execution unit 2 in execution group 1 use the same group shared memory 1, and execution unit 3 and execution unit 4 in execution group 2 use the same group shared memory 2.

[0096] Taking the execution unit 1 executing the above step 401-2 as an example, the execution unit 1 writes its first operation result 1 to the second storage address in the target group shared memory 1 corresponding to the execution group 1 to which the execution unit 1 belongs; the execution result of the first operation corresponding to the execution unit 2 is called the first operation result 2, and the execution unit 2 also writes the first operation result 2 to the second storage address in the target group shared memory 1, that is, the first operation result 1 and the second operation result 2 are written to the same storage address in the target group shared memory; then, the first operation result 1 corresponding to the execution unit 1 and the first operation result 2 corresponding to the execution unit 2 are accumulated in the target group shared memory 1 to obtain the second accumulated result 1, and the second accumulated result 1 is stored in the second storage address in the target group shared memory 1. In this example, if the first operation result 1 corresponding to the execution unit 1 is first written into the target group shared memory 1, and the first operation result 2 corresponding to the execution unit 2 is later written into the target group shared memory 1, then when the execution unit 2 writes the first operation result 2 into the target group shared memory 1, the first operation result 2 can be accumulated into the first operation result 1 that has been written into the target group shared memory 1 to obtain the second accumulated result 1. If the first operation result 2 corresponding to execution unit 2 is written into the target group shared memory 1 first, and the first operation result 1 corresponding to execution unit 1 is written into the target group shared memory 1 later, then when executing unit 1 writes the first operation result 1 into the target group shared memory 1, it can add the first operation result 1 to the first operation result 2 that has been written into the target group shared memory 1 to obtain the second accumulated result 1 stored in the second storage address in the target shared memory 1.

[0097] Taking the execution unit 3 executing the above step 401-2 as an example, the execution unit 3 writes its first operation result 3 to the third storage address in the target group shared memory 2 corresponding to the execution group 2 to which the execution unit 3 belongs; the execution result of the first operation corresponding to the execution unit 4 is called the first operation result 4, and the execution unit 4 also writes the first operation result 4 to the third storage address in the target group shared memory 2; then, the first operation result 3 corresponding to the execution unit 3 and the first operation result 4 corresponding to the execution unit 4 are accumulated in the target group shared memory 2 to obtain the second accumulation result 2, and the second accumulation result 2 is stored in the third storage address in the target group shared memory 2. In this example, the first operation result written later in the target group shared memory 2 is accumulated to the first operation result already written in the target shared memory 2, thereby obtaining the second accumulation result 2 stored in the third storage address in the target group shared memory 2.

[0098] Furthermore, the on-chip cache also includes a secondary cache, which is referred to as L2Cache hereinafter. After accumulating the first operation result with the execution result of the first operation written into the target group shared memory by other execution units except one execution unit in the target execution group to obtain the second accumulated result, the first execution unit in the target execution group can load the second accumulated result from the target group shared memory; wherein, an execution unit as a core in each execution group is called the first execution unit, and the first execution unit in the target execution group writes the loaded second accumulated result into L2 Cache, and accumulates the second accumulated result with the second accumulated results corresponding to the other M-1 execution groups in the M execution groups written into L2 Cache; the result obtained by accumulating the second accumulated results corresponding to the M execution groups is used as the first accumulated result. The above step 402 can be implemented in the following way: the reference execution unit in the N execution units loads the first accumulated result from the L2 Cache.

[0099] For the first execution unit in the target execution group, the second accumulation result can be loaded from the target group shared memory and written to the fourth storage address in the L2 Cache. The other M-1 execution groups except the target execution group among the M execution groups also write their respective corresponding second accumulation results to the fourth storage address in the L2 Cache. In other words, the second accumulation results corresponding to the M execution groups are all written to the same storage address in the L2 Cache, namely, the fourth storage address, so that the second accumulation results corresponding to the M execution groups are accumulated in the L2 Cache to obtain the first accumulation result stored at the fourth storage address in the L2 Cache.

[0100] The reference execution unit among the N execution units may load the first accumulation result from the L2 Cache, determine a first average result according to the first accumulation result, and write the first average result to the video memory.

[0101] The N execution units are divided into M execution groups. For the M execution groups, each execution group has a first execution unit as a core, and the M execution groups correspond to M first execution units. The reference execution unit in the above N execution units can be an execution unit as a core in any execution group, that is, the reference execution unit in the above N execution units can be any first execution unit in the M first execution units.

[0102] In an embodiment of the present application, the above-mentioned M execution groups may be located in the same SPC, that is, all threads included in each execution unit on an SPC on the artificial intelligence chip are used to process the first channel data.

[0103] In some other embodiments, the M execution groups may be located on multiple SPCs, and each SPC includes at least one of the M execution groups, that is, the first channel data is processed using the execution units on the multiple SPCs on the artificial intelligence chip, and each execution unit includes multiple threads. All threads included in all execution units in the multiple SPCs on the artificial intelligence chip can be used to process the first channel data in parallel. Compared with using the threads included in all execution units in one SPC to process the first channel data, the parallelism of using multiple SPCs to process the first channel data is higher, so that the execution efficiency of the mean operator is higher.

[0104] The following takes the processing of the first channel data by multiple SPCs in an artificial intelligence chip as an example to introduce in detail the mean operator execution method provided in the embodiment of the present application.

[0105] The AI ​​chip can include multiple SPCs, e.g. Figure 7 As shown in SPC0 and SPC1, each SPC may include multiple execution groups (CUs), for example Figure 7 The SPC0 shown includes CU00 and CU01, SPC1 includes CU10 and CU11, each execution group (CU) may include multiple execution units (EU), for example, CU00 includes EU00 and EU01, CU01 includes EU10 and EU11, CU10 includes EU20 and EU21, CU11 includes EU30 and EU31; each execution unit includes multiple threads, for example Figure 7It is shown that each execution unit EU includes 40 threads, for example, EU00 includes threads 000 to 039, EU01 includes threads 100 to 139, EU10 includes threads 200 to 239, EU11 includes threads 300 to 339, etc. Figure 7 The threads included in other EUs in are not listed here one by one. It should be understood that Figure 7 There is no limitation on the number of SPCs, CUs, EUs, threads, etc. included in the artificial intelligence chip in the embodiment of the present application. The artificial intelligence chip in the embodiment of the present application may also include more SPCs, more CUs, more EUs, and more threads. Figure 7 In addition, the artificial intelligence chip may also include HMB, Figure 7 The HBM used to store the first channel data and the HBM used to store the average result obtained by the averaging operation are actually the same HBM, wherein the first channel data and the average result can be stored in different storage addresses in the same HBM.

[0106] The process of AI chip processing the first channel data includes:

[0107] S1, each of all threads included in the multiple SPCs reads multiple first channel values ​​from the HBM respectively. Here, the number of first channel values ​​allocated to each thread can be determined according to the total number of all threads included in the multiple SPCs and the number of first channel values ​​in the first channel data. For example, if the total number of all threads is 512 and the number of first channel values ​​is 20480, then 40 first channel values ​​can be allocated to each thread for processing, and each thread accumulates the 40 first channel values ​​into its own register, that is, the register corresponding to each thread stores the accumulated result of the 40 first channel values; for example, the accumulated result of the 40 first channel values ​​stored in the register corresponding to thread 001 is called result 001, the accumulated result of the 40 first channel values ​​stored in the register corresponding to thread 002 is called result 002, ... the accumulated result of the 40 first channel values ​​stored in the register corresponding to thread 039 is called result 039.

[0108] S2. For an execution unit (EU), the results in the registers of each thread in the execution unit are accumulated to the register corresponding to a reference thread in the execution unit. For example, results 001 to 039 in the registers corresponding to threads 001 to 039 in EU00 are accumulated to the register corresponding to thread 000, and the result after accumulating results 001 to 039 and result 000 in the register corresponding to thread 000 is called result 040. In this way, only the register corresponding to thread 000 in EU00 has data, that is, result 040. Similarly, results 101 in the register corresponding to thread 101, results 102 in the register corresponding to thread 102, ... and result 139 in the register corresponding to thread 139 in EU01 are accumulated to the register corresponding to thread 100. The result after accumulating results 101 to 139 and result 100 in the register corresponding to thread 100 is called result 140. In this way, only the register corresponding to thread 100 in EU01 has data, that is, result 140. Similarly, in EU10, only the register corresponding to thread 200 contains data, that is, the result of the accumulation of results 200 to 239 corresponding to threads 200 to 239 is called result 240; in EU11, only the register corresponding to thread 300 contains data, that is, the result of the accumulation of results 300 to 339 corresponding to threads 300 to 339 is called result 340.

[0109] S3, for a CU, the reference threads corresponding to the multiple EUs inside the CU read data from their own registers and accumulate them into the group shared memory corresponding to the CU. Figure 7 In the shown SPC0, thread 000 in EU00 in CU00 reads result 040 from its own register and writes it into GSM0 corresponding to CU00; thread 100 in EU01 in CU00 reads result 140 from its own register and writes it into GSM0 to be accumulated with result 040, and the accumulated result is called result 041. Similarly, thread 200 in EU10 in CU01 reads result 240 from its own register and writes it into GSM1 corresponding to CU01; thread 300 in EU11 in CU01 reads result 340 from its own register and writes it into GSM1 to be accumulated with result 240, and the accumulated result is called result 042. Similarly, the accumulated result stored in GSM2 corresponding to CU10 in SPC1 is called result 043, and the accumulated result stored in GSM3 corresponding to CU11 in SPC1 is called result 044.

[0110] S4, in SPC0, the reference thread in EU00 in CU00 is thread 000, thread 000 reads result 041 from GSM0 to its own register, and then writes it to address 1 of L2 Cache; the reference thread in EU10 in CU01 is thread 100, thread 100 reads result 042 from GSM1 to its own register, and then writes it to address 1 of L2 Cache. The results written into L2 Cache are accumulated, that is, result 041 and result 042 are accumulated in L2 Cache to obtain a result, called result 045. The reference thread in EU00 in CU10 in SPC1 is thread 400, thread 400 reads result 043 from GSM2 to its own register, and then writes it to address 1 of L2 Cache. The result after accumulating result 043 and result 045 is called result 046, that is, the content stored in address 1 of L2 Cache is result 046. The reference thread in EU10 in CU11 in SPC1 is thread 600. Thread 600 reads result 044 from GSM3 to its own register, and then writes it to address 1 of L2Cache. The result after adding result 044 and result 046 is called result 047. The content finally stored in address 1 of L2Cache is result 047. Of course, Figure 7 The artificial intelligence chip shown may also include more CUs, so the results written into the L2 Cache by more CUs are sequentially accumulated with the result stored in address 1 to obtain the final accumulated result. In step S5, the average operation is performed using the final accumulated result 047 as an example.

[0111] S5, thread 000 in EU00 in CU00 in SPC0 loads result 047 from address 1 in L2 cache into the register corresponding to thread 000, and performs an average operation on result 047 in the register to obtain the average result corresponding to the first channel data; thread 000 writes the average result to HBM.

[0112] In the embodiment of the present application, the first channel data is evenly loaded into all threads, and the communication and data exchange problems between EUs are solved through GSM, and the communication and data exchange problems between CU and SPC are solved through L2 Cache, thereby increasing the computing efficiency, balancing the computing and memory access, and optimizing the performance of the image mean operator.

[0113] In the above embodiment, the first channel data in the input image is used as the input data of the mean operator, and N execution units in the artificial intelligence chip are used to execute the mean operator in parallel. For an input image including multiple channel data, there are the following two possible implementation methods.

[0114] In a possible implementation, the artificial intelligence chip processes the multiple channel data in series. Taking the input image including the first channel data and the second channel data as an example, the artificial intelligence chip processes the multiple channel data in series according to the above Figure 4 After the method steps shown in the figure process the first channel data, refer to Figure 4 The method steps shown process the second channel data.

[0115] Exemplarily, after the above step 403, the artificial intelligence chip further performs the following steps 801 to 803:

[0116] Step 801: N execution units execute a third operation in parallel; wherein the third operation executed by one execution unit includes the following steps 801-1 and 801-2:

[0117] Step 801-1, load the second data block in the second channel data from the video memory, perform a third operation on the second data block, and obtain a third operation result; the second data block is any data block among the N data blocks obtained after the second channel data is segmented, and different execution units among the N execution units load different data blocks.

[0118] Figure 8 The third operation involved in the above may be a sum operation or an average operation. The specific implementation may refer to the above Figure 4 The description of the first operation in will not be repeated here.

[0119] Step 801 - 2 , writing the third operation result into the on-chip cache, and accumulating the third operation result with the execution results of the third operation written into the on-chip cache by other N−1 execution units among the N execution units.

[0120] Step 802, a reference execution unit among the N execution units loads a third accumulated result from an on-chip cache, where the third accumulated result includes an accumulated result of execution results of third operations respectively corresponding to the N execution units;

[0121] Step 803: The reference execution unit among the N execution units determines a second average result according to the third accumulation result, and writes the second average result to the video memory.

[0122] The second mean result here is the mean result corresponding to the second channel data.

[0123] The specific implementation details of the above steps 801 to 803 can be found in the above Figure 4 The relevant descriptions of steps 401 to 403 in the above description will not be repeated here.

[0124] In this implementation, multiple channel data included in the input image can be processed by multiplexing hardware resources, and compared with processing one channel data at the granularity of threads or thread blocks, the parallelism of the calculation can be increased, making full use of the computing power of the artificial intelligence chip.

[0125] In another possible implementation, the artificial intelligence chip processes multiple channel data in parallel. The processing process of each channel data can be seen in the above Figure 4 The method steps are shown.

[0126] The artificial intelligence chip processes multiple channel data in series. Taking the input image including the first channel data and the second channel data as an example, the artificial intelligence chip uses N execution units according to the above Figure 4 The method steps shown in the figure are used to process the first channel data. At the same time, the artificial intelligence chip also uses F execution units to process the second channel data in parallel. The process of processing the second channel data can be referred to in the above Figure 4 The value of F here can be the same as or different from the value of N mentioned above, and this application does not limit this.

[0127] Exemplarily, the artificial intelligence chip performs the above Figure 4 When performing the method steps shown, the following steps 901 to 903 are also performed in parallel:

[0128] Step 901: F execution units execute a fourth operation in parallel; wherein the fourth operation executed by one of the F execution units includes the following steps 901-1 and 901-2:

[0129] Step 901-1, loading a third data block in the second channel data in the input image from the video memory, performing a fourth operation on the third data block, and obtaining a fourth operation result; the third data block is any data block among the F data blocks segmented from the second channel data, and different execution units among the F execution units load different data blocks;

[0130] Fig. 9 The fourth operation involved in the above may be a sum operation or an average operation. The specific implementation may refer to the above Figure 4 The description of the first operation in will not be repeated here.

[0131] Step 901-2, writing the fourth operation result into the on-chip cache, and accumulating the fourth operation result with the execution results of the fourth operation written into the on-chip cache by the other F-1 execution units among the F execution units;

[0132] Step 902 : A reference execution unit among the F execution units loads a third accumulated result from an on-chip cache, where the third accumulated result includes an accumulated result of execution results of fourth operations of the F execution units.

[0133] Step 903: The reference execution unit among the F execution units determines a third average result according to the third accumulation result, and writes the third average result to the video memory.

[0134] The third mean result here is the mean result corresponding to the second channel data.

[0135] The specific implementation details of the above steps 901 to 903 can be found in the above Figure 4 The relevant descriptions of steps 401 to 403 in the above description will not be repeated here.

[0136] Compared with the method of processing a channel data at the granularity of threads or thread blocks, this implementation method can increase the parallelism of calculations and make full use of the computing power of artificial intelligence chips.

[0137] Based on the same technical concept, an embodiment of the present application provides an artificial intelligence chip, including N execution units, a video memory, and an on-chip cache shared by the N execution units, where N is a positive integer;

[0138] The N execution units are used to execute a first operation in parallel; wherein the first operation executed by one execution unit includes: loading a plurality of to-be-processed data in the input image from the video memory, executing a first operation on the plurality of to-be-processed data, and obtaining a first operation result; the plurality of to-be-processed data loaded by different execution units in the N execution units are different; writing the first operation result into an on-chip cache, and accumulating the first operation result with the execution results of the first operation written into the on-chip cache by other N-1 execution units in the N execution units;

[0139] The reference execution unit among the N execution units is used to load a first accumulation result from the on-chip cache, determine a mean result corresponding to the first channel data according to the first accumulation result, and write the mean result corresponding to the first channel data to the video memory, wherein the first accumulation result includes an accumulation result of execution results of the first operations respectively corresponding to the N execution units, and the reference execution unit is an execution unit serving as a core among the N execution units.

[0140] Optionally, the multiple data to be processed include first channel values ​​corresponding to multiple pixel points in the input image, each of the N execution units includes K threads, and each thread corresponds to a register; when the one execution unit loads the multiple data to be processed in the input image from the video memory, and performs a first operation on the multiple data to be processed to obtain a first operation result, it is specifically used to: perform a second operation in parallel through K threads; wherein the second operation performed by one thread includes: loading part of the first channel values ​​in the multiple data to be processed from the video memory into the register corresponding to the one thread; performing the first operation on the part of the first channel values ​​in the register corresponding to the one thread, and loading the result of the first operation performed by one thread from the register corresponding to the one thread to the register corresponding to the reference thread; the reference thread is a thread that serves as a core among the K threads in the one execution unit, and the execution results of the first operation corresponding to the K threads are accumulated in the register corresponding to the reference thread to obtain the first operation result.

[0141] Optionally, the N execution units are divided into M execution groups, and the on-chip cache includes group shared memories corresponding to the M execution groups respectively; when the one execution unit writes the first operation result into the on-chip cache, and accumulates the first operation result with the execution results of the first operation written into the on-chip cache by other N-1 execution units among the N execution units, it is specifically used to: write the first operation result into the target group shared memory corresponding to the target execution group to which the one execution unit belongs; accumulate the first operation result with the execution results of the first operation written into the target group shared memory by other execution units in the target execution group except the one execution unit, to obtain a second accumulated result.

[0142] Optionally, the on-chip cache also includes a secondary cache; the first execution unit within the target execution group is also used to load the second accumulated result from the target group shared memory; write the loaded second accumulated result into the secondary cache, and accumulate the second accumulated result with the second accumulated results corresponding to the other M-1 execution groups in the M execution groups written into the secondary cache, the first execution unit being the core execution unit within the target execution group; the result obtained by accumulating the second accumulated results corresponding to the M execution groups is used as the first accumulated result; the reference execution unit among the N execution units loads the first accumulated result from the on-chip cache, specifically used to: load the first accumulated result from the secondary cache.

[0143] Optionally, the multiple data to be processed include first channel values ​​corresponding to multiple pixel points in the input image, and the first operation is a summation operation; the one execution unit performs a first operation on the multiple data to be processed to obtain a first operation result, specifically for: performing a summation operation on all first channel values ​​included in the multiple data to be processed to obtain the first operation result; the reference execution unit determines the first mean result based on the first accumulation result, specifically for: the reference execution unit performs an averaging operation on the first accumulation result according to the number of first channel values ​​loaded by the N execution units to obtain the first mean result.

[0144] Optionally, the multiple data to be processed include first channel values ​​corresponding to multiple pixel points in the input image, and the first operation is an averaging operation; when the one execution unit performs the first operation on the multiple data to be processed to obtain a first operation result, it is specifically used to: perform an averaging operation on the sum of the first channel values ​​included in the multiple data to be processed according to the number of first channel values ​​included in the multiple data to be processed, to obtain the first operation result; when the reference execution unit determines the first mean result based on the first accumulated result, it is specifically used to: perform an averaging operation on the first accumulated result according to N, to obtain the first mean result.

[0145] Optionally, the input image includes first channel data, and the N data blocks obtained after the first channel data is segmented respectively include different multiple data to be processed; when the one unit loads multiple data to be processed in the input image from the video memory, and performs a first operation on the multiple data to be processed to obtain a first operation result, it is specifically used to: load a first data block in the first channel data from the video memory, and perform a first operation on the first data block to obtain a first operation result; the first data block is any data block among the N data blocks obtained by segmenting the first channel data, and different execution units among the N execution units load different data blocks.

[0146] Optionally, the multiple data to be processed belong to the first channel data in the input image, and the input image also includes the second channel data; the N execution units are also used to: perform a third operation in parallel; wherein the third operation performed by one execution unit includes: loading a second data block in the second channel data in the input image from the video memory, performing a third operation on the second data block, and obtaining a third operation result; the second data block is any data block in the N data blocks obtained after the second channel data is segmented, and different execution units in the N execution units load different data blocks; writing the third operation result into an on-chip cache, and accumulating the third operation result with the execution results of the third operation written into the on-chip cache by other N-1 execution units in the N execution units; the reference execution unit in the N execution units is used to: load a third accumulated result from the on-chip cache, determine the second mean result based on the third accumulated result, and write the second mean result out to the video memory, and the third accumulated result includes the accumulated result of the execution results of the third operation of the N execution units.

[0147] Optionally, the multiple data to be processed belong to the first channel data in the input image, and the input image also includes the second channel data; the artificial intelligence chip also includes F execution units, which are used to: perform a fourth operation in parallel; wherein the fourth operation performed by one of the F execution units includes: loading a third data block in the second channel data in the input image from the video memory, performing a fourth operation on the third data block, and obtaining a fourth operation result; the third data block is any data block in the F data blocks obtained by segmenting the second channel data, and different execution units in the F execution units load different data blocks; writing the fourth operation result into an on-chip cache, and accumulating the fourth operation result with the execution results of the fourth operation written into the on-chip cache by other F-1 execution units in the F execution units;

[0148] The reference execution unit among the F execution units is used to load the fourth accumulated result from the on-chip cache, determine the third average result according to the fourth accumulated result, and write the third average result to the video memory; wherein the fourth accumulated result includes the accumulated result of the execution results of the fourth operations of the F execution units.

[0149] Based on the same inventive concept, the present application provides a computer device, which includes a memory, a processor chip, and a computer program stored in the memory and executable on the processor chip. When the processor chip executes the computer program, the method steps in any of the above method embodiments are implemented.

[0150] Based on the same technical concept, an embodiment of the present application provides a computer-readable storage medium, including computer-executable instructions. When the computer-executable instructions are executed on a computer, the computer executes the method steps in any of the above method embodiments.

[0151] Based on the same technical concept, an embodiment of the present application provides a computer program product, which stores a computer program. The computer program includes program instructions. When the program instructions are executed by a computer, the computer executes the method steps in any of the above method embodiments.

[0152] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, devices (equipment), systems, chips, computer-readable storage media, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware, which are collectively referred to as "modules" or "systems" herein.

[0153] The present application is described with reference to at least one of the following diagrams of the method, apparatus (device) or system of the present application: a flowchart, a block diagram. It should be understood that at least one of the following can be implemented by computer program instructions: each process in the flowchart, each box in the block diagram, and a combination of a process in the flowchart and a box in the block diagram.

[0154] These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device for implementing at least one of the following: a device for the functions specified in one or more flows of the flowchart, a device for the functions specified in one or more blocks of the block diagram.

[0155] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device, which implements at least one of the following: a function specified in one or more processes in a flowchart, or a function specified in one or more boxes in a block diagram.

[0156] These computer program instructions may also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable device provide steps for implementing at least one of the following: steps for specifying a function in one or more processes in a flowchart, steps for specifying a function in one or more boxes in a block diagram.

[0157] Although the present invention has been described in conjunction with specific features and embodiments thereof, it is apparent that various modifications and combinations may be made thereto without departing from the spirit and scope of the present invention. Accordingly, this specification and the accompanying drawings are merely exemplary illustrations of the present invention as defined by the appended claims and are deemed to cover any and all modifications, variations, combinations or equivalents within the scope of the present invention. Obviously, those skilled in the art may make various modifications and variations to the present invention without departing from the spirit and scope of the present invention. Thus, the present invention is intended to include such modifications and variations if they fall within the scope of the claims of the present invention and their equivalents.

Claims

1. A method for executing a mean operator, characterized in that: Applied to an artificial intelligence chip, the artificial intelligence chip includes N execution units, a video memory, and an on-chip cache, the N execution units are divided into M execution groups, the on-chip cache includes a secondary cache and group shared memories corresponding to the M execution groups, each execution group includes multiple execution units, and each execution unit belonging to the same execution group uses the same group shared memory; N is a positive integer, and M is a positive integer less than N; The method comprises: The N execution units execute the first operation in parallel; wherein the first operation executed by one execution unit includes: loading a plurality of to-be-processed data in the input image from the video memory, executing a first operation on the plurality of to-be-processed data to obtain a first operation result, the plurality of to-be-processed data loaded by different execution units among the N execution units are different, and the plurality of to-be-processed data respectively loaded by the N execution units all belong to the first channel data in the input image; writing the first operation result into the target group shared memory corresponding to the target execution group to which the one execution unit belongs, and accumulating the first operation result with the execution results of the first operation written into the target group shared memory by other execution units in the target execution group except the one execution unit to obtain a second accumulated result; A first execution unit in the target execution group loads the second accumulation result from the target group shared memory, the first execution unit being an execution unit serving as a core in the target execution group; The first execution unit writes the loaded second accumulation result into the secondary cache, and accumulates the second accumulation result with the second accumulation results corresponding to the other M-1 execution groups in the M execution groups written into the secondary cache; the result obtained by accumulating the second accumulation results corresponding to the M execution groups is used as the first accumulation result; A reference execution unit among the N execution units loads the first accumulated result from the secondary cache, wherein the first accumulated result includes an accumulated result of execution results of the first operations respectively corresponding to the N execution units, and the reference execution unit is an execution unit serving as a core among the N execution units; The reference execution unit determines a first average result according to the first accumulated result, and writes the first average result to the video memory.

2. The method according to claim 1, characterized in that Each of the N execution units includes K threads, and each thread corresponds to a register; the loading of a plurality of to-be-processed data in the input image from the video memory, and performing a first operation on the plurality of to-be-processed data to obtain a first operation result, includes: The K threads in one execution unit execute the second operation in parallel; wherein the second operation executed by one thread includes: Loading a portion of first channel values ​​in the plurality of to-be-processed data from the video memory into a register corresponding to the one thread; A first operation is performed on the part of the first channel values ​​in the register corresponding to the one thread, and a result of the first operation performed by the one thread is loaded from the register corresponding to the one thread to the register corresponding to the reference thread; the reference thread is a core thread among the K threads in the one execution unit, and the execution results of the first operations corresponding to the K threads are accumulated in the register corresponding to the reference thread to obtain the first operation result.

3. The method according to claim 1, characterized in that The plurality of data to be processed include first channel values ​​respectively corresponding to a plurality of pixel points in the input image, and the first operation is a sum operation; and performing the first operation on the plurality of data to be processed to obtain a first operation result includes: Performing a sum operation on all first channel values ​​included in the plurality of data to be processed to obtain the first operation result; The reference execution unit determines the first mean value result according to the first accumulation result, including: The reference execution unit performs an averaging operation on the first accumulated result according to the number of first channel values ​​loaded by the N execution units to obtain the first average result.

4. The method according to claim 1, characterized in that The plurality of data to be processed include first channel values ​​respectively corresponding to a plurality of pixel points in the input image, and the first operation is an average operation; The performing a first operation on the plurality of data to be processed to obtain a first operation result includes: According to the number of first channel values ​​included in the plurality of data to be processed, performing an average operation on the sum of the first channel values ​​included in the plurality of data to be processed to obtain the first operation result; The reference execution unit determines the first mean value result according to the first accumulation result, including: The reference execution unit performs an averaging operation on the first accumulated result according to the N to obtain the first average result.

5. The method according to claim 1, characterized in that The input image includes first channel data, and the N data blocks obtained after the first channel data is segmented respectively include a plurality of different data to be processed; The step of loading a plurality of data to be processed in the input image from the video memory and performing a first operation on the plurality of data to be processed to obtain a first operation result includes: A first data block in the first channel data is loaded from the video memory, and a first operation is performed on the first data block to obtain a first operation result; the first data block is any one of the N data blocks obtained by segmenting the first channel data, and different execution units among the N execution units load different data blocks.

6. The method according to claim 1, characterized in that The input image also includes second channel data; The reference execution unit determines a first average result according to the first accumulation result, and after writing the first average result to the video memory, further includes: The N execution units execute the third operation in parallel; wherein the third operation executed by one execution unit includes: loading the second data block in the second channel data from the video memory, and executing the third operation on the second data block to obtain a third operation result; the second data block is any data block of the N data blocks obtained after the second channel data is segmented, and different execution units in the N execution units load different data blocks; writing the third operation result into the on-chip cache, and accumulating the third operation result with the execution results of the third operation written into the on-chip cache by other N-1 execution units in the N execution units; A reference execution unit among the N execution units loads a third accumulated result from the on-chip cache, wherein the third accumulated result includes an accumulated result of execution results of third operations respectively corresponding to the N execution units; The reference execution unit determines a second average result according to the third accumulated result, and writes the second average result to the video memory.

7. The method according to claim 1, characterized in that, The input image also includes second channel data; the artificial intelligence chip also includes F execution units; the method also includes: The F execution units execute the fourth operation in parallel; wherein the fourth operation executed by one of the F execution units includes: loading a third data block in the second channel data from the video memory, and executing a fourth operation on the third data block to obtain a fourth operation result; the third data block is any data block in the F data blocks obtained by segmenting the second channel data, and different execution units in the F execution units load different data blocks; writing the fourth operation result into the on-chip cache, and accumulating the fourth operation result with the execution results of the fourth operation written into the on-chip cache by the other F-1 execution units in the F execution units; The reference execution unit among the F execution units loads the fourth accumulated result from the on-chip cache, determines a third average result based on the fourth accumulated result, and writes the third average result to the video memory; wherein the fourth accumulated result includes the accumulated result of the execution results of the fourth operations respectively corresponding to the F execution units.

8. An artificial intelligence chip, characterized in that: The system comprises N execution units, a video memory and an on-chip cache, wherein the N execution units are divided into M execution groups, the on-chip cache comprises a secondary cache and group shared memories respectively corresponding to the M execution groups, each execution group comprises a plurality of execution units, and each execution unit belonging to the same execution group uses the same group shared memory, and M is a positive integer less than N; Said N is a positive integer; The N execution units are used to execute a first operation in parallel; wherein the first operation executed by one execution unit includes: loading a plurality of data to be processed in the input image from the video memory, and executing a first operation on the plurality of data to be processed to obtain a first operation result; the plurality of data to be processed loaded by different execution units among the N execution units are different, and the plurality of data to be processed loaded by the N execution units respectively belong to the first channel data in the input image; writing the first operation result into a target group shared memory corresponding to the target execution group to which the one execution unit belongs, and accumulating the first operation result with the execution results of the first operation written into the target group shared memory by other execution units other than the one execution unit in the target execution group by other N-1 execution units among the N execution units to obtain a second accumulated result; The first execution unit in the target execution group is used to load the second accumulated result from the target group shared memory, and the first execution unit is an execution unit as a core in the target execution group; write the loaded second accumulated result into the secondary cache, and accumulate the second accumulated result with the second accumulated results corresponding to the other M-1 execution groups in the M execution groups written into the secondary cache; and accumulate the second accumulated results corresponding to the M execution groups as the first accumulated result; The reference execution unit among the N execution units is also used to load the first accumulated result from the secondary cache, determine the first average result based on the first accumulated result, and write the first average result to the video memory; the reference execution unit is the execution unit serving as the core among the N execution units, and the first accumulated result includes the accumulated result of the execution results of the first operations respectively corresponding to the N execution units.

9. A computer device comprising a memory, a processor chip, and a computer program stored in the memory and executable on the processor chip, characterized in that: When the processor chip executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that: The method comprises computer executable instructions, which, when executed on a computer, cause the computer to execute the steps of the method according to any one of claims 1 to 7.

11. A computer program product, characterized in that The computer program product stores a computer program, wherein the computer program includes program instructions. When the program instructions are executed by a computer, the computer is caused to execute the steps of the method according to any one of claims 1 to 7.