Data processing method and device, computer equipment and program product

By deploying communication units in computing units and setting matching computing units and communication configurations for the target neural network, the problems of power consumption and cost in large-scale neural network operation are solved, and low-cost, efficient data processing and accuracy are achieved.

CN120597960APending Publication Date: 2025-09-05BEIJING PINGXIN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510666094.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-02-13
Filing Date
2025-05-22
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

The operation of large-scale neural networks results in huge power consumption and cost, and existing solutions usually reduce communication flexibility and affect communication efficiency.

Method used

Communication units are deployed in computing units. By pre-setting matching computing units for each network layer of the target neural network and setting communication lines and configuration files for them, flexible communication and orderly data processing between computing units can be achieved.

Benefits of technology

It reduces the inference cost and power consumption of large models, improves the interconnection performance and communication flexibility between multiple computing units, and ensures the accuracy of data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120597960A_ABST
    Figure CN120597960A_ABST
Patent Text Reader

Abstract

The invention provides a data processing method and device, computer equipment and a program product, and the method comprises the steps: carrying out the data processing of input data through a processing unit in a calculation unit which is matched with any network layer of a target neural network, and obtaining output data; sending the output data to at least one other computing unit by using the communication unit in the computing unit and the communication units in the other computing units according to communication connection lines and communication configuration files set for the communication units in the different computing units; wherein the other computing units comprise other computing units matched with the network layer or computing units matched with the next network layer; and returning to the step of performing data processing on the input data by using the processing unit in the calculation unit to obtain the output data until the output data of the last calculation unit matched with the last network layer of the target neural network is obtained, and taking the output data as a data processing result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to a data processing method, apparatus, computer device, and program product. Background Art

[0002] The operation of large-scale neural network models requires massive computing power, which in turn relies on massive storage capacity. This massive storage capacity, in turn, relies on massive data transmission. This massive data transmission leads to significant power consumption and costs. Therefore, the operation of large-scale neural networks can incur significant power consumption and costs. Addressing these power and cost issues requires addressing the issue of massive data transmission. While some solutions have been proposed, these often reduce communication flexibility and efficiency, resulting in significant drawbacks. Summary of the Invention

[0003] The embodiments of the present disclosure at least provide a data processing method, apparatus, computer device, and program product.

[0004] In a first aspect, an embodiment of the present disclosure provides a data processing method, including:

[0005] For a computing unit that matches any network layer of the target neural network, using a processing unit in the computing unit, processing the input data to obtain output data;

[0006] According to the communication lines and communication configuration files set for the communication units in different computing units, the output data is sent to at least one other computing unit using the communication unit in the computing unit and the communication units in other computing units; wherein the other computing units include other computing units matching the network layer or computing units matching the next network layer;

[0007] Return to the step of using the processing unit in the computing unit to process the input data to obtain output data, until the output data of the last computing unit matching the last network layer of the target neural network is obtained, and the output data is used as the data processing result.

[0008] In one possible implementation, the network layer includes multiple functional layers;

[0009] The multiple functional layers at least include a normalization processing functional layer located in the attention layer, a matrix-vector multiplication functional layer related to the query matrix, the key matrix and the value matrix, a multi-head functional layer, a matrix-vector multiplication functional layer related to the result output, an accumulation processing functional layer, and a normalization processing functional layer located in the feedforward neural network layer, a matrix-vector multiplication functional layer related to multiple weight matrices, and an accumulation processing functional layer;

[0010] A computing unit matching any network layer of the target neural network is associated with a functional layer among the multiple functional layers.

[0011] In a possible implementation, the processing unit in the computing unit processes the input data to obtain output data, including:

[0012] In a case where the computing unit is associated with any matrix-vector multiplication functional layer, using a processing unit in the computing unit, performing matrix-vector multiplication processing on the input data and target data stored in a storage unit in the computing unit to obtain output data;

[0013] Among them, the target data includes any preset matrix corresponding to the functional layer matching the computing unit or a submatrix obtained by matrix splitting any preset matrix, and the preset matrix includes a query matrix, a key matrix, a value matrix, a result output matrix and a weight matrix.

[0014] In a possible implementation, the processing of the input data by the processing unit in the computing unit to obtain output data further includes:

[0015] When the computing unit is independent of the matrix-vector multiplication functional layer, the processing unit in the computing unit is used to perform corresponding functional processing on the input data according to the network function of the functional layer related to the computing unit to obtain output data.

[0016] In a possible implementation, before using the processing unit in the computing unit to process the input data to obtain output data, the method further includes:

[0017] In a case where the input data includes output data sent from multiple other computing units, the output data of each other computing unit is received and buffered respectively using multiple data receiving interfaces in the communication unit included in the computing unit;

[0018] In response to each of the data receiving interfaces completing buffering of the corresponding output data, the calculation unit is used to trigger a reception completion signal to determine that the acquisition of the input data is completed.

[0019] In one possible implementation, the method further includes:

[0020] Performing functionality testing and performance testing on the computing unit;

[0021] In the case of a failure in the test, the computing unit is replaced by a spare computing unit.

[0022] In a possible implementation, the weight matrix includes a feature mapping matrix, a feature transformation matrix, and a feature compression matrix;

[0023] In a case where the preset matrix is ​​any one of the query matrix, the key matrix, the value matrix, the feature mapping matrix, and the feature transformation matrix, the submatrix included in the target data is a matrix obtained by performing row splitting on the preset matrix;

[0024] In a case where the preset matrix is ​​any one of a result output matrix and a feature compression matrix, the submatrix included in the target data is a matrix obtained by performing column splitting on the preset matrix.

[0025] In a possible implementation, if the preset matrix is ​​split into multiple sub-matrices, the multiple sub-matrices are respectively stored in storage units of different computing units;

[0026] The sizes of the sub-matrices corresponding to a preset matrix are related to at least one of the size of the preset matrix, the storage space of the storage unit in the calculation unit, and the computing capability of the processing unit in the calculation unit.

[0027] In one possible implementation, the utilizing the processing unit in the computing unit to perform corresponding functional processing on the input data according to a network function of a functional layer related to the computing unit to obtain output data includes:

[0028] In a case where the computing unit is a computing unit that matches a multi-head functional layer in the attention layer, using a communication unit in the computing unit to obtain a matrix cache associated with a key matrix and a value matrix from a double rate DDR storage unit corresponding to the computing unit;

[0029] The processing unit is used to perform rotation position encoding processing on the input data and the matrix buffer according to the network function of the multi-head function layer to obtain output data.

[0030] In one possible implementation, when the computing unit is a computing unit that matches an accumulation processing functional layer located in an attention layer, the input data includes the first output data sent by the last computing unit that matches the previous network layer, and the second output data sent by the computing unit that matches the matrix-vector multiplication functional layer related to the result output located in the attention layer;

[0031] The utilizing the processing unit in the computing unit to perform corresponding functional processing on the input data according to the network function of the functional layer related to the computing unit to obtain output data includes:

[0032] Utilizing the processing unit in the computing unit, the first output data and the second output data are cumulatively processed according to the network function of the cumulative processing function layer in the attention layer to obtain output data.

[0033] In a possible implementation, when the computing unit is a computing unit matching an accumulation processing functional layer located in a feedforward neural network layer, the input data includes third output data sent by a computing unit matching a matrix-vector multiplication functional layer associated with a feature compression matrix, and fourth output data sent by a computing unit matching an accumulation processing functional layer located in an attention layer;

[0034] The utilizing the processing unit in the computing unit to perform corresponding functional processing on the input data according to the network function of the functional layer related to the computing unit to obtain output data includes:

[0035] Utilizing the processing unit in the computing unit, the third output data and the fourth output data are accumulated according to the network function of the accumulation processing function layer in the feedforward neural network layer to obtain output data.

[0036] In a possible implementation, when the computing unit is a computing unit matched with a matrix-vector multiplication function layer associated with a feature compression matrix, the input data includes fifth output data sent by a computing unit matched with a matrix-vector multiplication function layer associated with a feature mapping matrix, and sixth output data sent by a computing unit matched with a matrix-vector multiplication function layer associated with a feature change matrix;

[0037] The step of performing matrix-vector multiplication on the input data and the target data stored in the storage unit of the computing unit by using the processing unit in the computing unit to obtain output data includes:

[0038] The processing unit in the computing unit is used to perform matrix-vector multiplication processing on the fifth output data, the sixth output data, and the target data stored in the storage unit in the computing unit to obtain output data.

[0039] In one possible implementation, the sending of the output data to at least one other computing unit using the communication unit in the computing unit and the communication units in the other computing units according to the communication connections and communication configuration files set for the communication units in different computing units includes:

[0040] determining a target functional layer having a communication relationship with the functional layer matched with the computing unit;

[0041] determining each other computing unit associated with the target functional layer;

[0042] According to the communication connections and communication configuration files set for the communication units in different computing units, the output data is sent to each other computing unit respectively using the data sending interface included in the communication unit in the computing unit and the data receiving interface included in the communication units in each other computing unit.

[0043] In a second aspect, an embodiment of the present disclosure further provides a data processing device, including:

[0044] a processing module, configured to process input data using a processing unit in a computing unit that matches any network layer of a target neural network to obtain output data;

[0045] a sending module, configured to send the output data to at least one other computing unit using the communication unit in the computing unit and the communication units in other computing units, in accordance with the communication connections and communication configuration files set for the communication units in different computing units; wherein the other computing units include other computing units matching the network layer or computing units matching the next network layer;

[0046] The loop module is used to return to the step of using the processing unit in the computing unit to process the input data and obtain output data, until the output data of the last computing unit matching the last network layer of the target neural network is obtained, and the output data is used as the data processing result.

[0047] In a third aspect, an optional implementation of the present disclosure further provides a computer device, a processor, and a memory, wherein the memory stores machine-readable instructions executable by the processor, and the processor is used to execute the machine-readable instructions stored in the memory, and when the machine-readable instructions are executed by the processor, the steps of the above-mentioned first aspect or any possible implementation of the first aspect are performed.

[0048] In a fourth aspect, an optional implementation of the present disclosure further provides a computer program product, including a computer program, which, when executed, implements the above-mentioned first aspect, or the steps in any possible implementation of the first aspect.

[0049] For a description of the effects of the above-mentioned data processing apparatus, computer equipment, and computer program product, please refer to the description of the above-mentioned data processing method, which will not be repeated here.

[0050] The data processing method, apparatus, computer equipment and program product provided by the embodiments of the present disclosure deploy communication units in computing units, which can achieve data reception for computing units and flexible communication interconnection between multiple computing units at a relatively small unit area cost, thereby improving the interconnection performance and communication flexibility between multiple computing units. Moreover, compared with the prior art of using a graphics processor to complete large-scale calculations in the model, by pre-setting at least one matching computing unit for each network layer of the target neural network and ensuring that the computing unit can focus on the reasoning calculation of the corresponding network layer, it is possible to use lower-cost computing units to replace expensive graphics processors to complete the reasoning calculation of the model, thereby reducing both the reasoning cost of large models and the reasoning power consumption of large models. In addition, by matching different computing units to different target neural networks and setting communication lines and communication configuration files for each matched computing unit, each computing unit can perform data processing in an orderly manner according to the indicated processing order, thereby improving the accuracy of data processing.

[0051] Furthermore, the data processing methods, devices, computer equipment and program products provided by the embodiments of the present disclosure, since matrix-vector multiplication operations in the process of large-scale neural network inference require the use of frequent operators, and such operations will have a large-scale calculation ratio when calculated by the number of operands, it is sometimes difficult to cache the required huge matrix using a single computing unit. Therefore, the present application proposes to split the matrix using row or column splitting methods, and pre-store the sub-matrices obtained by the split in different computing units respectively, while ensuring that a computing unit only resides in the sub-matrix corresponding to a certain type of matrix. This effectively solves the calculation problem of large-scale matrices and reduces the power consumption caused by large-scale data transmission, and enables each computing unit to focus on the calculation of the corresponding network layer, thereby ensuring the accuracy of model inference.

[0052] In order to make the above-mentioned objectives, features and advantages of the present disclosure more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the following briefly introduces the drawings required for use in the embodiments. The drawings herein are incorporated into and constitute a part of the specification. These drawings illustrate embodiments consistent with the present disclosure and, together with the specification, are used to illustrate the technical solutions of the present disclosure. It should be understood that the following drawings only illustrate certain embodiments of the present disclosure and should not be regarded as limiting the scope. For those of ordinary skill in the art, other relevant drawings can be obtained based on these drawings without inventive effort.

[0054] Figure 1A flow chart of a data processing method provided by an embodiment of the present disclosure is shown;

[0055] Figure 2 A schematic diagram of the structure of a computing unit provided by an embodiment of the present disclosure is shown;

[0056] Figure 3 A schematic diagram of the structure of a network layer provided by an embodiment of the present disclosure is shown;

[0057] Figure 4 A schematic diagram of a data processing device provided by an embodiment of the present disclosure is shown;

[0058] Figure 5 A schematic structural diagram of a computer device provided by an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0059] In order to make the purpose, technical solutions and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, rather than all of the embodiments. The components of the embodiments of the present disclosure generally described and shown here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present disclosure is not intended to limit the scope of the present disclosure for protection, but merely represents the selected embodiments of the present disclosure. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present disclosure.

[0060] In addition, the terms "first," "second," and the like in the description and claims of the embodiments of the present disclosure and in the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, such that the embodiments described herein can be practiced in an order other than that shown or described herein.

[0061] In this document, "multiple or several" refers to two or more. "And / or" describes the relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, or B exists alone. The character " / " generally indicates that the associated objects are in an "or" relationship.

[0062] Research has found that in order to solve the power consumption and cost problems caused by large-scale neural network models during data inference, the traditional method has proposed using multiple chips to pre-store data to solve the data transmission problem. Although this method can reduce power consumption and cost, it reduces the flexibility of communication and has certain disadvantages.

[0063] Based on the above research, the present disclosure provides a data processing method, apparatus, computer equipment and program product. By deploying a communication unit in a computing unit, data reception for the computing unit can be achieved at a relatively small unit area cost, and flexible communication and interconnection between multiple computing units can be achieved, thereby improving the interconnection performance and communication flexibility between multiple computing units. Moreover, compared with the prior art of using a graphics processor to complete large-scale calculations in the model, by pre-setting at least one matching computing unit for each network layer of the target neural network and ensuring that the computing unit can focus on the reasoning calculation of the corresponding network layer, it is possible to use lower-cost computing units to replace expensive graphics processors to complete the reasoning calculation of the model, thereby reducing both the reasoning cost of large models and the reasoning power consumption of large models. In addition, by matching different computing units to different target neural networks and setting communication lines and communication configuration files for each matched computing unit, each computing unit can perform data processing in an orderly manner according to the indicated processing order, thereby improving the accuracy of data processing.

[0064] The defects in the above solutions are the results obtained by the inventors after practice and careful research. Therefore, the process of discovering the above problems and the solutions proposed by this disclosure for the above problems below should be the contributions made by the inventors to this disclosure during the disclosure process.

[0065] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings.

[0066] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0067] To facilitate understanding of this embodiment, a data processing method disclosed in an embodiment of the present disclosure is first introduced in detail. The executor of the data processing method provided in the embodiment of the present disclosure is generally a terminal device or other processing device with certain computing capabilities, where the terminal device can be a user equipment (User Equipment, UE), a mobile device, a user terminal, a terminal, a personal digital assistant (Personal Digital Assistant, PDA), a handheld device, a computer device, a computing unit (such as a chip), etc.; in some possible implementations, the data processing method can be implemented by a processor calling computer-readable instructions stored in a memory.

[0068] The data processing method provided by the embodiment of the present disclosure is described below by taking the execution subject as a computing unit as an example.

[0069] like Figure 1 FIG. 1 is a flowchart of a data processing method provided by an embodiment of the present disclosure, which may include the following steps:

[0070] S101: For a computing unit that matches any network layer of a target neural network, use a processing unit in the computing unit to process input data to obtain output data.

[0071] Here, the target neural network can be a network model with any reasoning function, for example, a large language model, a large language model based on Transformer, or other large-scale network models based on Transformer. In other words, the data processing method provided in the embodiments of the present disclosure can be applied to large language model reasoning scenarios.

[0072] The target neural network can be a single-layer network model including one network layer, or a multi-layer network model including multiple network layers, which is not specifically limited in the embodiments of the present disclosure. For each network layer, at least one computing unit that matches the network layer can be predetermined, and the processing order of each computing unit is related to the inference order of the target neural network. For example, a matching computing unit cluster (system) can be set up for the target neural network in advance, and the cluster includes various computing units that match the various network layers of the target neural network, and each computing unit is relatively independent and completely consistent in design. The control and coordination between the various computing units is achieved through different pre-set communication connections and communication configuration files, ultimately realizing the functions of a large-scale neural network model.

[0073] The computing units in the disclosed embodiments are small-scale units with lower cost, power consumption, and storage capacity than conventional image processors. However, by configuring a matching computing unit cluster for the target neural network, the image processor's functionality can be replaced during model inference, reducing power consumption and cost while ensuring model inference. Alternatively, the computing units in the disclosed embodiments can be pre-developed chips.

[0074] like Figure 2As shown, it is a structural diagram of a computing unit provided in an embodiment of the present disclosure, and the computing unit includes a processing unit, a storage unit and a communication unit. The processing unit is a unit with a computing function, which can be used for data calculation. For example, the processing unit can implement matrix-vector multiplication calculation, nonlinear calculation, element-wise calculation, etc. Optionally, the processing unit can be implemented using in-memory calculation or near-memory calculation. The storage unit is used for data storage. For example, the sub-matrix mentioned later can be pre-stored. The storage unit can be, for example, a static random access memory (SRAM).

[0075] The communication unit is used to realize the communication between different computing units, that is, the computing units can be organized together by communication units to jointly solve large-scale computing problems; in terms of overall structure and area, the storage unit occupies a larger area, followed by the processing unit and the computing unit. The core of this disclosure is how to organize the computing units through the design of communication to jointly complete the single-layer and multi-layer reasoning of the large language model. The communication unit is in Figure 2 Specifically, it may include an inter-chip interconnect interface and a double data rate (DDR) interface. The inter-chip interconnect interface is used to implement communication interconnection between computing units, and may include, for example, a serial peripheral interface (SPI) and multiple communication interfaces. Multiple communication interfaces may include, for example, N groups of transmit interfaces and N groups of receive interfaces, or N groups of transmit and receive interfaces, where N can be 8, 16, 32, etc. The SPI interface is used to implement interconnection between computing units. For example, different computing units can be connected according to their interface functions through the SPI interface, thereby implementing communication between units. The transmit interface is used to implement data transmission, and the receive interface is used to implement data interface. The transmit and receive interface is an interface with flexible configurable transmit and receive modes. For any group of transmit and receive interfaces, the transmit and receive mode of the interface can be controlled by software. For example, when a computing unit needs to receive data, the transmit and receive interface can be configured to receive mode through software; when a computing unit needs to send data, the transmit and receive interface can be configured to send mode through software. The DDR interface is used to connect to a storage unit outside the computing unit to store data processed by the computing unit in the external storage unit, where the external storage unit may be, for example, a DDR storage unit.

[0076] The communication unit can send the calculation results of the processing unit to one or more transmission interfaces, and then send the calculation results to other computing units through the communication connection (such as SPI connection) on the circuit board corresponding to the computing unit and the receiving interface of other computing units in advance. In addition, through the connection on the circuit board, different master / slave relationships can be configured for the same computing unit, thereby achieving orderly data flow between multiple computing units.

[0077] The target neural network may include one or more network layers, and each network layer has at least one matching computing unit. For any computing unit matched by any network layer, the input data is the data input to the computing unit, and the data can be processed by the processing unit in the computing unit. The input data of computing units matched by different network layers are different. For example, the input data of a computing unit may be at least one of the output data sent by other computing units matched by the previous network layer through the internal communication unit, the initial input data for the target neural network, and the output data sent by other computing units matched by the network layer through the communication unit. Among them, the initial input data can be regarded as the original data that needs to be input to the target neural network.

[0078] In one embodiment, each network layer of the target neural network may include multiple functional layers, and different functional layers have different network functions. Specifically, the multiple functional layers include at least a normalization processing functional layer (norm) located in the attention layer, a matrix-vector multiplication functional layer (matmul(QKV)) associated with the query matrix, the key matrix, and the value matrix, a multi-head functional layer (multihead), a matrix-vector multiplication functional layer (matmul(Output)) associated with the result output, an accumulation processing functional layer (Add), and a normalization processing functional layer (Norm) located in the feedforward neural network layer (FNN), a matrix-vector multiplication functional layer associated with multiple weight matrices, and an accumulation processing functional layer (Add). The weight matrix may specifically include a feature mapping matrix (hereinafter referred to as the W1 matrix), a feature transformation matrix (hereinafter referred to as the W3 matrix) and a feature compression matrix (hereinafter referred to as the W2 matrix). The matrix-vector multiplication function layers respectively associated with the multiple weight matrices may include a matrix-vector multiplication function layer (matmul(W1)) associated with the W1 matrix, a matrix-vector multiplication function layer (matmul(W2)) associated with the W2 matrix and a matrix-vector multiplication function layer (matmul(W3)) associated with the W3 matrix. The computing unit that matches any network layer of the target neural network is associated with a functional layer among the multiple functional layers. That is, in the embodiment of the present disclosure, each functional layer may have at least one matching computing unit, and each computing unit that matches any network layer of the target neural network is a computing unit respectively associated with each functional layer in the network layer. The computing units that match each functional layer can be pre-set before using the target neural network for model inference.

[0079] like Figure 3 As shown, it is a schematic diagram of the structure of a network layer provided by an embodiment of the present disclosure. Figure 3In , a structural diagram of the current network layer is shown. The last Block can be understood as the previous network layer of the current network layer, and the next Block can be understood as the next network layer of the current network layer. The current network layer includes the Attention layer and the FFN layer. The Attention layer includes the input (Input) layer, the Norm layer, the matmul (QKV) layer, the multihead layer, the matmul (Output) layer, and the Add layer in sequence. The processing order and data input relationship between each layer are shown by the arrows. In the case where the computing unit matches the input layer of the non-first network layer of the target neural network, the input data can be the data sent by the computing unit that matches the previous network layer of the non-first network layer through the internal communication unit, and the data can be the output data of the computing unit that matches the previous network layer. For example, in Figure 3 In , the input data of the Input layer can be the output data sent by the computing unit that matches the previous network layer through the internal communication unit. In the case where the computing unit matches the input layer of the first network layer of the target neural network, the input data can be the initial input data for the target neural network. In the case where the computing unit matches the non-input layer of any network layer of the target neural network, the input data can be the output data sent by the computing unit that matches other functional layers of the network layer through the internal communication unit. For example, in Figure 3 In the matmul(QKV) layer, the input data of any computing unit matched by the Norm layer can be the output data sent by the computing unit matched by the Norm layer. For example, Figure 3 The input data of the computing unit matched by the Add layer in the attention layer may include the output data of the Input layer and the output data sent by the computing unit matched by the matmul(Output) layer through the communication unit.

[0080] The FFN layer includes the Norm layer, matmul(W1) layer, matmul(W3) layer, matmul(W2) layer, Add layer, and output layer, which appear in sequence. The processing order and data input relationship between each layer are indicated by arrows. The input of the FFN layer is the output of the entire Attention layer, and the output of the FFN layer is the output of the entire network layer. The input data of the matmul(W3) layer can include the output data sent by the matching computation unit of the Norm layer. The input data of the matmul(W2) layer can include the output data sent by the matching computation unit of the matmul(W1) layer and the output data sent by the matching computation unit of the matmul(W3) layer. The input data of the Add layer can include the output data of the Attention layer and the output data sent by the matching computation unit of the Add layer of the FFN layer. It is understood that the Input and Output layers are only used for data input and output and do not perform data calculation. Therefore, the Input and Output layers do not need to have matching computation units.

[0081] Exemplarily, a corresponding computing unit cluster can be set up in advance for the target neural network, and the communication unit can be used to realize the interconnection between the various computing units. Then, before executing S101, for any computing unit that matches any network layer of the target neural network, the communication unit in the computing unit can be used to obtain input data that matches the computing unit; the input data includes at least one of the output data sent by other computing units that match the previous network layer through the communication unit, the initial input data for the target neural network, and the output data sent by other computing units that match the network layer through the communication unit. For example, for any computing unit, the data receiving interface in the communication unit in the computing unit can be used to obtain input data related to the computing unit. Furthermore, the processing unit in the computing unit can be used to process the input data to obtain output data.

[0082] The data processing required by the processing unit in the computing unit is related to the network function of the functional layer matched with the computing unit. The data processing may include matrix-vector multiplication, accumulation processing, rotation position encoding, nonlinear calculation, etc.

[0083] During specific implementation, for any computing unit, after obtaining the input data using the communication unit of the computing unit, the processing unit in the computing unit can continue to be used to perform corresponding data processing on the input data according to the network function of the functional layer matched by the computing unit to obtain the data processing result, and use the result as the output data of the computing unit.

[0084] S102: According to the communication connections and communication configuration files set for the communication units in different computing units, the output data is sent to at least one other computing unit using the communication unit in the computing unit and the communication units in other computing units; wherein the other computing units include other computing units that match the network layer, or computing units that match the next network layer.

[0085] Here, the communication profile is used to indicate the communication parameters and communication configuration between the computing units. The communication profile may include a communication profile corresponding to each computing unit, or it may be an overall profile for all computing units that match the target neural network. Different target neural networks require different computing unit clusters to be set up, and the communication connections and communication profiles between the computing units in the computing unit clusters are different. That is, for different target neural networks, different numbers of computing units can be set up in advance, and different communication connections and communication profiles can be set up for each computing unit. Exemplarily, for any target neural network, the number of computing units required to be set up for the target neural network, the communication connections between the computing units, and the communication profiles can be determined according to the function of the target neural network.

[0086] Other computing units are computing units that need to perform the next step of computational processing. Other computing units can be determined based on pre-set communication connections and communication configuration files. For example, after the computing unit matching the Norm layer obtains output data, the other computing units can be the computing units matching the matmul (QKV) layer. For another example, after the Output layer of the FFN layer obtains output data, the other computing units can be the computing units matching the input layer of the next network layer. In this way, data processing can be achieved through direct communication between computing units, which can improve data processing efficiency.

[0087] In a specific implementation, after any computing unit determines the output data, it can determine the other computing units that need to perform the next step of computational processing based on the communication configuration file, the preset communication connection that matches the computing unit, the functional layer that matches the computing unit, and the processing order between the functional layers in the target neural network. Then, the communication unit in the computing unit and the communication units in each other computing unit can be used to send the output data of the computing unit to each other computing unit respectively, so that each other computing unit can perform the next step of model inference.

[0088] S103: Return to the step of using the processing unit in the computing unit to process the input data to obtain output data, until the output data of the last computing unit matching the last network layer of the target neural network is obtained, and the output data is used as the data processing result.

[0089] Here, the last computing unit matched by the last network layer can specifically be the output layer of the last network layer in the target neural network. The data processing result is the final processing result output by the target neural network after performing data inference on the initial input data.

[0090] Exemplarily, after the output data of any computing unit is sent to other computing units, the output data can be regarded as the input data of the other computing units. Then, for each other computing unit, the step of S101 can be returned to be executed, thereby continuing to use other computing units to complete further reasoning processing of the data. After obtaining the output data, the other computing units can continue to send the output data to new other computing units through their internal communication units for further processing. After being processed in sequence by each computing unit that matches the target neural network, the data processing result output by the output layer of the last network layer is obtained.

[0091] It is understandable that before the computing unit receives the data or before the computing unit sends the data to other computing units, the input data can also be processed separately using the provided additional processing unit. The separate processing here is only for the input data itself, and no interaction between the input data and other data occurs. For example, the separate processing can include square root calculation, natural exponential calculation, natural logarithm calculation, silicon rectified linear unit (silu) calculation, etc. of the input data.

[0092] In this way, by deploying the communication unit in the computing unit, it is possible to achieve data reception for the computing unit and flexible communication interconnection between multiple computing units at a relatively small unit area cost, thereby improving the interconnection performance and communication flexibility between multiple computing units. Moreover, compared with the prior art of using a graphics processor to complete large-scale calculations in the model, by pre-setting at least one matching computing unit for each network layer of the target neural network and ensuring that the computing unit can focus on the inference calculation of the corresponding network layer, it is possible to use lower-cost computing units to replace expensive graphics processors to complete the inference calculation of the model, thereby reducing both the inference cost of large models and the inference power consumption of large models. In addition, by matching different computing units to different target neural networks and setting communication lines and communication configuration files for each matched computing unit, each computing unit can perform data processing in an orderly manner according to the indicated processing order, thereby improving the accuracy of data processing.

[0093] In one embodiment, the specific implementation process of S101 may be implemented in different ways depending on whether the computing unit is related to a matrix-vector multiplication function layer. The matrix-vector multiplication function layer may include, for example, the matmul(QKV) layer, matmul(Output) layer, matmul(W1) layer, matmul(W3) layer, and matmul(W2) layer mentioned above.

[0094] Specifically, when the computing unit is associated with any matrix-vector multiplication functional layer, the processing unit in the computing unit is used to perform matrix-vector multiplication processing on the input data and the target data stored in the storage unit in the computing unit to obtain output data; wherein the target data includes any preset matrix corresponding to the functional layer matching the computing unit or a submatrix obtained by matrix splitting any preset matrix, and the preset matrix includes a query matrix, a key matrix, a value matrix, a result output matrix and a weight matrix.

[0095] Here, for any network layer, any functional layer in the network layer has a corresponding network function. Moreover, for any matrix-vector multiplication functional layer, there is a corresponding preset matrix, and the matrix size of the preset matrix is ​​related to the processing power of the target neural network. Matrix-vector multiplication refers to the process of multiplying a vector of length N and a matrix of size M×N to obtain a vector of length M. In this process, the size of the matrix is ​​MxN. When M and N are very large (such as 10000X10000), it is often difficult to cache such a large matrix on the chip. Therefore, it is necessary to split it into multiple matrices, calculate the partial sum of multiple matrices separately, and realize multiplication and addition calculations.

[0096] For example, for the matmul(QKV) layer, the corresponding preset matrix may include a preset query matrix (i.e., Q matrix), a preset key matrix (i.e., K matrix), and a preset value matrix (i.e., V matrix). For the matmul(Output) layer, the corresponding preset matrix may include a preset result output matrix (i.e., Output matrix). For the matmul(W1) layer, the corresponding preset matrix may include a preset W1 matrix, for the matmul(W2) layer, the corresponding preset matrix may include a preset W2 matrix, and for the matmul(W3) layer, the corresponding preset matrix may include a preset W3 matrix.

[0097] Taking the Q matrix, K matrix and V matrix as examples, the Q matrix, K matrix and V matrix can each include at least one matching computing unit. For example, for the Q matrix, if the computing unit is able to process the matrix-vector multiplication operation corresponding to the matrix, the matrix can be deployed in a computing unit, and the computing unit is only used to process operations related to the Q matrix. If the computing unit cannot process the matrix-vector multiplication operation corresponding to the matrix, the Q matrix can be split into multiple sub-matrices according to the processing capability of the computing unit, and the multiple sub-matrices are respectively deployed in different computing units. These computing units are only used to process operations related to the deployed sub-matrices, and the processing results of these computing units are accumulated to obtain the complete results corresponding to the Q matrix.

[0098] Therefore, for any computing unit related to any matrix-vector multiplication functional layer, the storage unit of the computing unit can pre-store a preset matrix corresponding to the matrix-vector multiplication functional layer or a sub-matrix split from the preset matrix, wherein the preset matrix or sub-matrix stored in the storage unit can be called target data.

[0099] Exemplarily, for any computing unit that matches the Q matrix, the processing unit in the computing unit can be used to perform matrix-vector multiplication on the input data and the target data stored in the storage unit of the computing unit (the Q matrix or a submatrix of the Q matrix) to obtain output data. For any computing unit that matches the K matrix, the processing unit in the computing unit can be used to perform matrix-vector multiplication on the input data and the target data stored in the storage unit of the computing unit (the K matrix or a submatrix of the K matrix) to obtain output data. For any computing unit that matches the V matrix, the processing unit in the computing unit can be used to perform matrix-vector multiplication on the input data and the target data stored in the storage unit of the computing unit (the V matrix or a submatrix of the V matrix) to obtain output data. For any computing unit that matches the Output matrix, the processing unit in the computing unit can be used to perform matrix-vector multiplication on the input data and the target data stored in the storage unit of the computing unit (the Output matrix or a submatrix of the Output matrix) to obtain output data. For any computational unit that matches the Wx matrix, the processing unit in the computational unit can be used to perform matrix-vector multiplication on the input data and the target data (the Wx matrix or a submatrix of the Wx matrix) stored in the storage unit of the computational unit to obtain the output data. Here, the value of x is 1, 2, or 3.

[0100] In this way, each computing unit is related to a preset matrix corresponding to a certain functional layer, and is only used to perform operations related to the stored target data in a certain reasoning operation, which refines the processing operations of each computing unit and reduces the processing difficulty of the computing unit.

[0101] In another embodiment, when the computing unit is independent of the matrix-vector multiplication functional layer, the processing unit in the computing unit is used to perform corresponding functional processing on the input data according to the network function of the functional layer related to the computing unit to obtain output data.

[0102] Here, the computing units that are not related to the matrix-vector multiplication function layer may include computing units matched by the Norm layer, computing units matched by the multihead layer, computing units matched by the Add layer, and the like.

[0103] Network functions may include, for example, the normalization processing function of the Norm layer, the multi-head attention processing function of the multihead layer, and the accumulation processing function of the Add layer.

[0104] For example, for any computing unit that matches the Norm layer, the processing unit in the computing unit can be used to normalize the input data to obtain output data. For any computing unit that matches the multihead layer, the processing unit in the computing unit can be used to perform multi-head attention processing on the input data to obtain output data. For any computing unit that matches the Add layer, the processing unit in the computing unit can be used to accumulate the input data to obtain output data.

[0105] In one embodiment, the output data of a certain computing unit may need to be sent to multiple computing units matched by the next functional layer. For example, the output data of the computing unit matched by the Norm layer may need to be sent synchronously to the computing units in the matmul (QKV) layer that match the Q matrix, K matrix, and V matrix respectively. In this case, the communication unit in the embodiment of the present disclosure may have a broadcast function, based on which a copy of the output data can be sent to the communication units in multiple computing units, thereby ensuring that each computing unit can use the output data. Similarly, a computing unit may need to synchronously receive output data from multiple other computing units. For example, for the computing unit matched by the multihead layer, it may need to synchronously receive output data from computing units that match the Q matrix, K matrix, and V matrix respectively. Therefore, the step of obtaining input data for the computing unit in S101 above can be implemented as follows:

[0106] In the case where the input data includes output data sent from multiple other computing units, the output data of each computing unit is received and cached separately using multiple data receiving interfaces in the communication unit included in the computing unit; in response to each data receiving interface completing the caching of the corresponding packet data, the computing unit is used to trigger a reception completion signal to determine that the input data acquisition is completed.

[0107] Here, for any computing unit, each data receiving interface included in the communication unit in the computing unit has a communication caching function. In the case where multiple groups of data receiving interfaces of a computing unit need to simultaneously receive output data from multiple computing units matched to a certain functional layer, the output data of multiple computing units can be received separately by using multiple data receiving interfaces in the communication unit according to the communication connection and communication configuration file. That is, the output data of one computing unit is received using one data receiving interface. At the same time, for each data receiving interface, the received output data can be received and cached. After each data receiving interface has cached the output data of the corresponding computing unit, the computing unit can trigger a reception completion signal to use the signal to indicate that the synchronous reception of the output data from multiple computing units is completed, thereby achieving complete acquisition of the input data.

[0108] That is, in the embodiment of the present disclosure, when the multiple data receiving interfaces of the computing unit are in a working state of simultaneously receiving data, a reception completion signal is further provided to indicate that the synchronous reception of the multiple data sets is complete. The reception completion signal is triggered only when all the data corresponding to the multiple data receiving interfaces have been received.

[0109] Optionally, when sending output data of any computing unit to other computing units, the computing unit may split the output data into multiple groups of data. In this case, the other computing units will use multiple data receiving interfaces to receive each group of data separately and receive and buffer them. When all the multiple groups of data are received, a reception completion signal is triggered.

[0110] In one embodiment, to provide a backup, the present disclosure can also implement board-level redundancy. Specifically, at least one backup computing unit can be provided to enable prompt replacement in the event of a failure. Therefore, the present disclosure can also perform functional and performance testing on the computing units; if the testing fails, the backup computing unit can be used to replace the computing unit.

[0111] Here, the functional test is used to test whether the function of the computing unit is normal, and the performance test is used to test whether the performance of the computing unit can meet the preset standard.

[0112] Exemplarily, a detection program can be provided. For any computing unit that matches the target neural network, the detection program can be used to perform functional detection and performance detection on the computing unit. If both tests pass, it means that the computing unit can be used normally. If one of the two tests fails, it can be said that the computing unit has a fault and fails the test. At this time, the computing unit can be replaced with a spare computing unit through the SPI interface and routing protocol to ensure the normal operation of subsequent model reasoning. Optionally, the spare computing unit can be a computing unit that can only receive a single input and a single output, and has poorer flexibility of use than the replaced computing unit.

[0113] In one embodiment, the weight matrix includes a feature mapping matrix, a feature transformation matrix, and a feature compression matrix. When the preset matrix is ​​any one of a query matrix, a key matrix, a value matrix, a feature mapping matrix, and a feature transformation matrix, the submatrix included in the target data is a matrix obtained by splitting the preset matrix into rows. When the preset matrix is ​​any one of a result output matrix and a feature compression matrix, the submatrix included in the target data is a matrix obtained by splitting the preset matrix into columns.

[0114] That is, the embodiment of the present disclosure provides two matrix splitting methods, namely row splitting (i.e. splitting by rows) and column splitting (i.e. splitting by columns). Different splitting methods will be used for the corresponding preset matrices in different functional layers. For row splitting, the communication unit can be used to broadcast the output data of other computing units, so as to send the output data to the computing unit storing each sub-matrix obtained by row splitting. Each computing unit can complete M rows of matrix-vector multiplication operations and output M results. Among them, the number of M is related to the number of rows included in the sub-matrix. Then, the M results output by each computing unit can be input to the next node (i.e. other computing units) to realize the splicing of the M results output by each computing unit to obtain the matrix-vector multiplication result for the preset matrix.

[0115] For column splitting, the communication unit can be used to broadcast the output data of other computing units, so as to realize sending the output data to the computing unit that stores each sub-matrix obtained by column splitting. Each computing unit can use the stored sub-matrix to complete the matrix-vector multiplication operation, obtain the sub-result, and then send the sub-result to the adjacent computing unit (i.e., the computing unit matched by the Add layer) for accumulation. If it is a one-way accumulation, the matrix-vector multiplication result for the preset matrix can be obtained by accumulating the sub-results sent by the computing units that store the sub-matrices of each column split. For each computing unit that stores the sub-matrix of the column split, after the computing unit sends the sub-result to the adjacent computing unit, the calculation result of the computing unit will be cleared. In this way, after the one-way transmission of each computing unit, the accurate matrix-vector multiplication result can be accumulated.

[0116] It is understandable that the embodiments of the present disclosure also involve splitting rows first and then columns, wherein row splitting can give partial numerical values ​​of the result, and the output of row splitting is used as the input of column splitting. After the column splitting completes the calculation, the results of column splitting are accumulated to obtain the final result of matrix-vector multiplication. Among them, row first and column second can be reflected in that after the calculation units respectively matching the Q matrix, K matrix and V matrix obtain the output data, they will eventually be transmitted to the calculation unit matching the Output matrix for further processing; and it can be reflected in that after the calculation units respectively matching the W1 matrix and the W3 matrix obtain the output data, they will be transmitted to the calculation unit matching the W2 matrix for further processing.

[0117] In one embodiment, if a preset matrix is ​​split into multiple sub-matrices, the multiple sub-matrices are stored in storage units of different computing units. A storage unit of a computing unit will only store one sub-matrix associated with a particular preset matrix, and the computing unit will only be used to perform matrix-vector multiplication operations on that sub-matrix and will not perform matrix-vector multiplication operations on other sub-matrices. The size of each sub-matrix corresponding to a preset matrix is ​​related to at least one of the size of the preset matrix, the storage space of the storage unit in the computing unit, and the computing power of the processing unit in the computing unit.

[0118] Taking the Q matrix as an example, how to split the Q matrix into rows can be determined based on one or more of the size of the Q matrix, the storage space of the storage unit in the computing unit, and the computing power of the processing unit in the computing unit, thereby obtaining multiple sub-matrices corresponding to the Q matrix. The size of each sub-matrix is ​​smaller than the storage space of the storage unit in the computing unit, and the calculation can be completed when the matrix-vector multiplication is performed on the sub-matrix and the input data using the processing unit in the computing unit.

[0119] Taking the W2 matrix as an example, how to split the W2 matrix into columns can be determined based on one or more of the size of the W2 matrix, the storage space of the storage unit in the computing unit, and the computing power of the processing unit in the computing unit, so as to obtain multiple sub-matrices corresponding to the W2 matrix.

[0120] Continue to refer Figure 3 , briefly explain the principle of matrix splitting, Figure 3In the example, the input data is an 8×1 vector. After the 8×8 preset matrix is ​​split into rows, a 2×8 sub-matrix (such as a sub-matrix filled with black) is obtained. In this way, the 8×8 preset matrix will be split into 4 sub-matrices and stored in the storage space of 4 different computing units. After broadcasting the 8×1 vector to these 4 computing units, each computing unit can perform a matrix-vector multiplication operation on the 8×1 vector and the 2×8 sub-matrix to obtain the output data (i.e., a 2×1 vector). By splicing the output data of the 4 computing units, the 8×1 matrix-vector multiplication result can be obtained.

[0121] For column splitting, the input data is the output data of the previous calculation unit (i.e., a 2×1 vector). After the 8×8 preset matrix is ​​split into columns, an 8×2 sub-matrix is ​​obtained. Such sub-matrices include 4 and are stored in the storage space of 4 different calculation units. After broadcasting the 2×1 vector to these 4 calculation units, each calculation unit can perform a matrix-vector multiplication operation on the 2×1 vector and the 8×2 sub-matrix to obtain the output data (i.e., an 8×1 vector). By accumulating the output data of the 4 calculation units, the final matrix-vector multiplication result can be obtained.

[0122] exist Figure 3 , the Q matrix, K matrix and V matrix in the matmul(QKV) layer are split into rows and stored in storage units in different computing units. The Output matrix in the matmul(Output) layer is split into columns and stored in storage units in different computing units. The W1 matrix in the matmul(W1) layer is split into rows and stored in storage units in different computing units. The W3 matrix in the matmul(W3) layer is split into rows and stored in storage units in different computing units. The W2 matrix in the matmul(W2) layer is split into columns and stored in storage units in different computing units. Afterwards, the data inference process can be implemented using each computing unit according to the data inference process described in the above embodiments.

[0123] In one embodiment, to facilitate understanding of the embodiments of the present disclosure, the processing of the computing unit that matches some functional layers will be described below. For example, with respect to the step of "using the processing unit in the computing unit to perform corresponding functional processing on the input data according to the network function of the functional layer related to the computing unit to obtain output data", when the computing unit is a computing unit that matches the multi-head functional layer located in the attention layer, the communication unit in the computing unit is used to obtain the matrix cache related to the key matrix and the value matrix from the double-rate DDR storage unit corresponding to the computing unit. The processing unit can then be used to perform rotational position encoding processing on the input data and the matrix cache according to the network function of the multi-head functional layer to obtain output data.

[0124] Here, multihead can be implemented in a single-layer network function layer, and the data processing operations in multihead may include rotation position encoding operations, which mainly perform bitwise multiplication of vectors and vector addition and subtraction. One part of the two-part vector comes from the output of each computing unit matched by the matmul (QKV) layer, and the other part needs to be actively loaded by the computing unit matched by the multihead layer. The data is cache data related to the K matrix and the V matrix, which can be called KVCache. KVCache can be obtained by the computing unit matched by the matmul (QKV) layer during the process of matrix-vector multiplication and sent to the DDR storage unit through the DDR interface for storage. In the embodiment of the present disclosure, there is an association relationship between the computing unit and the DDR storage unit, and multiple computing units can be associated with the same DDR storage unit. The DDR storage unit is located outside the computing unit, and each computing unit can communicate with the associated DDR storage unit through the internal DDR interface.

[0125] Exemplarily, for the computing unit that matches the multihead layer, the computing unit can receive the output data sent by each computing unit that matches the matmul (QKV) layer, and use these output data as input data. At the same time, the DDR interface in the communication unit inside itself can be used to obtain all the KVCache stored in the unit from the corresponding DDR storage unit. Then, the processing unit can be used to perform rotational position encoding processing on the input data and the obtained KVCache according to the multi-head attention processing function corresponding to the multihead layer to obtain the output data corresponding to the multihead layer. Optionally, when using the computing unit for data reasoning, the embodiment of the present disclosure can also use the Group Query Attention (GQA) method to reduce the KVCache usage in the calculation process.

[0126] Optionally, in the nonlinear spike framework within the multihead layer, the computing units that match the multihead layer can also be used to perform processing such as normalization (softmax) and scalar division on the input data. There may also be matrix-vector multiplication operations within the multihead layer, but the matrix-vector multiplication operations here may be smaller in scale and can be completed using computing units that perform rotational position encoding. Of course, it can also be implemented using other computing units preset for the multihead layer, or it can be completed using computing units that match the matmul (QKV) layer. Here, since the computing units that match the multihead layer have already been processed and are idle when the computing units that match the multihead layer are processing, the calculation can be completed by reusing the idle computing units.

[0127] In one embodiment, when the computing unit is a computing unit that matches the accumulation processing functional layer located in the attention layer, the input data includes the first output data sent by the last computing unit that matches the previous network layer, and the second output data sent by the computing unit that matches the matrix-vector multiplication functional layer related to the result output located in the attention layer.

[0128] Here, combined Figure 3 It can be seen that the previous network layer of the Add layer in the Attention layer is the previous Block. The first output data sent by the last computing unit that matches the previous network layer is the output data of the Output layer in the previous Block, and is also the data of the Input layer of the current network layer in the input. The matrix-vector multiplication function layer related to the result output in the attention layer is the matmul (Output) layer, and the second output data is the output data of the computing unit matched with the matmul (Output) layer. It should be noted that when data transmission between different computing units is involved in the embodiments of the present disclosure, it is necessary to rely on the communication unit set inside the computing unit to complete.

[0129] Exemplarily, for any computing unit that matches the Add layer in the attention layer, the input data of the computing unit includes the output data of each computing unit matched by the matmul(Output) layer and the output data of the Output layer of the previous network layer.

[0130] Regarding the step of "using the processing unit in the computing unit to perform corresponding functional processing on the input data according to the network function of the functional layer related to the computing unit to obtain output data", the processing unit in the computing unit can be used to perform accumulation processing on the first output data and the second output data according to the network function of the accumulation processing functional layer in the attention layer to obtain output data.

[0131] Illustratively, for a computing unit that matches the Add layer, the computing unit may utilize a processing unit to perform accumulation processing on the first output data and each second output data according to an accumulation function, thereby obtaining the output data of the computing unit.

[0132] In another embodiment, when the computing unit is a computing unit that matches the accumulation processing functional layer located in the feedforward neural network layer, the input data includes the third output data sent by the computing unit that matches the matrix-vector multiplication functional layer related to the feature compression matrix, and the fourth output data sent by the computing unit that matches the accumulation processing functional layer located in the attention layer.

[0133] Here, combined Figure 3 It can be seen that the third output data can be the output data of each computing unit matching the matmul(W2) layer; the fourth output data is the output data of the computing unit matching the Add layer in the Attention layer.

[0134] Exemplarily, for any computing unit that matches the Add layer in the FFN, the input data of the computing unit may include the output data of each computing unit that matches the matmul (W2) layer, and the computing unit that matches the Add layer in the Attention layer.

[0135] Regarding the step of "using the processing unit in the computing unit to perform corresponding functional processing on the input data according to the network function of the functional layer related to the computing unit to obtain output data", the processing unit in the computing unit can be used to perform accumulation processing on the third output data and the fourth output data according to the network function of the accumulation processing functional layer in the feedforward neural network layer to obtain the output data.

[0136] Illustratively, for a computing unit that matches the Add layer in the FFN layer, the computing unit may utilize a processing unit to perform accumulation processing on each third output data and fourth output data according to an accumulation function, thereby obtaining output data of the computing unit.

[0137] In another embodiment, when the computing unit is a computing unit matched with a matrix-vector multiplication functional layer associated with a feature compression matrix, the input data includes the fifth output data sent by the computing unit matched with a matrix-vector multiplication functional layer associated with a feature mapping matrix, and the sixth output data sent by the computing unit matched with a matrix-vector multiplication functional layer associated with a feature change matrix.

[0138] Here, combined Figure 3It can be seen that for any computing unit matching the matmul(W2) layer, the fifth output data can be the output data of each computing unit matching the matmul(W1) layer; the sixth output data is the output data of each computing unit matching the matmul(W3) layer.

[0139] Exemplarily, the output data of each computing unit matching the matmul(W1) layer can be sent to each computing unit matching the matmul(W2) layer by broadcasting, and the output data of each computing unit matching the matmul(W3) layer can also be sent to each computing unit matching the matmul(W2) layer by broadcasting, and the output data of each computing unit matching the matmul(W3) layer can also be sent to each computing unit matching the matmul(W2) layer by broadcasting.

[0140] Furthermore, for the step of “using the processing unit in the computing unit to perform matrix-vector multiplication on the input data and the target data stored in the storage unit in the computing unit to obtain output data”, the processing unit in the computing unit can be used to perform matrix-vector multiplication on the fifth output data, the sixth output data and the target data stored in the storage unit in the computing unit to obtain output data.

[0141] Exemplarily, for any computing unit matched with the matmul(W2) layer, the processing unit in the computing unit can be used to perform matrix-vector multiplication processing on each fifth output data, each sixth output data, and the target data stored in the storage unit of the computing unit to obtain output data. Specifically, each fifth output data can be subjected to matrix-vector multiplication processing with the target data, and each sixth output data can be subjected to matrix-vector multiplication processing with the target data to obtain output data related to each fifth output data and output data related to each sixth output data. Afterwards, each output data calculated by the computing unit can be sent to the Add layer in the FFN layer for accumulation processing to obtain the output data of the corresponding network layer.

[0142] In one embodiment, S102 may be implemented as follows:

[0143] Determine a target functional layer that has a communication relationship with the functional layer matched with the computing unit; determine each other computing unit associated with the target functional layer; and according to the communication connections and communication configuration files set for the communication units in different computing units, use the data sending interface included in the communication unit in the computing unit and the data receiving interface included in the communication unit in each other computing unit to send the output data to each other computing unit respectively.

[0144] Here, the target functional layer that communicates with any functional layer is the functional layer that the output data of the functional layer needs to be input. For example, the target functional layer corresponding to the matmul(QKV) layer can be the multihead layer; the target functional layer corresponding to the Add layer in the Attention layer can be the Norm layer in the FFN layer and the Add layer in the FFN layer; the target functional layer corresponding to the Norm layer in the FFN layer can be the matmul(W1) layer and the matmul(W3) layer.

[0145] Optionally, the target functional layer may be determined according to at least one of a communication configuration file, a communication connection, and a processing sequence between functional layers.

[0146] After determining the target functional layer, the computing units that match the target functional layer can be determined and used as the other computing units with which the current computing unit needs to communicate. Then, according to the communication connections set for the communication units in the other computing units and the communication configuration files associated with the other computing units, the output data of the current computing unit is sent to the other computing units using the data transmission interface of the communication unit in the current computing unit, the data reception interface of the communication units in the other computing units, and the SPI connection between the current computing unit and the other computing units, so that the other computing units can further process the data.

[0147] To facilitate understanding of the embodiments of the present disclosure, the following will be explained using an example of using a computing unit to implement reasoning of a single-layer target neural network. In this example, the computing unit is a chip, and the chip's storage unit SRAM = 32MB + 1MB. The parameters of the single-layer target neural network are as follows: hidden size hidden_size = 8192, intermediate size intermediate_size = 28672, head head = 64, key-value pair head nukmakvheads = 8, and the corresponding parameter quantities of the single-layer target neural network are:

[0148] 8192*8192(Q)+8192*(64heads*16)*2(KV)+8192*8192(Attention)+8192*28672*3(FFN)=8192*8192*12.675=850M. This shows that the corresponding parameter count for this single-layer target neural network is 850M parameters. Assuming that the parameters are stored in 8 bits, the memory required for each head in the Attention layer (when calculating the KVCache, assuming the string length len = 4096) is: 8192*128+8192*16*2+8192*128+len*16*2=2.5M. In the Attention layer, 12 heads can be placed on the same chip, so a total of 6 chips are required. In the input stage, the input data can be sent to the six chips through broadcasting; in the process of calculating the Q, K, and V matrices, row splitting can be used to calculate the final results of some positions; in the process of calculating the output, column splitting is used to calculate the partial results of all positions, and then finally in the output, the entire result of the attention is output in the form of addition.

[0149] In the FFN layer, row splitting is used in both the w1 and w3 matrices to calculate the total value of the partial values. Column splitting is used in the w2 matrix calculation to calculate the partial sum of all values ​​to obtain the final FFN result. For each chip, the computational load of 28672 is split into 23 parts, each of size 1280. This means that the amount of data required to be stored in each chip is: 8192*1280 + 8192*1280 + 1280*8192 = 31MB. In other words, each chip needs to store 31MB of sub-matrix data. Because chips are significantly cheaper than GPUs, using multiple chips instead of traditional GPUs for model inference can effectively reduce model inference costs.

[0150] Those skilled in the art will understand that in the above-mentioned method of the specific implementation method, the writing order of each step does not mean a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.

[0151] Based on the same inventive concept, a data processing device corresponding to the data processing method is also provided in the embodiment of the present disclosure. Since the principle of solving the problem by the device in the embodiment of the present disclosure is similar to the above-mentioned data processing method in the embodiment of the present disclosure, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be repeated.

[0152] like Figure 4 FIG. 1 is a schematic diagram of a data processing device provided by an embodiment of the present disclosure, comprising:

[0153] Processing module 401, for processing input data using a processing unit in a computing unit that matches any network layer of a target neural network to obtain output data;

[0154] a sending module 402, configured to send the output data to at least one other computing unit using the communication unit in the computing unit and the communication units in other computing units, according to the communication connections and communication configuration files set for the communication units in different computing units; wherein the other computing units include other computing units matching the network layer or computing units matching the next network layer;

[0155] The loop module 403 is used to return to the step of using the processing unit in the computing unit to process the input data and obtain output data, until the output data of the last computing unit matching the last network layer of the target neural network is obtained, and the output data is used as the data processing result.

[0156] In one possible implementation, the network layer includes multiple functional layers;

[0157] The multiple functional layers at least include a normalization processing functional layer located in the attention layer, a matrix-vector multiplication functional layer related to the query matrix, the key matrix and the value matrix, a multi-head functional layer, a matrix-vector multiplication functional layer related to the result output, an accumulation processing functional layer, and a normalization processing functional layer located in the feedforward neural network layer, a matrix-vector multiplication functional layer related to multiple weight matrices, and an accumulation processing functional layer;

[0158] A computing unit matching any network layer of the target neural network is associated with a functional layer among the multiple functional layers.

[0159] In a possible implementation, the processing module 401, when processing the input data using the processing unit in the computing unit to obtain output data, is configured to:

[0160] In a case where the computing unit is associated with any matrix-vector multiplication functional layer, using a processing unit in the computing unit, performing matrix-vector multiplication processing on the input data and target data stored in a storage unit in the computing unit to obtain output data;

[0161] Among them, the target data includes any preset matrix corresponding to the functional layer matching the computing unit or a submatrix obtained by matrix splitting any preset matrix, and the preset matrix includes a query matrix, a key matrix, a value matrix, a result output matrix and a weight matrix.

[0162] In a possible implementation, the processing module 401, when processing the input data using the processing unit in the computing unit to obtain output data, is configured to:

[0163] When the computing unit is independent of the matrix-vector multiplication functional layer, the processing unit in the computing unit is used to perform corresponding functional processing on the input data according to the network function of the functional layer related to the computing unit to obtain output data.

[0164] In a possible implementation, the device further includes:

[0165] The acquisition module 404 is used to:

[0166] In a case where the input data includes output data sent from multiple other computing units, the output data of each other computing unit is received and buffered respectively using multiple data receiving interfaces in the communication unit included in the computing unit;

[0167] In response to each of the data receiving interfaces completing buffering of the corresponding output data, the calculation unit is used to trigger a reception completion signal to determine that the acquisition of the input data is completed.

[0168] In a possible implementation, the apparatus further includes a detection module 405 configured to:

[0169] Performing functionality testing and performance testing on the computing unit;

[0170] In the case of a failure in the test, the computing unit is replaced by a spare computing unit.

[0171] In a possible implementation, the weight matrix includes a feature mapping matrix, a feature transformation matrix, and a feature compression matrix;

[0172] In a case where the preset matrix is ​​any one of the query matrix, the key matrix, the value matrix, the feature mapping matrix, and the feature transformation matrix, the submatrix included in the target data is a matrix obtained by performing row splitting on the preset matrix;

[0173] In a case where the preset matrix is ​​any one of a result output matrix and a feature compression matrix, the submatrix included in the target data is a matrix obtained by performing column splitting on the preset matrix.

[0174] In a possible implementation, if the preset matrix is ​​split into multiple sub-matrices, the multiple sub-matrices are respectively stored in storage units of different computing units;

[0175] The sizes of the sub-matrices corresponding to a preset matrix are related to at least one of the size of the preset matrix, the storage space of the storage unit in the calculation unit, and the computing capability of the processing unit in the calculation unit.

[0176] In a possible implementation, the processing module 401, when performing corresponding functional processing on the input data using the processing unit in the computing unit according to the network function of the functional layer related to the computing unit to obtain output data, is configured to:

[0177] In a case where the computing unit is a computing unit that matches a multi-head functional layer in the attention layer, using a communication unit in the computing unit to obtain a matrix cache associated with a key matrix and a value matrix from a double rate DDR storage unit corresponding to the computing unit;

[0178] The processing unit is used to perform rotation position encoding processing on the input data and the matrix buffer according to the network function of the multi-head function layer to obtain output data.

[0179] In one possible implementation, when the computing unit is a computing unit that matches an accumulation processing functional layer located in an attention layer, the input data includes the first output data sent by the last computing unit that matches the previous network layer, and the second output data sent by the computing unit that matches the matrix-vector multiplication functional layer related to the result output located in the attention layer;

[0180] The processing module 401, when performing corresponding functional processing on the input data using the processing unit in the computing unit according to the network function of the functional layer related to the computing unit to obtain output data, is used to:

[0181] Utilizing the processing unit in the computing unit, the first output data and the second output data are cumulatively processed according to the network function of the cumulative processing function layer in the attention layer to obtain output data.

[0182] In a possible implementation, when the computing unit is a computing unit matching an accumulation processing functional layer located in a feedforward neural network layer, the input data includes third output data sent by a computing unit matching a matrix-vector multiplication functional layer associated with a feature compression matrix, and fourth output data sent by a computing unit matching an accumulation processing functional layer located in an attention layer;

[0183] The processing module 401, when performing corresponding functional processing on the input data using the processing unit in the computing unit according to the network function of the functional layer related to the computing unit to obtain output data, is used to:

[0184] Utilizing the processing unit in the computing unit, the third output data and the fourth output data are accumulated according to the network function of the accumulation processing function layer in the feedforward neural network layer to obtain output data.

[0185] In a possible implementation, when the computing unit is a computing unit matched with a matrix-vector multiplication function layer associated with a feature compression matrix, the input data includes fifth output data sent by a computing unit matched with a matrix-vector multiplication function layer associated with a feature mapping matrix, and sixth output data sent by a computing unit matched with a matrix-vector multiplication function layer associated with a feature change matrix;

[0186] The processing module 401, when performing matrix-vector multiplication on the input data and the target data stored in the storage unit of the computing unit using the processing unit of the computing unit to obtain output data, is configured to:

[0187] The processing unit in the computing unit is used to perform matrix-vector multiplication processing on the fifth output data, the sixth output data, and the target data stored in the storage unit in the computing unit to obtain output data.

[0188] In one possible implementation, the sending module 402, when sending the output data to at least one other computing unit using the communication unit in the computing unit and the communication units in other computing units according to the communication connections and communication configuration files set for the communication units in different computing units, is configured to:

[0189] determining a target functional layer having a communication relationship with the functional layer matched with the computing unit;

[0190] determining each other computing unit associated with the target functional layer;

[0191] According to the communication connections and communication configuration files set for the communication units in different computing units, the output data is sent to each other computing unit respectively using the data sending interface included in the communication unit in the computing unit and the data receiving interface included in the communication units in each other computing unit.

[0192] For descriptions of the processing flow of each module in the device and the interaction flow between each module, reference can be made to the relevant descriptions in the above method embodiment, which will not be described in detail here.

[0193] Based on the same technical concept, the embodiment of the present application also provides a computer device. Figure 5 FIG. 1 is a schematic diagram of a computer device according to an embodiment of the present invention, comprising:

[0194] Processor 501, memory 502, and bus 503. The memory 502 stores machine-readable instructions executable by the processor 501, and the processor 501 is configured to execute the machine-readable instructions stored in the memory 502. When the machine-readable instructions are executed by the processor 501, the processor 501 performs the following steps: S101: for a computing unit that matches any network layer of a target neural network, using a processing unit in the computing unit, performs data processing on input data to obtain output data; S102: according to communication lines and communication configuration files set for communication units in different computing units, using a communication unit in the computing unit and communication units in other computing units, sends the output data to at least one other computing unit; wherein the other computing units include other computing units that match the network layer, or computing units that match the next network layer; and S103: returning to the step of using the processing unit in the computing unit to perform data processing on the input data to obtain output data, until the output data of the last computing unit that matches the last network layer of the target neural network is obtained, and the output data is used as the data processing result.

[0195] The above-mentioned memory 502 includes a memory 5021 and an external memory 5022; the memory 5021 here is also called an internal memory, which is used to temporarily store the calculation data in the processor 501, as well as the data exchanged with the external memory 5022 such as a hard disk. The processor 501 exchanges data with the external memory 5022 through the memory 5021. When the computer device is running, the processor 501 and the memory 502 communicate through the bus 503, so that the processor 501 executes the execution instructions mentioned in the above method embodiment.

[0196] The present disclosure also provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the computer program executes the steps of the data processing method described in the above method embodiment. The storage medium can be a volatile or non-volatile computer-readable storage medium.

[0197] The embodiments of the present disclosure also provide a computer program product, which carries program code. The instructions included in the program code can be used to execute the steps of the software update method described in the above method embodiment. For details, please refer to the above method embodiment and will not be repeated here.

[0198] The computer program product may be implemented in hardware, software, or a combination thereof. In one embodiment, the computer program product is implemented as a computer storage medium. In another embodiment, the computer program product is implemented as a software product, such as a software development kit (SDK).

[0199] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems and devices described above can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here. In the several embodiments provided in the present disclosure, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0200] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0201] In addition, each functional unit in each embodiment of the present disclosure may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0202] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium that is executable by a processor. Based on this understanding, the technical solution of the present disclosure, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present disclosure. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0203] If the technical solution of this application involves personal information, the product that applies the technical solution of this application has clearly informed the personal information processing rules and obtained the individual's voluntary consent before processing personal information. If the technical solution of this application involves sensitive personal information, the product that applies the technical solution of this application has obtained the individual's separate consent before processing sensitive personal information, and at the same time meets the "explicit consent" requirement. For example, on personal information collection devices such as cameras, a clear and prominent sign is set to inform that the personal information collection scope has been entered and personal information will be collected. If the individual voluntarily enters the collection scope, it is deemed that they agree to the collection of their personal information; or on the personal information processing device, when the personal information processing rules are notified by obvious signs / information, the individual's authorization is obtained through pop-up information or by asking the individual to upload their personal information; among which, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the type of personal information processed.

[0204] Finally, it should be noted that the above-described embodiments are only specific implementation methods of the present disclosure, which are used to illustrate the technical solutions of the present disclosure, rather than to limit them. The scope of protection of the present disclosure is not limited thereto. Although the present disclosure has been described in detail with reference to the above-mentioned embodiments, those skilled in the art should understand that any person skilled in the art can modify or easily conceive of changes to the technical solutions described in the above-mentioned embodiments within the technical scope disclosed in the present disclosure, or replace some of the technical features therein with equivalents. Such modifications, changes, or replacements do not deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and should be included in the scope of protection of the present disclosure. Therefore, the scope of protection of the present disclosure should be based on the scope of protection of the claims.

Claims

1. A data processing method, characterized in that: include: For a computing unit that matches any network layer of the target neural network, using a processing unit in the computing unit, processing the input data to obtain output data; According to the communication lines and communication configuration files set for the communication units in different computing units, the output data is sent to at least one other computing unit using the communication unit in the computing unit and the communication units in other computing units; wherein the other computing units include other computing units matching the network layer or computing units matching the next network layer; Return to the step of using the processing unit in the computing unit to process the input data to obtain output data, until the output data of the last computing unit matching the last network layer of the target neural network is obtained, and the output data is used as the data processing result.

2. The method according to claim 1, characterized in that The network layer includes multiple functional layers; The multiple functional layers at least include a normalization processing functional layer located in the attention layer, a matrix-vector multiplication functional layer related to the query matrix, the key matrix and the value matrix, a multi-head functional layer, a matrix-vector multiplication functional layer related to the result output, an accumulation processing functional layer, and a normalization processing functional layer located in the feedforward neural network layer, a matrix-vector multiplication functional layer related to multiple weight matrices, and an accumulation processing functional layer; A computing unit that matches any network layer of the target neural network is associated with a functional layer among the multiple functional layers.

3. The method according to claim 1, characterized in that The processing unit in the computing unit processes the input data to obtain output data, including: In a case where the computing unit is associated with any matrix-vector multiplication functional layer, using a processing unit in the computing unit, performing matrix-vector multiplication processing on the input data and target data stored in a storage unit in the computing unit to obtain output data; Among them, the target data includes any preset matrix corresponding to the functional layer matching the computing unit or a submatrix obtained by matrix splitting any preset matrix, and the preset matrix includes a query matrix, a key matrix, a value matrix, a result output matrix and a weight matrix.

4. The method according to claim 1, wherein The processing unit in the computing unit processes the input data to obtain output data, further comprising: When the computing unit is independent of the matrix-vector multiplication functional layer, the processing unit in the computing unit is used to perform corresponding functional processing on the input data according to the network function of the functional layer related to the computing unit to obtain output data.

5. The method according to claim 1, wherein Before using the processing unit in the computing unit to process the input data to obtain output data, the method further includes: In a case where the input data includes output data sent from multiple other computing units, the output data of each other computing unit is received and buffered respectively using multiple data receiving interfaces in the communication unit included in the computing unit; In response to each of the data receiving interfaces completing buffering of the corresponding output data, the calculation unit is used to trigger a reception completion signal to determine that the acquisition of the input data is completed.

6. The method according to claim 1, characterized in that The method further comprises: Performing functionality testing and performance testing on the computing unit; In the case of a failure in the test, the computing unit is replaced by a spare computing unit.

7. The method according to claim 3, characterized in that The weight matrix includes a feature mapping matrix, a feature transformation matrix and a feature compression matrix; In a case where the preset matrix is ​​any one of the query matrix, the key matrix, the value matrix, the feature mapping matrix, and the feature transformation matrix, the submatrix included in the target data is a matrix obtained by performing row splitting on the preset matrix; In a case where the preset matrix is ​​any one of a result output matrix and a feature compression matrix, the submatrix included in the target data is a matrix obtained by performing column splitting on the preset matrix.

8. The method according to claim 7, characterized in that If the preset matrix is ​​split into multiple sub-matrices, the multiple sub-matrices are respectively stored in storage units of different computing units; The sizes of the sub-matrices corresponding to a preset matrix are related to at least one of the size of the preset matrix, the storage space of the storage unit in the calculation unit, and the computing capability of the processing unit in the calculation unit.

9. The method according to claim 4, characterized in that The utilizing the processing unit in the computing unit to perform corresponding functional processing on the input data according to the network function of the functional layer related to the computing unit to obtain output data includes: In a case where the computing unit is a computing unit that matches a multi-head functional layer in the attention layer, using a communication unit in the computing unit to obtain a matrix cache associated with a key matrix and a value matrix from a double rate DDR storage unit corresponding to the computing unit; The processing unit is used to perform rotation position encoding processing on the input data and the matrix buffer according to the network function of the multi-head functional layer to obtain output data.

10. The method according to claim 4, characterized in that In the case where the computing unit is a computing unit that matches the accumulation processing functional layer located in the attention layer, the input data includes the first output data sent by the last computing unit that matches the previous network layer, and the second output data sent by the computing unit that matches the matrix-vector multiplication functional layer related to the result output located in the attention layer; The utilizing the processing unit in the computing unit to perform corresponding functional processing on the input data according to the network function of the functional layer related to the computing unit to obtain output data includes: Utilizing the processing unit in the computing unit, the first output data and the second output data are cumulatively processed according to the network function of the cumulative processing function layer in the attention layer to obtain output data.

11. The method according to claim 4, characterized in that In a case where the computing unit is a computing unit that matches an accumulation processing functional layer located in a feedforward neural network layer, the input data includes third output data sent by a computing unit that matches a matrix-vector multiplication functional layer associated with a feature compression matrix, and fourth output data sent by a computing unit that matches an accumulation processing functional layer located in an attention layer; The utilizing the processing unit in the computing unit to perform corresponding functional processing on the input data according to the network function of the functional layer related to the computing unit to obtain output data includes: Utilizing the processing unit in the computing unit, the third output data and the fourth output data are accumulated according to the network function of the accumulation processing function layer in the feedforward neural network layer to obtain output data.

12. The method according to claim 3, characterized in that In a case where the computing unit is a computing unit matched with a matrix-vector multiplication functional layer associated with a feature compression matrix, the input data includes fifth output data sent by a computing unit matched with a matrix-vector multiplication functional layer associated with a feature mapping matrix, and sixth output data sent by a computing unit matched with a matrix-vector multiplication functional layer associated with a feature change matrix; The method of using the processing unit in the computing unit to perform matrix-vector multiplication on the input data and the target data stored in the storage unit in the computing unit to obtain output data includes: The processing unit in the computing unit is used to perform matrix-vector multiplication processing on the fifth output data, the sixth output data, and the target data stored in the storage unit in the computing unit to obtain output data.

13. The method according to claim 1, wherein The step of sending the output data to at least one other computing unit by using the communication unit in the computing unit and the communication units in other computing units according to the communication lines and communication configuration files set for the communication units in different computing units includes: determining a target functional layer having a communication relationship with the functional layer matched with the computing unit; determining other computing units associated with the target functional layer; According to the communication connections and communication configuration files set for the communication units in different computing units, the output data is sent to each other computing unit respectively using the data sending interface included in the communication unit in the computing unit and the data receiving interface included in the communication units in each other computing unit.

14. A data processing device, characterized in that: include: a processing module, configured to process input data using a processing unit in a computing unit that matches any network layer of a target neural network to obtain output data; a sending module, configured to send the output data to at least one other computing unit using the communication unit in the computing unit and the communication units in other computing units, in accordance with the communication connections and communication configuration files set for the communication units in different computing units; wherein the other computing units include other computing units matching the network layer or computing units matching the next network layer; The loop module is used to return to the step of using the processing unit in the computing unit to process the input data and obtain output data, until the output data of the last computing unit matching the last network layer of the target neural network is obtained, and the output data is used as the data processing result.

15. A computer device, characterized in that: include: A processor and a memory, wherein the memory stores machine-readable instructions executable by the processor, and the processor is used to execute the machine-readable instructions stored in the memory. When the machine-readable instructions are executed by the processor, the processor performs the steps of the data processing method according to any one of claims 1 to 13.

16. A computer program product comprising a computer program, characterized in that When the computer program is executed by a computer device, the computer device performs the steps of the data processing method according to any one of claims 1 to 13.