Model inference method, and terminal device and program product

WO2026200226A1PCT designated stage Publication Date: 2026-10-01HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2026/072886
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-24
Filing Date
2026-01-15
Publication Date
2026-10-01

Smart Images

  • Figure CN2026072886_01102026_PF_FP_ABST
    Figure CN2026072886_01102026_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed are a model inference method, and a terminal device and a program product, which relate to the technical field of artificial intelligence (AI). When receiving an inference request, the terminal device controls an acceleration chip of the terminal device to sequentially execute first operators of a plurality of data processing layers of a model, and controls a general-purpose processing chip of the terminal device to sequentially execute second operators of the plurality of data processing layers of the model, so as to obtain an inference result of the inference request. Data is transferred between different chips within the same device, such that communication overheads can be reduced. In addition, by using model operators as the granularity and scheduling a plurality of operators onto different chips within the same terminal device, computing resources within the terminal device can be fully utilized, and the workload of each chip can be reduced, thereby reducing the execution time required for each chip, and further improving the inference speed.
Need to check novelty before this filing date? Find Prior Art

Description

The model's inference methods, terminal devices, and application products

[0001] This application claims priority to Chinese Patent Application No. 202510353242.7, filed on March 24, 2025, entitled "Inference Method for a Model, Terminal Equipment and Program Product", the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to the field of artificial intelligence (AI) technology, and more particularly to a reasoning method for a model, a terminal device, and a program product. Background Technology

[0003] Edge-based large-model inference refers to the inference process of executing large-scale machine learning models on terminal devices. Compared to cloud-based large-model inference, edge-based large-model inference can reduce latency and protect user data security. Furthermore, it can perform inference even when the terminal device is offline.

[0004] Large-scale model inference on edge devices can often be improved by using distributed parallelism across multiple devices to enhance inference speed and reliability. However, in this distributed parallelism approach, the input sequence or intermediate data during inference is transferred between multiple devices, increasing communication overhead and thus reducing inference time. Summary of the Invention

[0005] This application provides a model inference method, terminal device, and program product to reduce the time consumption of large-scale model inference on the device side and enable the application of large-scale models on terminal devices with limited computing resources.

[0006] Firstly, this application provides a model inference method applied to a terminal device equipped with a general-purpose processing chip and an acceleration chip. After receiving an inference request, the terminal device invokes the model to perform inference and obtains the inference result requested. The model includes multiple sequentially executed data processing layers, each including a first operator and a second operator executed sequentially. During model inference, the terminal device controls the acceleration chip to sequentially execute the first operators of the multiple data processing layers, and controls the general-purpose processing chip to sequentially execute the second operators of the multiple data processing layers. The input of the second operator of the first data processing layer is associated with the output of the first operator of the first data processing layer.

[0007] Compared to distributed parallel model inference methods across multiple devices, this application applies to a single terminal device. During model inference, data is transferred between chips within the same device, eliminating data transfer between different devices. This reduces communication overhead, thereby reducing inference time and increasing inference speed. Furthermore, this application schedules multiple operators of the same data processing layer to different chips within the same terminal device at the operator level, fully utilizing the terminal device's computing resources. Additionally, the general-purpose processing chip sequentially executes the first operators of multiple data processing layers, while the acceleration chip sequentially executes the second operators of multiple data processing layers. Each chip independently processes a portion of the operators in each data processing layer, reducing the load on a single chip during model inference and shortening the execution time of a single data processing layer. This, in turn, reduces the total time required for the terminal device to perform model inference, thus improving inference speed.

[0008] In one alternative implementation, the input to the second operator of the first data processing layer includes the output of the first operator of the first data processing layer.

[0009] In this way, the first and second operators of the same data processing layer process the input of the data processing layer in a serial manner, and the first and second operators are executed by different chips, reducing the load on a single chip to improve the processing speed of a single chip, thereby improving the processing speed of the data processing layer, and thus improving the inference speed of the entire model.

[0010] In one alternative implementation, the input of the second operator of the first data processing layer includes the output of the first operator of the first data processing layer and the input of the first operator of the first data processing layer.

[0011] In this way, the first and second operators of the same data processing layer process the input of the data processing layer serially, and there is a residual connection between the first and second operators. The input of the second operator retains the basic information of the input of the first operator, which enhances the data processing performance of the data processing layer, thereby improving the inference performance of the model and increasing the reliability of the inference results. Moreover, the first and second operators are executed by different chips, reducing the load on a single chip and increasing the processing speed of a single chip, thereby improving the processing speed of the data processing layer and thus improving the inference speed of the entire model.

[0012] In one alternative implementation, the terminal device creates a partitioning strategy. This partitioning strategy is used to indicate the execution chip corresponding to each of the multiple operators in the data processing layer when the model processes the inference request. The execution chip includes a general-purpose processing chip or an acceleration chip, wherein the multiple operators include a first operator and a second operator.

[0013] Optionally, the terminal device creates a partitioning strategy periodically. Alternatively, the terminal device creates a partitioning strategy after receiving an inference request.

[0014] Based on this optional implementation method, the terminal device creates a partitioning strategy, determines the execution chip corresponding to each operator in the model inference stage based on the partitioning strategy, and can quickly schedule multiple operators to the corresponding execution chips, thereby improving the model inference speed.

[0015] In one optional implementation, the terminal device determines the execution chip corresponding to each of the multiple operators in the data processing layer based on the computing power parameters and resource utilization of the general-purpose processing chip. Alternatively, the terminal device determines the execution chip corresponding to each of the multiple operators in the data processing layer based on the computing power parameters and resource utilization of the acceleration chip. Or, the terminal device determines the execution chip corresponding to each of the multiple operators in the data processing layer based on both the computing power parameters and resource utilization of the general-purpose processing chip and the acceleration chip.

[0016] Among them, the computing power parameter is used to indicate the computing performance of the execution chip, and the resource utilization rate is used to indicate the ratio between the actual usage time of the execution chip's resources within a preset time period and the total duration of the preset time period. The resources include computing resources, storage resources, or network resources.

[0017] Optionally, the computing power parameters include one or more of the following: processing precision, network bandwidth, or floating-point operations.

[0018] Based on this optional implementation method, the terminal device determines the partitioning strategy for multiple operators in the model by using the computing power parameters and resource utilization of the chips (general-purpose processing chips or accelerator chips). This distributes the computational load required to execute the model across multiple chips, achieving load balancing. Furthermore, by using one or more computing power parameters such as the processing precision, network bandwidth, and floating-point operations of the chips (general-purpose processing chips or accelerator chips), the executable operators of each chip (general-purpose processing chip or accelerator chip) can be determined, fully utilizing the computing power of the chips (general-purpose processing chips or accelerator chips) in the terminal device.

[0019] In one alternative implementation, the terminal device also includes shared memory that supports access by both general-purpose processing chips and acceleration chips.

[0020] Optionally, shared memory is used to store the results of the first operator processing.

[0021] Optionally, shared memory is used to store the results of the second operator processing.

[0022] Optionally, shared memory is used to store the results of the first operator processing and the results of the second operator processing.

[0023] In this way, terminal devices can achieve data pass-through between different chips by sharing memory, reducing communication overhead and thus shortening the model inference time.

[0024] In one alternative implementation, the data processing layer is a transformer layer, which includes a first normalization operator, a first linear operator, a self-attention operator, a second linear operator, and a second normalization operator.

[0025] The first operator includes at least one of the following operators: a first normalization operator, a first linear operator, or a self-attention operator. The second operator includes either a second linear operator or a second normalization operator.

[0026] Based on this optional implementation, the self-attention operators that require a large amount of computation in the transformer layer are scheduled to the acceleration chip, reducing the load on the general-purpose processing chip, thereby shortening the execution time required by a single processing layer, thus reducing the total time required for terminal devices to perform inference through the model, and thus improving the inference speed.

[0027] In one alternative implementation, the acceleration chip includes a graphics processing unit (GPU) and a neural network processing unit (NPU); the execution chip corresponding to the first normalization operator or the first linear operator is the NPU. The execution chip corresponding to the self-attention operator is the GPU.

[0028] Based on this optional implementation, the self-attention operator is scheduled separately on the GPU, reducing the load on a single acceleration chip during the model inference process, thereby shortening the execution time required by a single processing layer, thus reducing the total time required for terminal devices to perform inference through the model, and thus improving the inference speed.

[0029] In one optional implementation, the first operator of the accelerated chip sequentially executing multiple data processing layers includes: for the first data processing layer among the multiple data processing layers, the first operator of the first data processing layer sequentially processes the first data and the second data to obtain the first result corresponding to the first data and the second result corresponding to the second data, respectively.

[0030] The first and second data are data related to the reasoning request.

[0031] Optionally, the data related to the inference request includes data corresponding to the input sequence carried in the inference request. The data corresponding to the input sequence can be obtained by encoding the input sequence, or by extracting features from the input sequence. The first data and the second data are obtained by segmenting the data corresponding to the input sequence.

[0032] Optionally, the data related to the inference request includes data processed by the preceding data processing layer of the first data processing layer on the data corresponding to the input sequence. The preceding data processing layer includes at least one data processing layer.

[0033] Based on this optional implementation, for any data processing layer, the terminal device uses each data as an input to the data processing layer, so that the data processing layer can perform inference on each data separately. This can reduce the load on a single chip during the model inference process, thereby shortening the execution time required by a single processing layer, thus reducing the total time required for the terminal device to perform inference through the model, and thus improving the inference speed.

[0034] In one optional implementation, the second operator of multiple data processing layers is executed sequentially by the general-purpose processing chip, including: the second operator of the first data processing layer sequentially processes the first result and the second result to obtain the third result corresponding to the first result and the fourth result corresponding to the second result, respectively.

[0035] Based on this optional implementation, for any processing layer, during the processing of each data processing layer, the terminal device controls the acceleration chip and the general-purpose processing chip to sequentially execute the first operator and the second operator, using each result (either the first or the second result) as input to the data processing layer, so that the data processing layer can perform inference on each result (either the first or the second result) separately. This reduces the load on a single chip during model inference, thereby shortening the execution time required for a single processing layer, and thus reducing the total time required for the terminal device to perform inference through the model, thereby improving inference speed.

[0036] In one alternative implementation, the first data and the second data have the same length, and the lengths of the first data and the second data are related to the computing power of the general-purpose processing chip or the computing power of the accelerator chip.

[0037] In this way, using multiple data sets of the same length ensures that each chip processes the same amount of data per cycle, reducing fluctuations in chip processing time and improving the stability of the model inference process. Furthermore, by determining the data length based on the computing power of the general-purpose processing chip or the accelerator chip, the compatibility between the data length and the actual frequency of the chip is ensured, improving the reasonableness of the data length.

[0038] In one optional implementation, the first operator controlling the acceleration chip to sequentially execute multiple data processing layers includes: the terminal device controlling the acceleration chip to sequentially execute the first operator of the first data processing layer and the first operator of the second data processing layer. The first and second data processing layers are adjacent data processing layers among the multiple data processing layers.

[0039] In this way, the terminal device controls the acceleration chip to execute multiple data processing layers sequentially. During the processing of each data processing layer, the terminal device controls the acceleration chip to execute the first operator of each data processing layer in the order of the multiple data processing layers. This reduces the load on a single chip during the model inference process, thereby shortening the execution time required for a single processing layer, thus reducing the total time required for the terminal device to perform inference through the model, and thus improving the inference speed.

[0040] In one optional implementation, controlling the general-purpose processing chip to sequentially execute the second operator of multiple data processing layers includes: the terminal device controlling the general-purpose processing chip to sequentially execute the second operator of the first data processing layer and the second operator of the second data processing layer. The first and second data processing layers are adjacent data processing layers among the multiple data processing layers.

[0041] In this way, the terminal device controls the general-purpose processing chip to execute multiple data processing layers sequentially. During the processing of each data processing layer, the terminal device controls the general-purpose processing chip to execute the second operator of each data processing layer in the order of the multiple data processing layers. This reduces the load on a single chip during model inference, thereby shortening the execution time required for a single processing layer, thus reducing the total time required for the terminal device to perform inference through the model, and ultimately improving inference speed.

[0042] In one optional implementation, the terminal device controls the acceleration chip to sequentially execute first operators of multiple data processing layers, and controls the general-purpose processing chip to sequentially execute second operators of multiple data processing layers. This includes: for any given first data processing layer, the terminal device schedules the first operators of the first data processing layer to the acceleration chip and the second operators of the first data processing layer to the general-purpose processing chip according to an operator partitioning strategy. The terminal device inputs first input data from the first data processing layer into the acceleration chip, and controls the acceleration chip to execute the first operators to process the first input data, obtaining a first processing result of the first input data. The terminal device inputs the first processing result of the first input data into the general-purpose processing chip, and controls the general-purpose processing chip to execute the second operators of the first data processing layer to process the first processing result of the first input data, obtaining a second processing result of the first input data. Based on the second processing result of the first input data, the terminal device obtains a first output result of the first data processing layer and uses the first output result of the first data processing layer as the second input data of the second data processing layer. The first and second data processing layers are adjacent data processing layers among the multiple data processing layers.

[0043] Thus, for any given processing layer, the terminal device controls the general-purpose processing chip and the acceleration chip to sequentially execute multiple data processing layers. During the processing of each data processing layer, the terminal device controls the general-purpose processing chip and the acceleration chip to sequentially execute the first operator and the second operator. This reduces the load on a single chip during model inference, thereby shortening the execution time required for a single processing layer, and consequently reducing the total time required for the terminal device to perform inference through the model, thereby improving inference speed.

[0044] In one optional implementation, before inputting the first processing result of the first input data into the general-purpose processing chip, the terminal device controls the acceleration chip to write the first processing result of the first input data into shared memory, and controls the general-purpose processing chip to read the first processing result of the first input data from the shared memory. After using the first processing result of the first input data as input to the processing execution chip to obtain the second processing result of the first input data, the terminal device controls the processing chip to write the second processing result of the first input data into the shared memory.

[0045] In this way, data can be passed between different chips through shared memory, reducing communication overhead and thus shortening the model inference time.

[0046] In one optional implementation, the first data processing layer is a transformer layer, which includes a feature extraction operator, a first normalization operator, a first linear operator, a self-attention operator, a second linear operator, and a second normalization operator. The first operator includes the feature extraction operator, the first normalization operator, the first linear operator, and the self-attention operator; the second operator includes the second linear operator and the second normalization operator. The inference request includes an input sequence. During model inference, the terminal device inputs the data corresponding to the input sequence into the acceleration chip, and controls the acceleration chip to sequentially execute the first normalization operator, the first linear operator, and the self-attention operator of the first data processing layer to process the data corresponding to the input sequence and obtain a first processing result. The terminal device inputs the first processing result into a general-purpose processing chip, and controls the general-purpose processing chip to sequentially execute the second linear operator and the second normalization operator of the first data processing layer to process the first processing result and obtain a second processing result. The terminal device inputs the second processing result into the acceleration chip, and controls the acceleration chip to sequentially execute the feature extraction operator of the first data processing layer, and the first normalization operator, first linear operator, and self-attention operator of the second data processing layer to process the second processing result and obtain the third processing result. The terminal device inputs the third processing result into the general processing chip, and controls the general processing chip to sequentially execute the second linear operator and second normalization operator of the second data processing layer to process the third processing result. This process is repeated until the last data processing layer of multiple data processing layers to obtain the inference result of the inference request.

[0047] Based on this optional implementation, the general-purpose processing chip and the accelerator chip independently process different processing stages of the data corresponding to the input sequence in the execution order of accelerator chip → general-purpose processing chip → accelerator chip, reducing the load on a single chip. Furthermore, after completing a single processing, the chip (general-purpose processing chip or accelerator chip) continues to execute the next processing, reducing the idle time of the chip (general-purpose processing chip or accelerator chip), making full use of the computing power of the chip (general-purpose processing chip or accelerator chip), and thus improving the inference efficiency of the model.

[0048] In one optional implementation, the data corresponding to the input sequence includes a first subsequence and a second subsequence. The accelerator chip is controlled to sequentially execute the first normalization operator, the first linear operator, and the self-attention operator of the first data processing layer to process the data corresponding to the input sequence and obtain a first processing result. This includes: the terminal device inputs the first subsequence and the second subsequence into the accelerator chip; and the accelerator chip is controlled to sequentially execute the first normalization operator, the first linear operator, and the self-attention operator of the first data processing layer to process the first subsequence and the second subsequence, thereby obtaining the first processing result of the first subsequence and the first processing result of the second subsequence.

[0049] Based on this optional implementation, the data corresponding to the input sequence is divided into a first subsequence and a second subsequence. Each subsequence is used as the input to the model, allowing the model to perform inference on each subsequence separately. For each chip, this reduces the amount of data processed per chip and shortens the processing time of each chip, thereby shortening the overall inference time of the model and further improving the inference efficiency of the model.

[0050] In one alternative implementation, the first subsequence and the second subsequence have the same length, and the lengths of the first subsequence and the second subsequence are determined based on the computing power of the general-purpose processing chip or the computing power of the accelerator chip.

[0051] In this way, using multiple sub-data of the same length ensures that each chip processes the same amount of data per run, reducing fluctuations in chip processing time and improving the stability of the model inference process. Furthermore, by determining the length of the sub-data based on the computing power of the general-purpose processing chip or the accelerator chip, the compatibility between the sub-data length and the actual frequency of the chip is ensured, improving the rationality of the sub-data length.

[0052] In one optional implementation, the terminal device obtains the initial segmentation length of the data corresponding to the input sequence based on the number of processing layers in the model and the length of the data corresponding to the input sequence. The terminal device acquires the reference frequency and the actual frequency of the accelerator chip. If the actual frequency does not match the reference frequency, the initial segmentation length is adjusted based on the difference between the actual frequency and the reference frequency to obtain the length of the sub-data. If the actual frequency matches the reference frequency, the initial segmentation length is determined as the length of the sub-data.

[0053] Based on this optional implementation method, the length of the sub-data is determined according to the actual frequency of the chip, the number of processing layers in the model, and the length of the data corresponding to the input sequence, so as to ensure the compatibility between the length of the sub-data and the actual frequency of the chip and improve the rationality of the length of the sub-data.

[0054] In one optional implementation, for a first data processing layer among multiple data processing layers, the terminal device uses a first sub-sequence and a second sub-sequence as input to an acceleration chip. The acceleration chip processes the first and second sub-sequences sequentially to obtain a first processing result for the first sub-sequence and a first processing result for the second sub-sequence. A general-purpose processing chip processes the first processing result of the first sub-sequence and the first processing result of the second sub-sequence sequentially to obtain a second processing result for the first sub-sequence and a second processing result for the second sub-sequence sequentially. The acceleration chip processes the second processing result of the first and second sub-sequences sequentially to obtain the output result of the first data processing layer on the first sub-sequence and the output result of the first data processing layer on the second sub-sequence.

[0055] Optionally, after the acceleration chip obtains the first processing result of the first subsequence, the terminal device uses the first processing result of the first subsequence as the input of the general processing chip.

[0056] Optionally, after the general-purpose processing chip obtains the second processing result of the first subsequence, the terminal device uses the second processing result of the first subsequence as the input of the acceleration chip.

[0057] Optionally, after the acceleration chip obtains the first processing result of the second subsequence and the second processing result of the first subsequence, the acceleration chip uses the operator scheduled in the acceleration chip to process the second processing result of the first subsequence to obtain the output result of the first data processing layer on the first subsequence.

[0058] Based on this optional implementation, the general-purpose processing chip and the accelerator chip independently process different processing stages of each sub-sequence in a pipeline direction of accelerator chip → general-purpose processing chip → accelerator chip, reducing the load on a single chip. Furthermore, after completing a single processing step, the chip (general-purpose processing chip or accelerator chip) continues to execute the next processing step, reducing the idle time of the chip (general-purpose processing chip or accelerator chip) and making full use of the computing power of the chip (general-purpose processing chip or accelerator chip), thereby improving the inference efficiency of the model.

[0059] In one optional implementation, for a first data processing layer among multiple data processing layers, the terminal device uses a first sub-sequence as input to an acceleration chip. The acceleration chip executes a first operator of the first data processing layer to process the first sub-sequence and obtain a first processing result for the first sub-sequence. A general-purpose processing chip executes a second operator of the first data processing layer to process the first processing result of the first sub-sequence and obtain a second processing result for the first sub-sequence. The terminal device then uses a second sub-sequence as input to an acceleration chip. The acceleration chip executes a first operator of the first data processing layer to process the second sub-sequence and obtain a first processing result for the second sub-sequence. The general-purpose processing chip executes a second operator of the first data processing layer to process the first processing result of the second sub-sequence and obtain a second processing result for the second sub-sequence.

[0060] Based on this optional implementation, the general-purpose processing chip and the acceleration chip process each subsequence independently in the order of the first subsequence → the second subsequence, reducing the amount of data processed each time. This improves the inference efficiency of the model.

[0061] In one optional implementation, after receiving an inference request, the terminal device sends a first request to the server based on the input sequence included in the inference request. The receiving server, based on the feature sequence sent in the first request, inputs the feature sequence into an accelerator chip, controls the accelerator chip to sequentially execute the first operators of multiple data processing layers, and controls the general-purpose processing chip to sequentially execute the second operators of multiple data processing layers to process the feature sequence and obtain the inference result of the inference request.

[0062] Optionally, the first request is used to instruct the server to process the input sequence, generate a feature sequence corresponding to the input sequence, and send the feature sequence to the terminal device.

[0063] Based on this optional approach, the computationally intensive feature extraction operation is performed on the server side, reducing the computing power consumption of the terminal device and thus reducing the load on the terminal device during the inference process, thereby improving inference speed. Furthermore, the terminal device schedules multiple operators to different execution chips within the same computing device at the granularity of data processing layer operators. This fully utilizes the computing resources within the terminal device, reducing the load on individual chips and thus reducing the execution time required by a single chip. Consequently, it reduces the total time required for the terminal device to perform inference through the model, thereby improving inference speed.

[0064] Secondly, this application provides a model inference apparatus, which includes an acquisition module and a scheduling module. Wherein:

[0065] The acquisition module is used to receive inference requests. The inference requests are used to request the inference device to call the model for inference. The model includes multiple data processing layers that are executed sequentially. Each data processing layer includes a first operator and a second operator that are executed sequentially.

[0066] The scheduling module is used to control the acceleration chip to execute the first operator of multiple data processing layers sequentially, and to control the general processing chip to execute the second operator of multiple data processing layers sequentially. The input of the second operator of the first data processing layer in the multiple data processing layers is related to the output of the first operator of the first data processing layer.

[0067] Thirdly, this application provides a terminal device, including: a general-purpose processing chip and an acceleration chip.

[0068] The general-purpose processing chip is used to: receive inference requests, and in response to inference requests, schedule the first operator of the model to the acceleration chip and schedule the second operator of the model to the general-purpose processing chip; the model includes multiple data processing layers executed sequentially, and the data processing layers include the first operator and the second operator executed sequentially.

[0069] The acceleration chip is used to: sequentially execute the first operator of multiple data processing layers.

[0070] The general-purpose processing chip is used to sequentially execute second operators of multiple data processing layers. The input of the second operator of the first data processing layer is associated with the output of the first operator of the first data processing layer.

[0071] In one alternative implementation, the input to the second operator of the first data processing layer is the output of the first operator of the first data processing layer.

[0072] In this way, the first and second operators of the same data processing layer process the input of the data processing layer in a serial manner, and the first and second operators are executed by different chips, reducing the load on a single chip to improve the processing speed of a single chip, thereby improving the processing speed of the data processing layer, and thus improving the inference speed of the entire model.

[0073] In one alternative implementation, the input to the second operator of the first data processing layer is the output of the first operator of the first data processing layer and the input of the first operator of the first data processing layer.

[0074] In this way, the first and second operators of the same data processing layer process the input of the data processing layer serially, and there is a residual connection between the first and second operators. The input of the second operator retains the basic information of the input of the first operator, which enhances the data processing performance of the data processing layer, thereby improving the inference performance of the model and increasing the reliability of the inference results. Moreover, the first and second operators are executed by different chips, reducing the load on a single chip and increasing the processing speed of a single chip, thereby improving the processing speed of the data processing layer and thus improving the inference speed of the entire model.

[0075] In one alternative implementation, the general-purpose processing chip is further used to: create a partitioning strategy. This partitioning strategy is used to indicate the execution chip corresponding to each of the multiple operators in the data processing layer when the model processes the inference request. The execution chip includes a general-purpose processing chip or an acceleration chip, wherein the multiple operators include a first operator and a second operator.

[0076] Optionally, the general-purpose processing chip is also used to: periodically create partitioning strategies. Alternatively, the general-purpose processing chip is also used to: create partitioning strategies after receiving inference requests.

[0077] In one alternative implementation, the general-purpose processing chip is further specifically used to: determine the execution chip corresponding to each of the multiple operators in the data processing layer based on the computing power parameters and resource utilization of the general-purpose processing chip. Alternatively, the general-purpose processing chip is further specifically used to: determine the execution chip corresponding to each of the multiple operators in the data processing layer based on the computing power parameters and resource utilization of the acceleration chip. Or, the general-purpose processing chip is further specifically used to: determine the execution chip corresponding to each of the multiple operators in the data processing layer based on both the computing power parameters and resource utilization of the general-purpose processing chip and the computing power parameters and resource utilization of the acceleration chip.

[0078] Among them, the computing power parameter is used to indicate the computing performance of the execution chip, and the resource utilization rate is used to indicate the ratio between the actual usage time of the execution chip's resources within a preset time period and the total duration of the preset time period. The resources include computing resources, storage resources, or network resources.

[0079] Optionally, the computing power parameters include one or more of the following: processing precision, network bandwidth, or floating-point operations.

[0080] In one alternative implementation, the terminal device also includes shared memory, which is used to support access to both general-purpose processing chips and acceleration chips.

[0081] Optionally, shared memory is used to store the results of the first operator processing.

[0082] Optionally, shared memory is used to store the results of the second operator processing.

[0083] Optionally, the shared memory is used to store the results of the first operator processing and the results of the second operator processing.

[0084] In one alternative implementation, the data processing layer is a transformer layer, which includes a first normalization operator, a first linear operator, a self-attention operator, a second linear operator, and a second normalization operator.

[0085] The first operator includes at least one of the following operators: a first normalization operator, a first linear operator, or a self-attention operator. The second operator includes either a second linear operator or a second normalization operator.

[0086] In one alternative implementation, the acceleration chips include GPUs and NPUs; the execution chip corresponding to the first normalization operator or the first linear operator is an NPU. The execution chip corresponding to the self-attention operator is a GPU.

[0087] In one optional implementation, the control acceleration chip is specifically used to: for a first data processing layer among multiple data processing layers, run the first operator of the first data processing layer to sequentially process the first data and the second data, and obtain the first result corresponding to the first data and the second result corresponding to the second data, respectively.

[0088] The first and second data are data related to the reasoning request.

[0089] Optionally, the data related to the inference request includes the input sequence carried by the inference request. The first data and the second data are obtained by segmenting the input sequence.

[0090] Alternatively, the data related to the inference request may include data corresponding to the input sequence carried in the inference request. The data corresponding to the input sequence may be obtained by encoding the input sequence, or by extracting features from the input sequence. The first data and the second data are obtained by segmenting the data corresponding to the input sequence.

[0091] In one alternative implementation, the general-purpose processing chip is further specifically used to: run the second operator of the first data processing layer to sequentially process the first result and the second result, and obtain the third result corresponding to the first result and the fourth result corresponding to the second result, respectively.

[0092] In one alternative implementation, the first data and the second data have the same length, and the lengths of the first data and the second data are related to the computing power of the general-purpose processing chip or the computing power of the accelerator chip.

[0093] In one alternative implementation, the accelerator chip is further specifically used to: sequentially execute a first operator of a first data processing layer and a first operator of a second data processing layer. The first and second data processing layers are adjacent data processing layers among a plurality of data processing layers.

[0094] In one alternative implementation, the control of the general-purpose processing chip is further specifically used to: sequentially execute the second operator of the first data processing layer and the second operator of the second data processing layer. Here, the first data processing layer and the second data processing layer are adjacent data processing layers among a plurality of data processing layers.

[0095] In one optional implementation, the general-purpose processing chip is further specifically used to: schedule the first operator of the first data processing layer to the acceleration chip according to the operator partitioning strategy, schedule the second operator of the first data processing layer to the general-purpose processing chip, and input the first input data of the first data processing layer into the acceleration chip. The acceleration chip is further specifically used to: execute the first operator of the first data processing layer to process the first input data, obtain a first processing result of the first input data, and input the first processing result of the first input data into the general-purpose processing chip. The general-purpose processing chip is further specifically used to: execute the second operator of the first data processing layer to process the first processing result of the first input data, obtain a second processing result of the first input data. The acceleration chip is further specifically used to: obtain a first output result of the first data processing layer based on the second processing result of the first input data, and execute the first operator of the second data processing layer to process the first output result. The first data processing layer and the second data processing layer are adjacent data processing layers among multiple data processing layers.

[0096] In one optional implementation, the first data processing layer is a transformer layer, which includes a feature extraction operator, a first normalization operator, a first linear operator, a self-attention operator, a second linear operator, and a second normalization operator. The first operator includes the feature extraction operator, the first normalization operator, the first linear operator, and the self-attention operator; the second operator includes the second linear operator and the second normalization operator. The inference request includes an input sequence. During model inference, the general-purpose processing chip is also used to input the data corresponding to the input sequence into the acceleration chip. Specifically, the acceleration chip is used to sequentially execute the first normalization operator, the first linear operator, and the self-attention operator of the first data processing layer to process the data corresponding to the input sequence, obtain a first processing result, and input the first processing result into the general-purpose processing chip. Specifically, the general-purpose processing chip is used to sequentially execute the second linear operator and the second normalization operator of the first data processing layer to process the first processing result, obtain a second processing result, and input the second processing result into the acceleration chip. The acceleration chip is also specifically used to: sequentially execute the feature extraction operator of the first data processing layer, and the first normalization operator, the first linear operator, and the self-attention operator of the second data processing layer, process the second processing result to obtain the third processing result, and input the third processing result into the general-purpose processing chip. The general-purpose processing chip is also specifically used to: sequentially execute the second linear operator and the second normalization operator of the second data processing layer, and process the third processing result.

[0097] In one optional implementation, the data corresponding to the input sequence includes a first subsequence and a second subsequence. The general-purpose processing chip is specifically used to input the first and second subsequences into the acceleration chip. The acceleration chip is specifically used to sequentially execute the first normalization operator, the first linear operator, and the self-attention operator of the first data processing layer to process the first and second subsequences, thereby obtaining the first processing result of the first subsequence and the first processing result of the second subsequence.

[0098] In one alternative implementation, the first subsequence and the second subsequence have the same length, and the lengths of the first subsequence and the second subsequence are determined based on the computing power of the general-purpose processing chip or the computing power of the accelerator chip.

[0099] In one optional implementation, the general-purpose processing chip is also used to: obtain the initial segmentation length of the data corresponding to the input sequence based on the number of processing layers in the model and the length of the data corresponding to the input sequence; and obtain the reference frequency and the actual frequency of the accelerator chip. If the actual frequency does not match the reference frequency, the initial segmentation length is adjusted based on the difference between the actual frequency and the reference frequency to obtain the length of the subsequence. If the actual frequency matches the reference frequency, the initial segmentation length is determined as the length of the subsequence.

[0100] In one optional implementation, for the first data processing layer among multiple data processing layers, the general-purpose processing chip is further configured to: use the first sub-sequence and the second sub-sequence as inputs to the acceleration chip. Specifically, the acceleration chip is configured to: process the first sub-sequence and the second sub-sequence sequentially, obtaining a first processing result for the first sub-sequence and a first processing result for the second sub-sequence in turn. Specifically, the general-purpose processing chip is configured to: process the first processing result for the first sub-sequence and the first processing result for the second sub-sequence sequentially, obtaining a second processing result for the first sub-sequence and a second processing result for the second sub-sequence in turn. Specifically, the acceleration chip is further configured to: process the second processing result for the first sub-sequence and the second processing result for the second sub-sequence sequentially, obtaining the output result of the first data processing layer on the first sub-sequence and the output result of the first data processing layer on the second sub-sequence in turn.

[0101] Optionally, the acceleration chip is also used to: after obtaining the first processing result of the first subsequence, use the first processing result of the first subsequence as the input of the general processing chip.

[0102] Optionally, the general-purpose processing chip is also used to: after obtaining the second processing result of the first subsequence, use the second processing result of the first subsequence as the input of the acceleration chip.

[0103] Optionally, the acceleration chip is further configured to: after obtaining the first processing result of the second subsequence and the second processing result of the first subsequence, process the second processing result of the first subsequence to obtain the output result of the first data processing layer on the first subsequence.

[0104] In one optional implementation, for the first data processing layer among multiple data processing layers, the general-purpose processing chip is further configured to: use the first sub-sequence as input to the acceleration chip. Specifically, the acceleration chip is configured to: execute a first operator of the first data processing layer to process the first sub-sequence and obtain a first processing result for the first sub-sequence. Specifically, the general-purpose processing chip is configured to: execute a second operator of the first data processing layer to process the first processing result of the first sub-sequence and obtain a second processing result for the first sub-sequence. The general-purpose processing chip is further configured to: use the second sub-sequence as input to the acceleration chip. Specifically, the acceleration chip is configured to: execute a first operator of the first data processing layer to process the second sub-sequence and obtain a first processing result for the second sub-sequence. Specifically, the general-purpose processing chip is configured to: execute a second operator of the first data processing layer to process the first processing result of the second sub-sequence and obtain a second processing result for the second sub-sequence.

[0105] In one alternative implementation, the general-purpose processing chip is further configured to: receive an inference request; send a first request to the server based on the input sequence included in the inference request; receive a feature sequence sent by the server based on the first request; and input the feature sequence into the acceleration chip.

[0106] Optionally, the first request is used to instruct the server to process the input sequence, generate a feature sequence corresponding to the input sequence, and send the feature sequence to the terminal device.

[0107] Fourthly, this application provides a computer-readable storage medium, including: computer software instructions. When the computer software instructions are invoked by a terminal device, the terminal device implements the method described in the first aspect or any of the optional implementations of the first aspect.

[0108] Fifthly, this application provides a computer program product, which, when run on a terminal device, executes the method described in the first aspect or any of the optional implementations of the first aspect.

[0109] The technical effects brought about by the second to fifth aspects can be found in the first aspect or any of the optional implementations of the first aspect, and will not be repeated here. Based on the implementations provided in the above aspects, this application can further combine them to provide more implementations. Attached Figure Description

[0110] Figure 1 is a schematic diagram of the network structure of a large-scale end-side model;

[0111] Figure 2 is a schematic diagram of distributed inference among multiple devices;

[0112] Figure 3 is a schematic diagram of the data processing system provided in this application;

[0113] Figure 4 is a schematic diagram of the computing device provided in this application;

[0114] Figure 5 is a schematic diagram of a distributed reasoning method provided in this application;

[0115] Figure 6A is a schematic diagram of the inference process of a transformer layer 1 provided in this application;

[0116] Figure 6B is a schematic diagram of the pipelined parallel processing provided in this application;

[0117] Figure 7A is a schematic diagram of an operator scheduling provided in this application;

[0118] Figure 7B is a schematic diagram of an operator scheduling provided in this application;

[0119] Figure 7C is a schematic diagram of an operator scheduling provided in this application;

[0120] Figure 7D is a schematic diagram of an operator scheduling provided in this application;

[0121] Figure 7E is a schematic diagram of an operator scheduling provided in this application;

[0122] Figure 8 is a flowchart illustrating the reasoning method of a model provided in this application.

[0123] Figure 9A is a schematic diagram of the sequential execution of multiple data processing layers provided in this application;

[0124] Figure 9B is a schematic diagram of the data input model corresponding to the input sequence provided in this application;

[0125] Figure 10 is a flowchart illustrating the model inference stage provided in this application;

[0126] Figure 11A is a schematic diagram of data transfer between data processing layers provided in this application;

[0127] Figure 11B is a schematic diagram of data transfer between data processing layers provided in this application;

[0128] Figure 12 is a schematic diagram of the model reasoning stage provided in this application (II).

[0129] Figure 13 is a flowchart illustrating the reasoning method of a model provided in this application (II).

[0130] Figure 14 is a schematic diagram of the structure of a processing chip provided in this application. Detailed Implementation

[0131] Large language models (LLMs) are deep learning models trained on large amounts of text data that can generate natural language text or understand the meaning of language text. LLMs can handle various natural language tasks, such as text classification, question answering, and dialogue, and also show potential in handling complex tasks such as code generation and vulnerability detection. In this paper, the large language model may also be referred to as a large model. In one example, an LLM can be trained using a generative deep neural network model based on a transformer architecture.

[0132] Large-scale edge models are large-scale machine learning models, such as LLMs, deployed on edge devices. Edge devices execute these large models to perform tasks such as language translation, health diagnosis, content generation, message logging, and autonomous driving. Commonly used large-scale edge language models include, but are not limited to: transformer-based bidirectional encoder representation (BERT) models, transformer-based decoder representation (GPT) models, transformer-based text-to-text transfer-transformer (T5) models, models combining autoregression and autoencoders (XLNet), or transformer-based deepseek-R1 models. This application does not limit the specific type of large-scale edge language models.

[0133] Taking BERT as an example, BERT can pre-train deep bidirectional representations using unlabeled text by jointly adjusting the left and right contexts across all layers.

[0134] As shown in Figure 1(a), a typical edge-side large model may include an input layer, multiple transformer layers, a normalization layer, a linear layer, and an output layer. In some alternative approaches, the transformer layer may be called a processing layer, and the input layer, normalization layer, linear layer, and output layer may be called network layers.

[0135] The input layer encodes the input sequence to obtain an input vector. For example, it converts words in the input text into word vectors, images into image-encoded vectors, or speech into word vectors. In some optional methods, the input layer can also obtain the positional encoding vector of each word (pixel or image patch) in the input sequence, concatenating the positional encoding vector and the input vector as the input to the large-scale model on the edge. The input vector can also be called an input sequence vector, an input vector sequence, or other names; this application does not limit its usage. Furthermore, the positional encoding vector can also be called a position vector or a positional encoding; this application does not limit its usage either.

[0136] Multiple transformer layers are used to extract features from the output of the input layer and obtain contextual information to obtain the feature representation of the input sequence, and to perform inference calculations based on the feature representation of the input sequence to obtain the inference representation result of the input sequence.

[0137] The inference representation result can refer to the output of any one of the multiple transformer layers, or the output of the last transformer layer. This inference representation result is used to generate the inference result. This inference representation result can include, but is not limited to, text features, image features, or audio features.

[0138] The inference result indicates the output of the model after processing the input sequence. This inference result can include, but is not limited to: text, query results, generated multimodal content, graphs, tables, audio, and video. Multimodal content includes: images generated from text, translated text, videos generated from text, and style-transformed images.

[0139] In some alternative approaches, as shown in Figure 1(a), multiple transformer layers comprise L transformer layers: transformer layer 1, transformer layer 2, ..., and transformer layer L. These L transformer layers are executed sequentially; that is, the output of transformer layer 1 serves as the input to transformer layer 2, the output of transformer layer L-1 serves as the input to transformer layer L, and the output of transformer layer L serves as the inference representation of the input sequence. Here, L is a positive integer greater than 2.

[0140] In some alternative approaches, the residuals between L transformer layers are concatenated (not shown in Figure 1(a)). Specifically, the output of transformer layer 1 is used as the input of transformer layer 2, and the output of transformer layer 1 is used as the input of transformer layer j. The input of transformer layer j consists of the outputs of transformer layer j-1 and transformer layer 1. This increases the number of transformer layers. Here, j is a positive integer greater than 2 and less than or equal to L.

[0141] The normalization layer normalizes the outputs of the L transformer layers (i.e., the inference representations of the input sequences). The output of the normalization layer is then input into the linear layer for feature processing. The output of the linear layer is input into the output layer, which calculates the inference results of the large model.

[0142] The input sequence is an ordered collection of elements, where each element has a unique position. Elements in the input sequence can be of any type, such as numbers, strings, objects, arrays, tuples, linked lists (or lists), tokens, stacks, and queues. Strings are sequences of characters commonly used in text processing, cryptography, and image processing. Arrays are sequences of elements of the same type, often used for storing large amounts of data and performing numerical calculations. Linked lists are sequences of nodes, often used to implement dynamic data structures and efficient insertion and deletion operations. Tuples are immutable sequences, often used to group multiple values ​​into a single unit. Lists are mutable sequences, often used for storing and manipulating data. Stacks and queues are special types of sequences commonly used to implement data structures and algorithms.

[0143] Multiple transformer layers can have different structures, or they can have the same structure. Alternatively, at least two transformer layers can have different structures.

[0144] Taking multiple transformer layers with the same structure as an example,

[0145] Each transformer layer typically includes a self-attention sublayer and a linear sublayer. For example, as shown in Figure 1(b), each transformer layer includes: a first normalization sublayer, a first linear sublayer, a self-attention sublayer, a second normalization sublayer, a first accumulation sublayer, a second linear sublayer, an activation function sublayer, and a second accumulation sublayer. The inputs of each transformer layer are fed into the first normalization sublayer and the first accumulation sublayer, respectively. The output of the first normalization sublayer is fed into the first linear sublayer. The output of the first linear sublayer is fed into the self-attention sublayer. The output of the self-attention sublayer is fed into the second linear sublayer. The output of the second linear sublayer is fed into the first accumulation sublayer. The output of the first accumulation sublayer is fed into the second normalization sublayer. The outputs of the second normalization sublayer are fed into both the activation function sublayer and the second accumulation sublayer. The output of the activation function sublayer is fed into the second accumulation sublayer. The second accumulation sublayer accumulates the outputs of the activation function sublayer and the second normalization sublayer; the output of the second accumulation sublayer is the output of the transformer layer.

[0146] The structure of the edge-side large model shown in Figure 1 above is merely illustrative. In practical applications, the edge-side large model can also have other structures. For example, the edge-side large model can also include an encoder and a decoder. The encoder is used to extract features from the input sequence, and the decoder is used to perform inference based on the extracted features to obtain the inference result. As another example, as shown in Figure 5 below, the transformer layer in the edge-side large model can also have other structures, which are not limited in this application.

[0147] Large edge models require a lot of memory and computing resources when running in actual operation, which makes it very difficult to deploy them to mobile devices.

[0148] To improve the running efficiency of large models on mobile devices, a distributed parallel approach across multiple devices can be adopted to enhance inference speed and reliability.

[0149] The distributed parallel approach across multiple devices involves deploying different transformer layers in the large edge model on different devices and dividing the input sequence into multiple subsequences. Different subsequences are processed by transformer layers deployed on different devices. As shown in Figure 2, the large edge model deploys transformer layers 1 through 5 on devices 1 through 5, respectively. Transformer layers 1 through 5 run serially, meaning the output of transformer layer i is input to transformer layer i+1, where i is a positive integer greater than 0 and less than 5. The input sequence is divided into five subsequences: subsequence 1 through subsequence 5, with each device (device 1 through device 5) processing one subsequence. As shown in Figure 2, starting from transformer layer 2, the input of each transformer layer (transformer layer 2, transformer layer 3, transformer layer 4, or transformer layer 5) is the output of the previous transformer layer and the corresponding subsequence (subsequence 2, subsequence 3, subsequence 4, or subsequence 5). By having five devices collaboratively process multiple subsequences, inference speed is improved. As shown in Figure 2, transformer layer 1 and processing subsequence 1 are deployed on device 1; transformer layer 2 and processing subsequence 2 are deployed on device 2; transformer layer 3 and processing subsequence 3 are deployed on device 3; transformer layer 4 and processing subsequence 4 are deployed on device 4; and transformer layer 5 and processing subsequence 5 are deployed on device 5.

[0150] As shown in Figure 2, since different transformer layers in the large edge model run on different devices, the data transfer between transformer layers increases communication overhead, leading to increased inference time. Furthermore, the distributed parallel approach across multiple devices typically requires a single large edge model to handle inference requests from different devices, increasing the amount of data the large edge model needs to infer and consuming significant memory resources.

[0151] Based on this, to reduce the time consumption of large-scale model inference on the edge and realize the application of large-scale models on terminal devices with limited computing resources, this application provides a model inference method for terminal devices. This terminal device is equipped with a general-purpose processing chip and an acceleration chip. During model inference, data is transferred between chips within the same device, without involving data transfer between different devices, thereby reducing communication overhead, reducing inference time, and improving inference speed. Furthermore, this application schedules multiple operators of the same data processing layer to different chips within the same terminal device at the operator level, fully utilizing the computing resources within the terminal device. Furthermore, the general-purpose processing chip sequentially executes the first operators of multiple data processing layers, and the acceleration chip sequentially executes the second operators of multiple data processing layers. Each chip independently processes a portion of the operators in each data processing layer, reducing the load on a single chip during model inference, thereby shortening the execution time required for a single data processing layer, and further reducing the total time required for the terminal device to perform inference through the model, thus improving inference speed. In addition, this application classifies multiple operators of each data processing layer into multiple types of operators at the granularity of the model's operators, and schedules different types of operators to the corresponding execution chips, so that each type of operator is executed by a type of execution chip within the same terminal device, which can fully utilize the computing resources within the terminal device, thereby improving inference speed.

[0152] Specifically, the terminal device includes a general-purpose processing chip and an acceleration chip. Upon receiving an inference request, the terminal device invokes the model to perform inference and obtains the inference result. This model comprises multiple sequentially executed data processing layers, each including a first operator and a second operator executed sequentially. During model inference, the terminal device controls the acceleration chip to sequentially execute the first operators of the multiple data processing layers and controls the general-purpose processing chip to sequentially execute the second operators of the multiple data processing layers to obtain the inference result.

[0153] The first operator may also be referred to as a first type of operator, a first kind of operator, or other names, and this application does not limit this. Similarly, the second operator may also be referred to as a second type of operator, a second kind of operator, or other names, and this application does not limit this. In some optional embodiments, the multiple operators in the data processing layer may include one or more first operators, and correspondingly, the multiple operators in the data processing layer may include multiple or one second operator. Furthermore, the sequentially executed first and second operators may be executed in the order of first operator → second operator. Alternatively, the sequentially executed first and second operators may be executed in the order of second operator → first operator, and this application does not limit this. Moreover, in the case of multiple first operators and multiple second operators, the sequentially executed first and second operators may be executed in the order of second operator → first operator → second operator. Alternatively, the sequentially executed first and second operators may be executed in the order of first operator → second operator → first operator. This application does not limit this.

[0154] Furthermore, sequential execution of multiple data processing layers can mean that multiple data processing layers run serially, meaning that the input of the current data processing layer is related to the output of the previous processing layer, and the output of the current data processing layer is input to the next data processing layer. In some alternative approaches, the data processing layer may also be called a processing layer, a feature processing layer, or other names, which this application does not limit.

[0155] The technical solution provided in this application can be applied to terminal devices equipped with one or more acceleration chips. In the case of multiple acceleration chips, the execution chip corresponding to the first operator among the multiple operators and the execution chip corresponding to the second operator among the multiple operators can be different acceleration chips.

[0156] In one alternative approach, the model deployed in the terminal device can be a large model proposed by current AI technology, or it can be a large model proposed by future AI technology. Accordingly, the technical solution provided in this application is not only applicable to current model inference and distributed inference scenarios, but also to future model inference, training, distributed inference, and distributed training technologies.

[0157] The terminology used in the implementation section of this application is only for explaining specific embodiments of this application and is not intended to limit this application. A brief introduction to some concepts that may be involved in this application is given below.

[0158] Operators: The operators included in a model refer to computational units within the model, such as the normalization sub-layer, linear sub-layer, and self-attention sub-layer in the transformer layer mentioned above. Different computational units have different processing logic; therefore, computational units with different processing logic are also called different types of operators. Typically, a model includes multiple different types of operators, and the support of computing devices varies for these different types of operators. For example, this difference in support indicates the difference in the computational precision of the operators run on the computing device. If the computing device has poor support for an operator, the computational precision of that operator will be low; if the computing device has good support for an operator, the computational precision of that operator will be high.

[0159] Pipeline parallelism refers to dividing a processing procedure into multiple independent stages, each executed independently, with the result of the previous stage serving as the input for the current stage. In this application, pipeline parallelism divides the model's inference process into multiple independent stages at the operator level. Each independent stage includes at least one operator, and different independent stages are executed by different execution chips. Furthermore, each execution chip is responsible for executing the operator assigned to it and passing the result to the next execution chip, forming a pipeline.

[0160] The terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined with "first," "second," etc., may explicitly or implicitly include one or more of that feature. In the description of this application, unless otherwise stated, "a plurality of" means two or more.

[0161] The following description, in conjunction with the accompanying drawings, provides an exemplary introduction to some application scenarios and system architectures that may be involved in this application.

[0162] Figure 3 is a schematic diagram of the structure of a data processing system provided in this application. As shown in Figure 3, the data processing system 100 includes a computing device 110, a training device 120, a database 130, a data storage system 150, and a data acquisition device 160.

[0163] The training device 120 can be a terminal, or other computing devices that support integer calculations, such as servers or cloud devices. Alternatively, the training device 120 can be a cluster, such as a server cluster or a cloud server cluster.

[0164] The computing device 110 can be a terminal, such as a computer, mobile terminal, tablet computer, laptop computer, virtual reality (VR) device, augmented reality (AR) device, mixed reality (MR) device, extended reality (ER) device, camera, or in-vehicle computer, etc., or it can be an edge device (e.g., a box carrying a chip with processing capabilities). In this application, the computing device 110 can be a mobile phone or an in-vehicle terminal, etc.

[0165] In an alternative embodiment, the computing device 110 may include a general-purpose processing chip 110a and one or more acceleration chips. As shown in FIG3, it includes a general-purpose processing chip 110a, a first acceleration chip 111, and a second acceleration chip 112.

[0166] In the computing device 110 shown in Figure 3, the general-purpose processing chip 110a is used to deploy different operators in the model on different execution chips (general-purpose processing chip 110a, first acceleration chip 111, or second acceleration chip 112). These execution chips process different input data through locally deployed operators, thereby achieving computational power collaboration among multiple execution chips and improving the resource utilization of the computing device 110. The general-purpose processing chip 110a can be a central processing unit (CPU) used to receive inference requests, schedule different operators in the model to the corresponding execution chips, and send the data (such as feature vectors) corresponding to the input sequence of the inference request to the corresponding execution chips. In terms of hardware implementation, the general-purpose processing chip 110a may include one or more processors, which is not limited in this application.

[0167] The acceleration chips included in the embodiments of this application include one or more dies, such as the first acceleration chip 111 including die 1 and die 2, and the second acceleration chip 112 including die 3 and die 4. The first acceleration chip 111 may be, but is not limited to, a GPU chip, an NPU chip, a tensor processing unit (TPU) chip, a microprocessor chip, an application-specific integrated circuit (ASIC) chip, or one or more integrated circuit chips used to control the execution of the computer program product provided in this application. Each acceleration chip includes a die such as a GPU within a GPU chip, an NPU within an NPU chip, or a TPU within a TPU chip. For example, if the die is a GPU within a GPU chip, then the die can be used to perform mathematical and geometric calculations to achieve tasks such as image rendering.

[0168] In some alternative approaches, the first accelerator chip 111 and the second accelerator chip 112 can be the same type of accelerator chip, for example, the first accelerator chip 111 and the second accelerator chip 112 can both be GPUs, or the first accelerator chip 111 and the second accelerator chip 112 can both be NPUs.

[0169] In other alternatives, the first accelerator chip 111 and the second accelerator chip 112 can be different types of accelerator chips. For example, the first accelerator chip 111 can be a GPU and the second accelerator chip 112 can be an NPU. Alternatively, the first accelerator chip 111 can be an NPU and the second accelerator chip 112 can be a GPU.

[0170] In one alternative implementation, the different chips (processor chip 110a, first accelerator chip 111, or second accelerator chip 112) are networked in a full mesh mode, meaning that all chips are directly connected to each other. In some cases, the communication links between different chips can also be called inter-chip communication links, i.e., chip-to-chip, such as connecting different chips using one or more of the following methods: High-Speed ​​Custom Communication System (HCCS) interface, High-Speed ​​GPU Interconnect Bandwidth interface, Integrated Circuit interface, Controller Area Network bus, Serial Peripheral Device interface, Queued Serial Peripheral Device interface, Full-Duplex Asynchronous Serial Interface, or Half-Duplex Differential Serial Interface, etc.

[0171] The HCCS interface is a high-speed connection channel between dies, used to accelerate data and computation to produce executable results. For example, in a chip, different dies are connected pairwise using HCCS technology.

[0172] High-speed GPU interconnect bandwidth interface is a high-speed interconnect technology between GPUs, which is usually implemented by multiple pairs of wires printed on the computer board, with each pair of wires connecting to different GPUs at their respective ends.

[0173] The Inter-Integrated Circuit (I2C) bus is a source-synchronous serial bus used for short-distance communication between different integrated circuits. I2C uses two lines for data transmission: a serial data line (SDL) and a serial clock line (SCL). The Controller Area Network (CAN) bus is a serial communication protocol bus used for real-time applications; it can use twisted-pair cables to transmit signals. The Serial Peripheral Interface (SPI) bus is a 3-wire synchronous serial full-duplex communication interface with advantages such as simple circuitry, high speed, and reliable communication. The Queued Serial Peripheral Interface (QSPI) bus adds a queued transmission mechanism to SPI; QSPI uses a dedicated communication interface to connect single, dual, or four data lines. A full-duplex asynchronous serial interface, also known as a universal asynchronous receiver / transmitter (UART) interface, is a universal serial data bus used for asynchronous communication. This UART bus can be a bidirectional communication bus, converting the data to be transmitted between serial and parallel communication modes. For example, a UART interface refers to an RS-232 interface. A half-duplex differential serial interface is a serial communication bus interface that uses a two-wire, differential transmission, and half-duplex mode, such as an RS-485 interface.

[0174] It is worth noting that, unlike chip-to-chip communication links, within a single chip, the communication links between individual chips can use, but are not limited to, SIO interfaces for connection.

[0175] In this embodiment, the transmission bandwidth (SIO bandwidth) of the die-to-die communication link is greater than the chip-to-chip transmission bandwidth. Transmission bandwidth refers to the maximum amount of data that different devices can transmit per unit time. Taking the first accelerator chip 111 and the second accelerator chip 112 shown in Figure 3 as examples, the SIO bandwidth between die 1 and die 2 is greater than the transmission bandwidth between the first accelerator chip 111 and the second accelerator chip 112. For example, the SIO bandwidth between die 1 and die 2 is 392 gigabytes per second (GB / s), and the transmission bandwidth between the first accelerator chip 111 and the second accelerator chip 112 is 100 GB / s.

[0176] In a first possible implementation, computing device 110 and training device 120 are deployed on different physical devices (e.g., servers in a server or cluster), or computing device 110 and training device 120 are different physical devices. Exemplarily, computing device 110 and training device 120 are processors deployed on different physical devices. For example, computing device 110 may be a GPU, central processing unit, other general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic device, discrete hardware component, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. Training device 120 may be a graphics processing unit (GPU), NPU, microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits used to control the execution of the program according to the present application.

[0177] In the second alternative implementation, the computing device 110 and the training device 120 are deployed on the same physical device, or the computing device 110 and the training device 120 are on the same physical device.

[0178] The two optional implementation methods mentioned above are only different deployment methods for computing device 110 and training device 120. In other embodiments, computing device 110 and training device 120 may also be deployed in other ways, which are not limited in this application embodiment.

[0179] The data acquisition device 160 is used to collect training data and store it in the database 130. The data acquisition device 160, computing device 110, and training device 120 can be the same or different devices. Taking images as an example, the data acquisition device 160 can be a sensor, camera, mobile phone, tablet computer, or computing device or network device with image acquisition function. Taking text as an example, the data acquisition device 160 can be a voice acquisition device, mobile phone, tablet computer, or computing device or network device with text acquisition or extraction function.

[0180] Training device 120 is used to train the neural network using training data until the loss function in the neural network converges and the value of the loss function is less than a certain threshold, at which point the neural network training is complete. Then, training device 120 configures the trained neural network onto computing device 110, as shown in neural network 101 in Figure 3. Computing device 110 uses the locally deployed neural network 101 to perform inference and computation on the input sequence to obtain the inference result.

[0181] In this embodiment, the training method that the training device 120 can adopt is not limited. For example, the training device 120 can use one or more of the following training methods to train the neural network: centralized training, distributed training, joint learning, online learning, reinforcement learning, self-supervised learning, semi-supervised learning, unsupervised learning, transfer learning, incremental learning, or meta-learning, etc.

[0182] The aforementioned neural network model can be referred to as an AI model or simply a model, which can refer to a large language model, a large model, or other models.

[0183] In one optional example, the neural network refers to the aforementioned large language model. Large language models, leveraging their powerful computing capabilities and sophisticated algorithms, can effectively process massive amounts of data, providing users with efficient and accurate information processing and analysis services. Large language models not only excel in understanding and generating human language but also demonstrate immense potential in solving complex problems and tasks. For example, large language models are widely used in automated question-answering systems, text summarization, machine translation, and language generation, significantly improving efficiency and accuracy. Especially when dealing with large-scale datasets, large language models can uncover deep patterns and associations, supporting decision-making. Furthermore, the self-learning ability of large language models allows them to continuously evolve, constantly improving their performance and intelligence by learning from new data. In some optional cases, large language models may also be simply referred to as large models. The large model provided in this application can refer not only to a large language model, but also to a model with a certain number of parameters. Depending on the domain in which the large model is applied, it can also refer to a model that includes various functions such as image processing, human-computer interaction, semantic search, semantic query, or dialogue. This application does not limit the domain in which the large model can be applied or its specific name. In this document, for the sake of simplicity, it is referred to as a large model, but this should not be construed as a limitation of this application, and will not be elaborated further thereafter.

[0184] As shown in Figure 1 above, the large language model contains many transformer layers, and each transformer layer includes multiple sub-layers. Running the large language model on a single device requires a long inference time. Therefore, to accelerate the inference speed of the large model on a single device, a heterogeneous computing power collaborative approach is adopted to run the large model, so as to fully utilize the computing resources of the terminal device. Specific examples can be found in the embodiments provided in Figures 5 to 13 below, which will not be elaborated upon here.

[0185] In another alternative implementation, neural network refers to other types of networks. For example, neural network 101 is a convolutional neural network (CNN), a recurrent neural network (RNN), or a graph neural network (GNN), etc. Further implementations of CNN, RNN, or GNN can be found in the description of commonly used techniques, which will not be elaborated upon here.

[0186] In some embodiments, if the computing device 110 and the training device 120 are the same physical device, the training device 120 can configure the trained neural network 101 to itself and use the trained neural network 101 to achieve the target function to be achieved by the model, such as intelligent question answering, data retrieval, document translation, online translation, multimodal content generation and other functions on terminal devices, as well as functions such as signage or document verification in commercial areas, schools, parks and sports venues in cities, as well as operations such as target detection, object recognition or classification on data, and functions such as facial recognition payment and object classification (such as commodity classification).

[0187] In other embodiments, the training device 120 can configure the trained neural network 101 onto multiple computing devices, so that each computing device can achieve the target function to be achieved by the model described above. In addition to the target function of the above embodiments, the neural network 101 can also achieve some functions that can be achieved by a large language model, etc.

[0188] It should be noted that in practical applications, the training data maintained in database 130 may not all come from data acquisition device 160; it may also be received from other devices. Furthermore, training device 120 may not necessarily train the neural network entirely based on the training data maintained in database 130; it may also obtain training data from the cloud or other sources. The above description should not be construed as limiting the embodiments of this application.

[0189] Figure 3 is merely a schematic diagram of a system architecture provided in an embodiment of this application. The positional relationships between the devices, components, modules, etc. shown in Figure 3 do not constitute any limitation. Depending on the user's needs for model inference, the training device and computing device may include more or fewer hardware components, and this application does not limit this. For example, in Figure 3, the data storage system 150 is an external memory relative to the computing device 110. In other cases, the data storage system 150 may also be placed in the computing device 110.

[0190] This application can be applied to scenarios including but not limited to: content generation scenarios for large edge models, multi-turn dialogue scenarios, semantic query scenarios for human-computer interaction, model training scenarios, model tuning scenarios, and scenarios involving distributed inference of models.

[0191] The structure of the computing device 110 shown in Figure 3 will be described below. Figure 4 is a schematic diagram of the structure of a computing device 110 provided in this application. Referring to Figure 4, the computing device 110 includes a general-purpose processing chip 110a, a memory 115, an I / O interface 116, a storage device 117, and multiple acceleration chips, such as a first acceleration chip 111, a second acceleration chip 112, a third acceleration chip 113, and a fourth acceleration chip 114.

[0192] The general-purpose processing chip 110a, multiple acceleration chips, memory 115, and I / O interface 116 are connected using a full mesh, meaning that any two components communicate via direct connection. For example, an acceleration chip communicates with any other chip via a direct connection. The interfaces used for direct connections between different components are described in Figure 3 and will not be repeated here.

[0193] In the computing device 110 shown in Figure 4, the memory 115 is provided with shared memory, which can support access to different chips in the computing device 110 (general-purpose processing chip 110a, first acceleration chip 111, second acceleration chip 112, third acceleration chip 113 or third acceleration chip 114), and realize data pass-through between general-purpose processing chip 110a, first acceleration chip 111, second acceleration chip 112, third acceleration chip 113 and third acceleration chip 114.

[0194] The shared memory can be a portion of the storage space in memory 115. Alternatively, the shared memory can refer to all the storage space in memory 115. This application does not limit this aspect.

[0195] I / O interface 116 is used for data interaction with external devices. Users can send data to I / O interface 116 through external devices, such as instructions to instruct computing device 110 to start model inference.

[0196] Memory 117 may include volatile memory, such as random access memory (RAM). Memory 117 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).

[0197] The memory 117 is used to store the program code of the model and a set of computer instructions. During the model inference process, the general-purpose processing chip 110a executes a set of computer instructions and calls the model, scheduling different operators in the model to the general-purpose processing chip 110a, the first acceleration chip 111, the second acceleration chip 112, the third acceleration chip 113 or the third acceleration chip 114.

[0198] The general-purpose processing chip 110a is used to preprocess the data received from the I / O interface 116, and the cooperating acceleration chips (first acceleration chip 111, second acceleration chip 112, third acceleration chip 113, or third acceleration chip 114) perform inference on the preprocessed data. Optionally, the general-purpose processing chip 110a can perform preprocessing operations such as denoising to remove irrelevant information and recover useful real information from the input sequence, such as text sequence, image, or speech data. Optionally, the general-purpose processing chip 110a can encode the input sequence, such as text sequence, image, or speech, to obtain the preprocessing operation corresponding to the input sequence.

[0199] The general-purpose processing chip 110a can also be used to perform inference on data received from the I / O interface 116 in collaboration with acceleration chips (first acceleration chip 111, second acceleration chip 112, third acceleration chip 113 or third acceleration chip 114).

[0200] During the preprocessing of the input sequence by the computing device 110, or during the computation and related processing of the general-purpose processing chip 110a of the computing device 110, the computing device 110 can call data, code, etc. in the data storage system 150 for corresponding processing, or store the data and instructions obtained from the corresponding processing into the data storage system 150.

[0201] In the computing device 110 shown in Figure 4, similar to Figure 3 above, each acceleration chip (first acceleration chip 111, second acceleration chip 112, third acceleration chip 113, or third acceleration chip 114) includes one or more dies. As shown in Figure 4, the first acceleration chip 111 includes dies 1 and 2, the second acceleration chip 112 includes dies 3 and 4, the third acceleration chip 113 includes dies 5 and 6, and the fourth acceleration chip 114 includes dies 7 and 8. For example, different dies within the same acceleration chip are connected via a SIO interface. The transmission bandwidth (SIO bandwidth) of the communication link (or channel) connected via SIO is greater than the transmission bandwidth of the communication link between different chips.

[0202] The structure of the computing device 110 shown in Figure 4 above is only exemplary. In actual applications, the computing device 110 may include more or fewer acceleration chips than that in Figure 4 above, or the acceleration chips in the computing device 110 may include more or fewer chips than those in Figure 4 above, or the computing device 110 may include multiple general-purpose processing chips 110a. This application embodiment does not limit this.

[0203] Based on the computing device 110 shown in Figure 4, and taking the computing device 110 of this application embodiment running a large model of a transformer architecture as an example, a possible example of distributed inference of a model is provided below. Please refer to Figure 5, which is a schematic diagram of distributed inference of a model provided in this application. Taking the data processing layer as a transformer layer as an example, the model includes multiple transformer layers, such as transformer layer 1, transformer layer 2, and transformer layer L. As shown in Figure 1(b) above, each of these multiple transformer layers includes multiple sub-layers, where each sub-layer can be called an operator, that is, each transformer layer includes multiple operators.

[0204] Each transformer layer relies on the computing device 110 shown in Figure 3 or Figure 4 for implementation. The multiple operators corresponding to each transformer layer can be run by multiple execution chips (general-purpose processing chips or accelerator chips) or multiple chips. By utilizing the multiple execution chips in the computing device 110 to collaboratively run each transformer layer, the amount of computation required by each execution chip is reduced, and the computing resources of the computing device 110 are fully utilized, thereby improving the inference efficiency of the model. As shown in Figure 5, each transformer layer in transformer layer 1, transformer layer 2, ... and transformer layer L is run by general-purpose processing chip 110a and chips 1 to 8.

[0205] In an alternative scenario, other network layers in the model (such as the input layer, linear layer, normalization layer, and output layer described above) can be executed by a single chip (e.g., a general-purpose processing chip 110a, a GPU, or an NPU). Alternatively, different network layers can be scheduled onto different chips; for example, the input layer can be scheduled onto the general-purpose processing chip 110a, while the linear layer, normalization layer, and output layer can be scheduled onto the GPU or NPU. This application does not limit the scope of this embodiment.

[0206] In the first alternative approach, multiple transformer layers operate serially. During the model's inference process, the general-purpose processing chip 110a acquires a sequence and inputs the sequence into different transformer layers in the model. The different transformer layers process a portion of the words in the sequence and the output of the previous layer, and output the inference result corresponding to that portion of words (not shown in Figure 5).

[0207] In the second alternative approach, multiple transformer layers operate serially. During model inference, the general-purpose processing chip 110a acquires the sequence and inputs it into the first transformer layer. Multiple transformer layers process the sequence in a pipelined manner to obtain the inference result. As shown in Figure 5, the output data 1 of transformer layer 1 serves as the input of transformer layer 2, and the input of transformer layer L is the output data L-1 of transformer layer L-1.

[0208] In the third alternative approach, multiple transformer layers operate serially. During the model's inference process, the general-purpose processing chip 110a acquires the sequence and inputs it into different transformer layers in the model. The different transformer layers process the sequence and the output of the previous network layer to obtain the inference result of the sequence (not shown in Figure 5).

[0209] In the three optional approaches described above, for any transformer layer, different chips (or kernels) are used to process a portion of the words in the sequence (or the input data or the entire sequence) according to the operators deployed locally, and output the corresponding inference results. Understandably, different operators in the transformer layer are run by different kernels or chips, realizing parallel inference at the operator granularity.

[0210] In some alternative approaches, the sequence may also be referred to as the input sequence or other names, and this application embodiment does not limit this.

[0211] Referring to Figure 5, in some optional embodiments, the sequence can refer to a query statement. This query statement can be semantic information generated by the general-purpose processing chip 110a based on a sequence of data such as text or audio information, which may be input by the user, transmitted from other devices, or extracted by the computing device 110 based on an image input by the user. Referring to Figure 5, the query statement includes multiple tokens, such as tokens 1 to m corresponding to the black pattern. In this document, tokens can include, but are not limited to, words, characters, punctuation marks, and special symbols (such as calculation symbols). In some optional cases, words can also be referred to as tokens; this application does not limit this usage.

[0212] As users' demands for model processing power gradually increase, and the words in query statements continue to grow, the processing efficiency of a single device becomes increasingly inefficient when processing long sequences using a model deployed on a single device. In this example, a long sequence refers to a query statement containing a large number of words. The following is a brief explanation of the model's inference process using transformer layer 1 as an example.

[0213] The general-purpose processing chip 110a randomly allocates or schedules resources for multiple operators in transformer layer 1 to different cores or chips. For example, the general-purpose processing chip 110a schedules the self-attention operator to core 1 and the normalization operator to core 2. The general-purpose processing chip 110a randomly or sequentially inputs the query (Q) vector, key (K) vector, and value (V) vector corresponding to different words (tokens) in the query statement into core 1. Core 1 uses the self-attention operator to perform multiple calculations on different words (tokens) to obtain intermediate results for each word (token). Core 1 writes the intermediate results of each word (token) into shared memory. Core 2 reads the intermediate results of each word (token) from the shared memory and performs normalization processing on the intermediate results of each word (token). In this way, different chips (or cores) collaboratively implement the pipelined execution of transformer layer 1.

[0214] In this paper, for the sake of simplicity, key vectors and value vectors can be simply referred to as key-value vectors.

[0215] Please refer to Figure 5. Different cores represent different operators within each of the multiple transformer layers used to run the model. Operators within the same transformer layer are executed in a specific order, therefore the execution order of operators running on different cores differs. Furthermore, each transformer layer operates at the token level, performing computations on multiple tokens (token 1 to token m). In other words, multiple fully interconnected cores (or chips) are used to pipeline different operators within a transformer layer, and these cores (or chips) process each token (token 1 to token m) in parallel. This allows each core to run only one operator per computation, processing one token or its intermediate result. This fully utilizes the computing resources of the chips in computing device 110 while reducing the computational load per operation. Moreover, multiple cores can achieve data pass-through through shared memory or full interconnect, reducing communication overhead and consequently decreasing model inference time, thus improving the inference efficiency of large edge models.

[0216] In the first optional scenario, "processing each word (token 1 to token m) in parallel in a pipelined manner" can mean that transformer layer 1 processes each word individually, and during the processing of each word, it utilizes multiple operators in transformer layer 1 to process each word in a pipelined manner. In this way, after the first operator in transformer layer 1 completes the processing of one word, it can start processing the next word, thus enabling transformer layer 1 to process multiple words in parallel.

[0217] In the second optional scenario, "processing each word (token 1 to token m) in parallel in a pipelined manner" can mean that transformer layer 1 processes each word individually, obtains the processing result for each word, and passes the processing result of that word to the execution chip or core that executes the operators in transformer layer 2. In this way, after obtaining the output result of one word, transformer layer 1 in the model can start processing the next word, thus enabling the model to process multiple words in parallel.

[0218] The two optional scenarios mentioned above are merely different alternatives for processing each word (token 1 to token m) in parallel in a pipeline manner. In other embodiments, there may be other implementations for processing each word (token 1 to token m) in parallel in a pipeline manner, which are not limited in this application.

[0219] For example, transformer layer 1 includes attention layer 1 and multilayer perceptron (MLP) 1. Attention layer 1 is used to calculate the query vector, key vector, and value vector corresponding to input data 1 (such as token 1 to token m in a query statement), and finally output the result of attention layer 1. MLP 1 includes two linear layers, which are used to linearly process the output of attention layer 1 to obtain the output result of transformer layer 1 (output data 1). The output result of transformer layer 1 is input to transformer layer 2. The contents of other transformer layers shown in Figure 5 can be referred to the description of transformer layer 1. For example, transformer layer 2 includes attention layer 2 and MLP 2, and transformer layer L includes attention layer L and MLPL. Transformer layer 2 processes the output data 1 of transformer layer 1 to obtain output data 2, and transformer layer L processes the output data L-1 of transformer layer L-1 to obtain output data L.

[0220] Regarding the specific network structure for scheduling the transformer layer into multiple cores, the following example, with reference to Figure 5, illustrates the specific network of the transformer layer.

[0221] In the first optional case, transformer layer 1 includes a first normalization operator, a first linear operator, a self-attention operator, a second linear operator, and a second normalization operator.

[0222] In a first optional example, the first operator includes a first normalization operator, and the second operator includes a second linear operator. The general-purpose processing chip 110a schedules the first normalization operator to acceleration chips (such as first acceleration chip 111, second acceleration chip 112, third acceleration chip 113, or third acceleration chip 114), schedules the second linear operator to itself, and schedules the first linear operator, the self-attention operator, and the second normalization operator to the remaining acceleration chips. For example, the first operator is scheduled to the first acceleration chip 111, and the general-purpose processing chip 110a schedules the first linear operator, the self-attention operator, and the second normalization operator to the second acceleration chip 112, the third acceleration chip 113, and the third acceleration chip 114, respectively.

[0223] In a second alternative example, the first operator includes a first normalization operator and a first linear operator, and the second operator includes a second linear operator. The general-purpose processing chip 110a schedules the first normalization operator and the first linear operator to acceleration chips (such as first acceleration chip 111, second acceleration chip 112, third acceleration chip 113, or third acceleration chip 114), and schedules the self-attention operator and the second normalization operator to the remaining acceleration chips. For example, the first operator is scheduled to the first acceleration chip 111, and the general-purpose processing chip 110a schedules the self-attention operator and the second normalization operator to the second acceleration chip 112 and the third acceleration chip 113, respectively.

[0224] In a third alternative example, the first operator includes a first normalization operator, a first linear operator, and a self-attention operator, and the second operator includes a second linear operator. The general-purpose processing chip 110a schedules the first normalization operator, the first linear operator, and the self-attention operator to acceleration chips (such as first acceleration chip 111, second acceleration chip 112, third acceleration chip 113, or third acceleration chip 114), and schedules the second normalization operator to the remaining acceleration chips. For example, the first operator is scheduled to first acceleration chip 111, and the general-purpose processing chip 110a schedules the second normalization operator to second acceleration chip 112, third acceleration chip 113, or third acceleration chip 114.

[0225] In a fourth alternative example, the first operator includes a first normalization operator, a first linear operator, and a self-attention operator, and the second operator includes a second linear operator and a second normalization operator. The general-purpose processing chip 110a schedules the first normalization operator, the first linear operator, and the self-attention operator to acceleration chips (such as the first acceleration chip 111, the second acceleration chip 112, the third acceleration chip 113, or the third acceleration chip 114), and schedules the second linear operator and the second normalization operator to the general-purpose processing chip 110a.

[0226] The four optional examples above are merely different scheduling methods for the first and second operators when the transformer layer 1 includes a first normalization operator, a first linear operator, a self-attention operator, a second linear operator, and a second normalization operator. In other embodiments, the first and second operators may have other scheduling methods. This application does not limit these methods.

[0227] In the second optional case, transformer layer 1 includes a feature extraction operator, a first normalization operator, a first linear operator, a self-attention operator, a second linear operator, and a second normalization operator.

[0228] The first operator includes one or more of the following operators: a feature extraction operator, a first normalization operator, a first linear operator, and a self-attention operator. The second operator includes either a second linear operator or a second normalization operator. The scheduling method of the operators in transformer layer 1 can refer to the first optional scenario, which will not be elaborated upon here. In the third optional scenario, each data processing layer may include a first operator, a second operator, and a third operator executed sequentially. Taking transformer layer 1 as an example, which includes a feature extraction operator, a first normalization operator, a first linear operator, a self-attention operator, a second linear operator, and a second normalization operator, the first operator includes one or more of the first normalization operator, the first linear operator, or a self-attention operator. The second operator includes either a second linear operator or a second normalization operator. The third operator includes the remaining operators in transformer layer 1 other than the first and second operators.

[0229] For example, taking the first operator as including the first normalization operator, the first linear operator, and the self-attention operator, the second operator as including the second linear operator and the second normalization operator, and the third operator as including the t-feature extraction operator, for each data processing layer, the first operator, the second operator, and the third operator are executed in the order of first operator → second operator → third operator.

[0230] In some alternative approaches, the first operator, the second operator, and the third operator can be executed by different chips. For example, the first operator is executed by the first acceleration chip 111, the second operator is executed by the general-purpose processing chip 110a, and the third operator is executed by the second acceleration chip 111.

[0231] In some alternative approaches, at least two of the first, second, and third operators can be executed by different chips. For example, the first and third operators can be executed by a first acceleration chip 111, and the second operator can be executed by a general-purpose processing chip 110a.

[0232] The above three optional scenarios are merely different scheduling methods for the first and second operators. In other embodiments, the first and second operators may have other scheduling methods. This application does not limit these methods.

[0233] For example, taking the second operator as including the second linear operator and the second normalization operator, and the first operator as including the feature extraction operator, the first normalization operator, the first linear operator, and the self-attention operator as an example, please refer to Figure 5. The general processing chip 110a schedules the feature extraction operator, the first normalization operator, and the first linear operator to chip 1, schedules the self-attention operator to chip 3, and schedules the second linear operator and the second normalization operator to the general processing chip 110a.

[0234] The first linear operator performs linear calculations on the normalization result of the first normalization operator. The self-attention operator, based on the calculation result of the first linear operator, calculates the dot product between the Q corresponding to each word (token) and the other K and V associated with that Q, obtaining the relevance score between each word and other words. The second linear operator performs linear calculations on the calculation result of the self-attention operator. The feature extraction operator extracts features from the output of the second normalization operator. The first and second normalization operators prevent network layer degradation in the transformer layer and perform over-normalization on the activation values ​​of different operators. For example, the second normalization operator performs a weighted summation on the linear result of the second linear operator to obtain a high-dimensional representation of each word. Alternatively, the first normalization operator performs a weighted summation on the words (tokens) input to transformer layer 1 to obtain a new representation for each word.

[0235] Figure 5 above is provided for illustrative purposes only and should not be construed as limiting the reasoning method of the model provided in this application. In practical applications, the first operator may include more or fewer operators than those in Figure 5 above, and correspondingly, the second operator may also include more or fewer operators than those in Figure 5 above; this application does not impose any limitations in this regard.

[0236] Based on the scheduling method of each operator in transformer layer 1 shown in Figure 5, the following example illustrates the inference process of transformer layer 1, taking a sequence including tokens 1 to 4 and multiple operators in transformer layer 1 processing each word in a pipeline manner, as shown in Figure 6A. Figure 6A is a schematic diagram of the inference process of transformer layer 1 provided in this application. The inference process of transformer layer 1 includes stages ① to ⑩, where:

[0237] Phase ①: The general processing chip 110a sends token 1 to token 4 to chip 1 in sequence.

[0238] In stage ②, core 1 processes token i sequentially according to the order of token 1 → token 4. During the processing of token i, it is processed sequentially using the first normalization operator and the first linear operator to obtain the first processing result of token i.

[0239] Where i equals 1, 2, 3 or 4.

[0240] In stage 3, core 1 writes the first processing result of token i into shared memory.

[0241] In stage 4, core 3 reads the first processing result of token i from the shared memory in the order of token 1→token 4.

[0242] In stage 5, core 3 processes the first processing result of token i sequentially through the self-attention operator in the order of token 1→token 4 to obtain the second processing result of token i.

[0243] In stage 6, core 3 writes the second processing result of token i into shared memory.

[0244] In stage 7, core 5 reads the second processing result of token i from the shared memory in the order of token 1→token 4.

[0245] In stage ⑧, the general-purpose processing chip 110a processes the second processing result of token i sequentially, from token 1 to token 4. During the processing of the second processing result of token i, it is also processed sequentially by the second linear operator and the second normalization operator to obtain the third processing result of token i. The general-purpose processing chip 110a writes the third processing result of token i into shared memory.

[0246] In stage 9, core 1 reads the third processing result of token i from shared memory in the order of token 1→token 4.

[0247] In stage 10, core 1 extracts features sequentially from token 1 to token 4 using the feature extraction operator based on the third processing result of token i, thus obtaining the first output result of transformer layer 1 for token i.

[0248] In one alternative implementation, after core 1 obtains the first processing result of token 4, it uses a feature extraction operator to extract features from the third processing result of token i according to the order of token 1→token 4.

[0249] In another alternative implementation, after core 1 obtains the third processing result of token 1, it starts feature extraction of the third processing result of token 1.

[0250] In some optional ways, transformer layer 1 takes the first output result of each token as input to transformer layer 2, and the first normalization operator and the first linear operator in transformer layer 2 process the first output result of each token respectively or in sequence. Referring to the processing process of transformer layer 1 provided in Figure 6A above, the second output result of transformer layer 2 for each token is obtained.

[0251] In one optional cleaning process, multiple operators in transformer layer 1 process each word (token 1 to token m) in a pipeline manner. After core 1 initiates feature extraction of the third processing result of token 1, it sequentially processes the first output result of token 1 through the first normalization operator and the first linear operator to obtain the first processing result of the first output result of token 1.

[0252] In another optional scenario, transformer layer 1 processes each word (token 1 to token m) individually to obtain the processing result for each word. After core 1 initiates feature extraction of the third processing result of token 1, and after obtaining the first output result of token i, core 1 processes the first output result of token i sequentially through the first normalization operator and the first linear operator for tokens 1 to 4 to obtain the first processing result of the first output result of token i.

[0253] The following example illustrates the pipelined parallel processing method between transformer layer 1 and transformer layer 2, using the first optional scenario described above. As shown in Figure 6B, which is a schematic diagram of the pipelined parallel processing provided in this application, in the processing stage of transformer layer 1, after processing token 1 and obtaining the first processing result T1-1 of token 1, core 1 begins processing token 2, until obtaining the first processing result T1-1 of token 4.

[0254] After core 1 completes the processing of token 1, core 3 begins processing the first processing result T1-1 of token 1. After obtaining the second processing result T1-2 of token 1, it waits for the first processing result T1-1 of token 2 from core 1. After obtaining the first processing result T1-1 of token 2, it begins processing the first processing result T1-1 of token 2. This continues until the processing of the first processing result T1-1 of token 4 is completed.

[0255] After the general-purpose processing chip 110a completes the processing of the first processing result T1-1 of token 1 in chip 3, it starts processing the second processing result T1-2 of token 1. After obtaining the third processing result T1-3 of token 1, it waits for the second processing result T1-2 of token 2 in chip 3. After obtaining the second processing result T1-2 of token 2, it starts processing the second processing result T1-2 of token 2 until the processing of the second processing result T1-2 of token 4 is completed.

[0256] After processing token 4, core 1 obtains the third processing result T1-3 of token 1, processes the third processing result T1-3 of token 1, and obtains the first processing result T2-1 of the first output result of token 1. After obtaining the first processing result T2-1 of the first output result of token 1, it begins processing the third processing result T1-3 of token 2, and so on, until the third processing result T1-3 of token 4 is completed.

[0257] Correspondingly, after completing the processing of the first processing result T1-4 of token 4, core 3 starts processing the first processing result T2-1 of the first output result of token 1, and obtains the second processing result T2-2 of the first output result of token 1.

[0258] Similarly, after processing the second processing result T1-2 of token 4, the general-purpose processing chip 110a processes the second processing result T2-2 of the first output result of token 1.

[0259] As shown in the embodiment provided in Figure 6B, during the processing of transformer layer 1, since chip 1 processes tokens 1 to 4 sequentially, while chip 1 processes token 2, chip 3 can process token 1. Furthermore, while chip 1 processes token 4, chip 2 processes token 3, and the general-purpose processing chip 110a processes token 2. Thus, when chip 1, chip 3, and the general-purpose processing chip 110a collaboratively run different operators in the transformer layer, each token can be processed in a pipelined manner according to the execution order of the different operators in the transformer layer. This fully utilizes the computing resources of each chip (or chip). Moreover, multiple chips (or chips) can achieve data pass-through through shared memory or full interconnection, reducing communication overhead and thus reducing model inference time and improving the inference efficiency of large models on the edge.

[0260] Furthermore, after completing token 4 processing, chip 1 can begin the processing phase of transformer layer 2. This means that after processing the current word (or processing result), each chip (or core) can quickly begin processing the next word (or processing result). This reduces the idle time of each chip (or core), thus fully utilizing the computing resources of each core and improving the inference efficiency of the large-scale model on the edge.

[0261] As can be seen from the embodiments in Figures 5, 6A, and 6B, since the general-purpose processing chip 110a schedules different operators of a transformer layer to different cores (or chips), the computing power of different cores (or chips) varies, and the resource utilization of different cores (or chips) also varies. During the model inference process, some cores have a long computing time, which causes the waiting time of subsequent cores to be long, thus affecting the inference efficiency of the model.

[0262] Regarding the inference process of different transformer layers in the model, based on the computing device 110 shown in Figures 4 and 5, and taking a large language model as an example, the inference method of the model provided in this application will be introduced below. The inference method of this model includes an operator scheduling stage and an inference stage. The operator scheduling stage is used to determine the different operators or chips in the computing device 110 that execute each transformer layer in the model, forming a partitioning strategy for multiple operators. This partitioning strategy distributes the computational load required to execute the model across multiple chips or chips to achieve load balancing. The inference stage uses the execution order of operators deployed in different chips or chips to control the sequential execution of multiple transformer layers on the corresponding chips or chips, obtaining the inference result of the input sequence.

[0263] In an alternative approach, the partitioning strategy for multiple operators may also be called a partitioning strategy, a scheduling strategy, or other names, which are not limited in this application.

[0264] The operator scheduling phase of the model's inference method will be described first by example, followed by the inference phase. For the implementation of the inference phase, please refer to the embodiments shown in Figures 8 to 13 below.

[0265] The operator scheduling phase can be executed by a chip in the acceleration chip, such as chip 1 in the first acceleration chip 111 mentioned above; or, the operator scheduling phase can also be executed by a general-purpose processing chip, such as the general-purpose processing chip 110a mentioned above.

[0266] The following example illustrates the operator scheduling phase performed by the general-purpose processing chip 110a.

[0267] First, we will illustrate the triggering method of the operator scheduling phase with an example.

[0268] In a first alternative example, the general-purpose processing chip 110a can initiate the operator scheduling phase when the computing device 110 is first started. For example, taking a mobile phone as an example, when the mobile phone is first powered on, the general-purpose processing chip 110a executes the operator scheduling phase and creates a partitioning strategy for multiple operators.

[0269] In a second alternative example, the general-purpose processing chip 110a can initiate the execution operator scheduling phase upon detecting an operating system update in the computing device 110.

[0270] In a third alternative example, the general-purpose processing chip 110a can initiate the execution operator scheduling phase when an update of an AI-type application is detected in the computing device 110.

[0271] In a fourth alternative example, the general-purpose processing chip 110a can initiate the operator scheduling phase when it detects that a model has been deployed to the computing device 110.

[0272] In a fifth optional example, the general-purpose processing chip 110a can initiate the operator scheduling phase upon receiving an inference request. The inference request instructs the computing device 110 to invoke the model for inference. The inference request includes an input sequence. The triggering method for the inference request can be referred to the embodiment shown in Figure 8 below, and will not be elaborated upon here.

[0273] In the sixth alternative example, the general-purpose processing chip 110a can initiate the operator scheduling phase when a model is detected to be invoked.

[0274] The above six optional examples are merely different triggering conditions for the general-purpose processing chip 110a to enable the operator scheduling phase. In other embodiments, the general-purpose processing chip 110a may have other triggering conditions for enabling the operator scheduling phase, which are not limited in this application. For example, the general-purpose processing chip 110a may periodically enable the execution operator scheduling phase. As another example, the general-purpose processing chip 110a may enable the execution operator scheduling phase when the power consumption of the computing device 110 is greater than or equal to a power consumption threshold. The power consumption of the computing device 110 can be determined by one or more of the following parameters: the temperature of the computing device 110, the frequency of the chip (accelerator chip or general-purpose processing chip), and the resource utilization rate of the chip (accelerator chip or general-purpose processing chip).

[0275] Secondly, with reference to Figures 7A to 7E, the following provides an exemplary description of how the general-purpose processing chip 110a schedules different operators to the corresponding cores (or chips).

[0276] In one alternative implementation, the general-purpose processing chip 110a can schedule multiple operators to corresponding cores (or chips) in a manner that distributes the number of operators equally. Here, "equal distribution" can mean that each core or chip running operators executes the same number of operators.

[0277] Taking H as the number of transformer layers in the model and N as the number of chips in the computing device 110 as an example, for any transformer layer including M operators, the general-purpose processing chip 110a determines the number of operators configured in each chip (or chip) in the computing device 110 based on the ratio between M and N. For example, the number of operators in a chip (or chip) is M / N. When M operators cannot be evenly distributed among N chips (or chips), the number of chips (or chips) is reduced, or the number of operators to be executed in a specific chip (or chip) is reduced. This specific chip (or chip) can be the last chip (or chip) or other chips (or chips), and this application does not limit this.

[0278] Taking computing device 11 as an example, which only includes the first acceleration chip 111, as shown in Figure 7A, Figure 7A is a schematic diagram of operator scheduling provided in this application. As shown in Figure 5 above, any transformer layer i includes 6 operators: feature extraction operator, first normalization operator, first linear operator, self-attention operator, second linear operator, and second normalization operator. The acceleration chip in computing device 110 contains 2 (N=2) chip data, so 3 operators need to be scheduled in each chip.

[0279] In some alternative approaches, the six operators in transformer layer i can be scheduled to the first acceleration chip 111 and the general-purpose processing chip 110a according to the order of operator execution. For example, as shown in Figure 7A(a), the feature extraction operator, the first normalization operator, and the first linear operator are scheduled to the first acceleration chip 111, and the self-attention operator, the second linear operator, and the second normalization operator are scheduled to the general-purpose processing chip 110a. In this alternative example, the second operator includes the self-attention operator, the second linear operator, and the second normalization operator, and the first operator includes the feature extraction operator, the first normalization operator, and the first linear operator.

[0280] Alternatively, the first normalization operator, the first linear operator, and the self-attention operator can be scheduled to the general processing chip 110a, and the second linear operator, the second normalization operator, and the feature extraction operator can be scheduled to the first acceleration chip 111 (not shown in Figure (a) of Figure 7A).

[0281] In other alternative approaches, the six operators in transformer layer i can be scheduled sequentially to the first acceleration chip 111 and the general processing chip 110a according to the order of operator execution. For example, as shown in Figure 7A(b), the feature extraction operator is scheduled to the first acceleration chip 111, the first normalization operator is scheduled to the general processing chip 110a, the first linear operator is scheduled to the first acceleration chip 111, the self-attention operator is scheduled to the general processing chip 110a, the second linear operator is scheduled to the first acceleration chip 111, and the second normalization operator is scheduled to the general processing chip 110a. In this alternative example, the second operator includes the first normalization operator, the self-attention operator, and the second normalization operator, and the first operator includes the feature extraction operator, the second linear operator, and the first linear operator.

[0282] The two optional methods described above are simply different implementations of how the general-purpose processing chip 110a schedules multiple operators to the first acceleration chip 111 and the general-purpose processing chip 110a in an evenly distributed manner. In other embodiments, when the computing device 110 includes more acceleration chips, the general-purpose processing chip 110a may also employ other evenly distributed methods to schedule multiple operators.

[0283] For example, computing device 110 includes a first acceleration chip 111, a second acceleration chip 112, a third acceleration chip 113, and a fourth acceleration chip 114. Taking the structure of transformer layer i shown in Figure 5 as an example, general processing chip 110a can schedule the six operators in transformer layer i to general processing chip 110a and cores 1 to 5, with one operator deployed in each core (or chip). As shown in Figure 7A(c), the feature extraction operator is scheduled to core 1, the first normalization operator is scheduled to core 2, the first linear operator is scheduled to core 3, the self-attention operator is scheduled to core 4, the second linear operator is scheduled to core 5, and the second normalization operator is scheduled to general processing chip 110a. In this optional example, the second operator includes the second normalization operator, and the first operator includes the feature extraction operator, the first normalization operator, the self-attention operator, the second linear operator, and the first linear operator.

[0284] Figures (a), (b), and (c) in Figure 7A above illustrate the example of operators in transformer layer i being scheduled to the general-purpose processing chip 110a and the accelerator chip, respectively. In other embodiments, transformer layer i can also be entirely executed by the accelerator chip. The computing device 110 includes a first accelerator chip 111 and a second accelerator chip 112. Taking the structure of transformer layer i shown in Figure 5 above as an example, the general-purpose processing chip 110a can schedule the six operators in transformer layer i to the first accelerator chip 111 and the second accelerator chip 112, respectively. As shown in Figure (d) in Figure 7A, the feature extraction operator, the first normalization operator, and the first linear operator are scheduled to the first accelerator chip 111, and the self-attention operator, the second linear operator, and the second normalization operator are scheduled to the second accelerator chip 112. In this optional example, the second operator includes the feature extraction operator, the first normalization operator, and the first linear operator, and the first operator includes the self-attention operator, the second linear operator, and the second normalization operator.

[0285] Figure 7A above illustrates an implementation of the general-purpose processing chip 110a scheduling multiple operators in an evenly distributed manner. In other embodiments, the general-purpose processing chip 110a can also implement operator scheduling using load balancing. Exemplarily, the general-purpose processing chip 110a schedules multiple operators to corresponding cores or chips according to their computational load, thereby ensuring load balancing across each core or chip running an operator. Load balancing can mean that the resource utilization rate of each core or chip running an operator is the same or similar. Resource utilization rate, also known as resource utilization, computing power utilization, or other names, indicates the ratio between the actual usage time of a chip's resources within a preset time period and the total duration of the preset time period. In the embodiments of this application, resources include one or more of computing resources, storage resources, and network resources. In some optional embodiments, resource utilization rate can reflect the load on a chip or chip; a chip with a higher resource utilization rate has a larger load, and correspondingly, a chip with a lower resource utilization rate has a lower load.

[0286] The following provides an example of how operator scheduling can be implemented under load balancing.

[0287] In the first optional implementation, for any transformer layer, if the transformer layer includes M operators, the general processing chip 110a schedules the M operators to the corresponding cores (or chips) according to the amount of computation required for each of the M operators to perform and the computing power of each core (or chip) among the N cores (or chips), so as to balance the load of the cores (or chips) with scheduled operators.

[0288] The computational cost required to execute each operator can be the network bandwidth or floating-point operations required to execute the operator. Alternatively, the computational cost required to execute each operator can be both the network bandwidth and floating-point operations required to execute the operator.

[0289] The computing power of each chip (or kernel) includes at least one of the chip's frequency and bandwidth.

[0290] In the first alternative approach, the computational cost required for each operator execution is the network bandwidth. When the computational power of each core (or chip) includes frequency, a core (or chip) with a higher frequency can execute operators with greater network bandwidth, or a core (or chip) with a higher frequency can execute a greater number of operators. When the computational power of each core (or chip) includes bandwidth, a core (or chip) with higher bandwidth can execute operators with greater network bandwidth, or a core (or chip) with higher bandwidth can execute a greater number of operators.

[0291] In the second alternative approach, the computational cost required for each operator execution is in the form of floating-point operations. When the computational power of each chip (or kernel) includes frequency, a chip (or kernel) with a higher frequency can execute more floating-point operators, or a chip (or kernel) with a higher frequency can execute a greater number of operators. When the computational power of each chip (or kernel) includes bandwidth, a chip (or kernel) with higher bandwidth can execute more floating-point operators, or a chip (or kernel) with higher bandwidth can execute a greater number of operators.

[0292] In the third alternative approach, the computational load required for each operator execution includes network bandwidth and floating-point operations. When the computational power of each core (or chip) includes frequency, a core (or chip) with a higher frequency can execute operators with greater network bandwidth and more floating-point operations, or a core (or chip) with a higher frequency can execute a greater number of operators. When the computational power of each core (or chip) includes bandwidth, a core (or chip) with higher bandwidth can execute operators with greater network bandwidth and more floating-point operations, or a core (or chip) with higher bandwidth can execute a greater number of operators.

[0293] The above three optional methods are only optional methods for load balancing under different forms of computational load. In other embodiments, when the computational load includes other content, other methods can also be used to achieve load balancing. This application embodiment does not limit this.

[0294] Taking the computing device 11 as an example, which only includes the first acceleration chip 111, when the computing power of the first acceleration chip 111 and the general-purpose processing chip 110a is the same, the general-purpose processing chip 110a can group multiple operators according to the amount of computation to obtain two operator combinations. The amount of computation required by each operator combination is the same or similar, and the two operator combinations are scheduled to the general-purpose processing chip 110a and the first acceleration chip 111 respectively.

[0295] Taking the structure of transformer layer i shown in Figure 5 as an example, since the self-attention operator requires a large amount of computation, and the general-purpose processing chip 110a also needs to perform an operator scheduling stage, as shown in Figure 7B(a), the first linear operator, the self-attention operator, and the second linear operator are scheduled to the first acceleration chip 111, and the feature extraction operator, the first normalization operator, and the second normalization operator are scheduled to the general-purpose processing chip 110a. In this optional example, the second operator includes: the feature extraction operator, the first normalization operator, and the second normalization operator, and the first operator includes the first linear operator, the self-attention operator, and the second linear operator.

[0296] When the computing power of the first acceleration chip 111 and the general-purpose processing chip 110a is different, the general-purpose processing chip 110a can group multiple operators into two operator combinations according to the computing power ratio between the first acceleration chip 111 and the general-purpose processing chip 110a. The computing power ratio of the two operator combinations is the same as the computing power ratio between the first acceleration chip 111 and the general-purpose processing chip 110a. The two operator combinations are scheduled to the corresponding chips according to the computing power ratio between the first acceleration chip 111 and the general-purpose processing chip 110a.

[0297] Taking the structure of transformer layer i shown in Figure 5 as an example, when the computing power ratio between the first acceleration chip 111 and the general-purpose processing chip 110a is 2:1, as shown in Figure 7B(b), the self-attention operator, feature extraction operator, first normalization operator, and second normalization operator are scheduled to the first acceleration chip 111, and the first linear operator and second linear operator are scheduled to the general-purpose processing chip 110a. In this optional example, the second operator includes the first linear operator and the second linear operator, and the first operator includes the self-attention operator, feature extraction operator, first normalization operator, and second normalization operator.

[0298] In the second alternative implementation, the general-purpose processing chip 110a schedules multiple operators to corresponding chips (or chips) according to the resource utilization rate of each chip (or chip). This ensures that chips (or chips) with high resource utilization rates execute more operators, and chips (or chips) with high resource utilization rates execute fewer operators. Alternatively, it can ensure that chips (or chips) with high resource utilization rates execute computationally intensive operators, and chips (or chips) with low resource utilization rates execute computationally intensive operators.

[0299] For any transformer layer, if the transformer layer includes M operators, the general processing chip 110a schedules the M operators to the corresponding cores (or chips) according to the amount of computation required for each of the M operators and the resource utilization rate of each core (or chip) among the N cores (or chips), so as to balance the load of the cores (or chips) with scheduled operators.

[0300] Taking the computing device 11 as an example, which only includes the first acceleration chip 111, for instance, using the structure of transformer layer i shown in Figure 5 above, if the resource utilization rates of the first acceleration chip 111 and the general-purpose processing chip 110a are the same and both are less than the utilization rate threshold, the six operators in transformer layer i can be evenly distributed and scheduled to the first acceleration chip 111 and the general-purpose processing chip 110a. Here, "evenly distributed" can refer to equal distribution of computational load, meaning that the first acceleration chip 111 and the general-purpose processing chip 110a each execute three operators, as shown in Figure 7A above. "Equal distribution" can also refer to equal distribution of computational load, meaning that the first acceleration chip 111 and the general-purpose processing chip 110a each execute operators with similar computational loads, as shown in Figure 7B(a) above.

[0301] For example, taking the structure of transformer layer i shown in Figure 5 above as an example, when the resource utilization rates of the first acceleration chip 111 and the general processing chip 110a are different, the general processing chip 110a schedules the six operators in transformer layer i to the corresponding chips according to the ratio of the resource utilization rates of the first acceleration chip 111 and the general processing chip 110a, as shown in Figure 7B(b) above, so that the chip with low resource utilization rate executes more operators and the chip with high resource utilization rate executes fewer operators, thereby achieving load balancing.

[0302] In a third alternative implementation, multiple operators in the acceleration chip can be run by different granules. For example, taking the structure of transformer layer i shown in Figure 5 above as an example, when the computing device 110 includes a first acceleration chip 111, the general-purpose processing chip 110a schedules the six operators in transformer layer i to granules 1, 2, and 110a according to the computational load of each operator in transformer layer i, as well as the computational power or resource utilization of granules 1, 2, and 110a. The first linear operator, feature extraction operator, and first normalization operator are scheduled to granule 1; the first linear operator and self-attention operator are scheduled to granule 2; and the second linear operator and second normalization operator are scheduled to 110a.

[0303] The above three optional implementation methods are merely different ways in which the general-purpose processing chip 110a schedules multiple operators in a load-balancing manner. In other embodiments, the general-purpose processing chip 110a can also use other implementation methods to achieve load balancing between different chips or chips. This application does not limit this aspect.

[0304] In some alternative implementations, different accelerator chips have different data processing precision, and different operators also have different requirements for data processing precision. For example, the data processing precision of a self-attention operator is higher than that of a linear operator (either the first or second linear operator). To ensure the reliability of the final inference result, the general-purpose processing chip 110a can schedule multiple operators from each transformer layer to the appropriate chip based on the chip's (general-purpose processing chip or accelerator chip) processing precision. This ensures that the processing precision of each chip running an operator matches the data processing precision of the executed operator. Here, "matching" can mean that the processing precision of the chip is the same as the data processing precision of the executed operator, or that the processing precision of the chip is greater than the data processing precision of the executed operator.

[0305] Taking computing device 11 as an example, which includes only the first acceleration chip 111 and the second acceleration chip 112, the general-purpose processing chip 110a is a CPU with a processing precision of FP16, the first acceleration chip 111 is an NPU with a processing precision of INT8, and the second acceleration chip 112 is a GPU with a processing precision of FP16. The general-purpose processing chip 110a can schedule operators with integer precision in each transformer layer to the first acceleration chip 111, and schedule operators with floating-point precision in each transformer layer to both the first acceleration chip 111 and the general-purpose processing chip 110a.

[0306] Taking the structure of transformer layer i shown in Figure 5 above as an example, the implementation of operator scheduling based on the precision of processed data will be illustrated.

[0307] In the first optional approach, the data processing precision of the feature extraction operator, the first normalization operator, and the first linear operator is integer, while the data processing precision of the self-attention operator, the second linear operator, and the second normalization operator is floating-point, as shown in Figure 7C(a). The general-purpose processing chip 110a schedules the feature extraction operator, the first normalization operator, and the first linear operator to the first acceleration chip 111, and schedules the self-attention operator, the second linear operator, and the second normalization operator to the general-purpose processing chip 110a. In this optional example, the second operator includes: the self-attention operator, the second linear operator, and the second normalization operator, and the first operator includes the feature extraction operator, the first normalization operator, and the first linear operator.

[0308] In some examples, the feature extraction operator, the first normalization operator, and the first linear operator can be executed by chip 1 or chip 2 in the first acceleration chip 111. Alternatively, chip 1 and chip 2 in the first acceleration chip 111 can execute the feature extraction operator, the first normalization operator, and the first linear operator in a load-balanced manner.

[0309] Correspondingly, the self-attention operator and the second linear operator can also be executed by chip 3 or chip 4 in the second acceleration chip 112. Alternatively, chip 3 and chip 4 in the second acceleration chip 112 can execute the self-attention operator and the second linear operator in a load-balanced manner.

[0310] In the second optional approach, the feature extraction operator, the first normalization operator, and the first linear operator have integer processing precision, while the self-attention operator, the second linear operator, and the second normalization operator have floating-point processing precision. The general-purpose processing chip 110a can schedule the self-attention operator, the second linear operator, and the second normalization operator to the general-purpose processing chip 110a and the second acceleration chip 112 respectively, using a load balancing method. As shown in Figure 7C(b), compared to the schematic diagram shown in Figure 7C(a), in Figure 7C(b), the self-attention operator is scheduled to the second acceleration chip 112, and the second linear operator and the second normalization operator are scheduled to the general-purpose processing chip 110a. In this optional example, the second operator includes a second linear operator and a second normalization operator, and the first operator includes a self-attention operator, a feature extraction operator, a first normalization operator, and a first linear operator.

[0311] The two optional methods described above are merely different implementations of how the general-purpose processing chip 110a schedules multiple operators to different chips according to the processing data precision. In other embodiments, the general-purpose processing chip 110a may also use other implementation methods to schedule multiple operators to different chips according to the processing data precision, which are not limited in this application.

[0312] Figure 7C above illustrates, using operator scheduling based on data processing precision as an example, how the general-purpose processing chip 110a schedules different operators to corresponding chips. In other embodiments, the general-purpose processing chip 110a can also employ other implementation methods. For example, the general-purpose processing chip 110a can combine data processing precision and computing power to schedule multiple operators to corresponding chips. As another example, the general-purpose processing chip 110a can combine chip resource utilization and data processing precision to schedule multiple operators to corresponding chips. This application does not limit the scope of these embodiments.

[0313] The following examples illustrate how the general-purpose processing chip 110a schedules different operators to corresponding chips, using operator scheduling methods that combine data processing precision and computing power, and operator scheduling methods that combine chip resource utilization and data processing precision.

[0314] First, an example is given of an operator scheduling method that combines data processing precision and computing power.

[0315] In one alternative implementation, for any transformer layer, the general-purpose processing chip 110a can select one or more candidate execution chips from the computing device 110 corresponding to different processing data precisions, based on the processing data precision of each operator among the multiple operators in that transformer layer. For multiple operators under each processing data precision, the general-purpose processing chip 110a schedules the multiple operators under that processing data precision to the corresponding chip based on the computing power of the candidate execution chips corresponding to that processing data precision.

[0316] Taking the structure of transformer layer i shown in Figure 5 as an example, the data processing precision of the feature extraction operator, the first normalization operator, and the first linear operator is integer, while the data processing precision of the self-attention operator, the second linear operator, and the second normalization operator is floating-point. The computing device 110 includes a first acceleration chip 111, a second acceleration chip 112, a third acceleration chip 113, and a fourth acceleration chip 114. When the processing precision of the first acceleration chip 111 and the fourth acceleration chip 114 is INT8, and the processing precision of the general-purpose processing chip 110a, the second acceleration chip 112, and the third acceleration chip 113 is FP16, the candidate execution chips corresponding to operators with integer data processing precision are the first acceleration chip 111 and the fourth acceleration chip 114. The candidate execution chips corresponding to operators with floating-point data processing precision are the general-purpose processing chip 110a, the second acceleration chip 112, and the third acceleration chip 113.

[0317] Based on the computing power of the chips in the first acceleration chip 111 and the fourth acceleration chip 114, the general-purpose processing chip 110a schedules the feature extraction operator, the first normalization operator, and the first linear operator to the corresponding chips in a load-balanced manner, as shown in Figure 7D. The feature extraction operator is scheduled to chip 1, the first normalization operator is scheduled to chip 2, and the first linear operator is scheduled to chip 7.

[0318] Based on the computing power of the general-purpose processing chip 110a, the second acceleration chip 112, and the third acceleration chip 113, the corresponding general-purpose processing chip 110a schedules the self-attention operator, the second linear operator, and the second normalization operator to the corresponding chips in a load-balanced manner, as shown in Figure 7D. The self-attention operator is scheduled to the second acceleration chip 112, the second linear operator is scheduled to the third acceleration chip 113, and the second normalization operator is scheduled to the general-purpose processing chip 110a.

[0319] In this alternative example, the second operator includes a second normalization operator, which includes a feature extraction operator, a first normalization operator, a first linear operator, a self-attention operator, and a second linear operator.

[0320] Secondly, an example is given of an operator scheduling method that combines chip resource utilization and data processing accuracy.

[0321] In one alternative implementation, similar to the operator scheduling method combining processing data precision and computing power described above, the general-purpose processing chip 110a selects one or more candidate execution chips corresponding to different processing data precisions within the computing device 110. Unlike the operator scheduling method combining processing data precision and computing power described above, for multiple operators at each processing data precision, the general-purpose processing chip 110a schedules the multiple operators at that processing data precision to the corresponding chips based on the resource utilization rate of the candidate execution chips corresponding to that processing data precision.

[0322] Taking the structure of transformer layer i shown in Figure 5 and the schematic diagram shown in Figure 7D as examples, the candidate execution chips corresponding to operators with integer data processing precision are the first acceleration chip 111 and the fourth acceleration chip 114. The candidate execution chips corresponding to operators with floating-point data processing precision are the general-purpose processing chip 110a, the second acceleration chip 112, and the third acceleration chip 113.

[0323] The general-purpose processing chip 110a schedules the feature extraction operator, the first normalization operator, and the first linear operator to the corresponding chips based on the resource utilization rates of the first acceleration chip 111 and the fourth acceleration chip 114. This allows the chip with the lower resource utilization rate to execute more operators, thereby improving resource utilization. For example, if the resource utilization rate of the first acceleration chip 111 is lower than that of the fourth acceleration chip 114, the feature extraction operator and the first normalization operator are scheduled to the first acceleration chip 111. As shown in Figure 7E, if the resource utilization rate of the first acceleration chip 111 is lower than that of the fourth acceleration chip 114, and the resource utilization rate of the fourth acceleration chip 114 is greater than a resource utilization threshold, the feature extraction operator, the first normalization operator, and the first linear operator are scheduled to the first acceleration chip 111. For example, if the resource utilization rate of the first acceleration chip 111 is similar to that of the fourth acceleration chip 114, the feature extraction operator, the first normalization operator, and the first linear operator can be scheduled to the first acceleration chip 111 and the fourth acceleration chip 114, respectively.

[0324] Accordingly, the general-purpose processing chip 110a schedules the self-attention operator, the second linear operator, and the second normalization operator to the corresponding chips based on the resource utilization rates of the general-purpose processing chip 110a, the second acceleration chip 112, and the third acceleration chip 113. This allows chips with lower resource utilization rates to execute more operators, thereby improving resource utilization. For example, as shown in Figure 7E, if the resource utilization rate of the third acceleration chip 113 is greater than that of the second acceleration chip 112 and also greater than that of the general-purpose processing chip 110a, the self-attention operator and the second normalization operator can be scheduled to the second acceleration chip 112, and the second linear operator can be scheduled to the general-purpose processing chip 110a. Alternatively, if the resource utilization rates of the general-purpose processing chip 110a, the second acceleration chip 112, and the third acceleration chip 113 are similar, the self-attention operator, the second linear operator, and the second normalization operator can be scheduled to the general-purpose processing chip 110a, the second acceleration chip 112, and the third acceleration chip 113, respectively.

[0325] In the example provided in Figure 7E, the second operator includes a second linear operator, and the first operator includes a feature extraction operator, a first normalization operator, a first linear operator, a self-attention operator, and a second normalization operator.

[0326] Figures 7A to 7E above are only different implementations of the general-purpose processing chip 110a scheduling different operators to the corresponding chips. In other embodiments, the general-purpose processing chip 110a may also use other implementations to schedule different operators to the corresponding chips. This application does not limit this.

[0327] For example, the general-purpose processing chip 110a can also combine data processing precision, computing power, and resource utilization to schedule different operators to corresponding chips. This allows chips with low resource utilization and high computing power to execute more operators or operators with greater computational demands, while chips with high resource utilization and low computing power can execute fewer operators, operators with smaller computational demands, or none at all. In this way, the computing resources in the computing device 110 can be fully utilized, improving resource utilization.

[0328] For example, taking the structure of transformer layer i shown in FIG5 and the schematic diagram shown in FIG7D as examples, the candidate execution chips corresponding to operators with integer data processing precision are the first acceleration chip 111 and the fourth acceleration chip 114. The candidate execution chips corresponding to operators with floating-point data processing precision are the general-purpose processing chip 110a, the second acceleration chip 112, and the third acceleration chip 113.

[0329] Taking the scheduling of feature extraction operator, first normalization operator, and first linear operator as an example, the scheduling of self-attention operator, second linear operator, and second normalization operator can refer to the scheduling method of feature extraction operator, first normalization operator, and first linear operator.

[0330] When the resource utilization rates of the first acceleration chip 111 and the fourth acceleration chip 114 are similar, the general-purpose processing chip 110a schedules the feature extraction operator, the first normalization operator, and the first linear operator to the corresponding chips in a load-balanced manner, based on the computing power of the first acceleration chip 111 and the fourth acceleration chip 114 and the computing load of the feature extraction operator, the first normalization operator, and the first linear operator, referring to the example provided in Figure 7B above.

[0331] When the resource utilization rate of the first acceleration chip 111 is less than that of the fourth acceleration chip 114, and the resource utilization rate of the fourth acceleration chip 114 is greater than the resource utilization threshold, the general processing chip 110a schedules the feature extraction operator, the first normalization operator, and the first linear operator to the first acceleration chip 111. This reduces the resource utilization rate of the fourth acceleration chip 114.

[0332] The above embodiments are illustrated using the general-purpose processing chip 110a as an example. However, in some optional cases, the operator scheduling process can also be executed by any acceleration chip or a chip in the acceleration chip in the computing device 110, such as the first acceleration chip 111 or chip 1, etc. This application does not limit this.

[0333] Taking the operator scheduling process of chip 1 as an example, after receiving an inference request, chip 1 responds by calling the model deployed locally on computing device 110 and scheduling different operators in each transformer layer of the model to the corresponding chip (or chip). Alternatively, chip 1 schedules different operators in each transformer layer of the model to the corresponding chip (or chip) when it detects that the model has been called. The specific implementation of chip 1 scheduling different operators in each transformer layer of the model to the corresponding chip (or chip) can be referred to the embodiments related to Figures 7A to 7E above, and will not be repeated here.

[0334] Figures 7A to 7E above are only different ways of creating partitioning strategies for multiple operators. In other embodiments, the general-purpose processing chip 110a can also use other methods to create partitioning strategies for multiple operators. This application does not limit this.

[0335] As shown in Figures 7A to 7E, at the operator scheduling node, the general-purpose processing chip 110a divides multiple operators in each transformer layer of the model into at least two operator categories. This can be done by dividing the operators according to their data processing precision, computational complexity, or execution order, thus improving the rationality of the operator division. Different categories of operators are executed by different chips (or chips of different categories), with each chip matching the corresponding operator category. This ensures the reliability of the operator division strategy and improves the accuracy of the model's inference results. Since each operator can be scheduled to a chip that matches it, the execution efficiency of the operator is improved, thereby increasing the model's inference efficiency.

[0336] As shown in Figures 4 to 7E, the computing device 110, during the inference process using the model, employs different chips to execute different operators within the model. Multiple chips work together to execute each transformer layer, achieving heterogeneous scheduling of computing power across the chips. Furthermore, the chips communicate directly via shared memory, reducing communication overhead and shortening inference time, thus improving the model's inference efficiency.

[0337] The following example uses a computing device 110 including a general-purpose processing chip 110a, a first acceleration chip 111, and a second acceleration chip 112. Based on Figures 4 to 7E above, this application provides a model inference method, as shown in Figure 8, which is a flowchart illustrating one of the model inference methods provided in this application. The hardware implementation of the general-purpose processing chip 110a, the first acceleration chip 111, and the second acceleration chip 112 can be referred to the description of the foregoing embodiments, and will not be repeated here.

[0338] The reasoning method for the model provided in this application includes steps S810 to S820.

[0339] The S810, general-purpose processing chip 110a, receives inference requests.

[0340] In some alternative approaches, the inference request may be sent by an external device, which may be a smart wearable device, such as smart glasses, a smartwatch, AR, VR, or other wearable devices. The external device may also be a mobile terminal, such as a mobile phone or tablet. This application does not limit the specific form of the external device.

[0341] In other alternative approaches, the inference request can also be user-inputted. For example, computing device 110 provides a user interface for AI-like applications, where the user inputs an input sequence based on the user interface, and triggers an inference request based on the input sequence.

[0342] Furthermore, in some alternative methods, the inference request can also be automatically triggered by the computing device 110 when the AI ​​function is enabled. For example, when a camera application is running, the computing device 110 automatically triggers an inference request when it detects that AI functions such as AI photography, AI background removal, and AI imaging are enabled. As another example, when a query application is running, the computing device 110 automatically triggers an inference request when it detects that the AI ​​retrieval function is enabled.

[0343] The above three optional methods are merely different ways to trigger inference requests. In other embodiments, inference requests may have other triggering methods, which are not limited in this application.

[0344] In one alternative implementation, the general-purpose processing chip 110a receives and responds to an inference request to determine a partitioning strategy for multiple operators of the model.

[0345] In some alternative approaches, the general-purpose processing chip 110a can read partitioning strategies for multiple operators stored in memory. For example, the memory of the computing device 110 stores partitioning strategies for different models. The general-purpose processing chip 110a retrieves the partitioning strategies for multiple operators corresponding to the model indicated by the inference request from the partitioning strategies of different models. The model identifier can be the name of the model or the name of the application that triggered the inference request; the specific form of the model identifier is not limited in this embodiment.

[0346] In some alternative approaches, the general-purpose processing chip 110a may refer to the operator scheduling phase described above to create a partitioning strategy for multiple operators of the model.

[0347] Please refer to Figure 8. As shown in Figure 1(a) above, the model includes multiple data processing layers, normalization layers, linear layers, and output layers, such as processing layers 1 to L. For each data processing layer, the network structure is the same as the structure of the transformer layer provided in Figure 5 above. Taking the first data processing layer as an example, the second operator includes a second linear operator and a second normalization operator, and the first operator includes a feature extraction operator, a first normalization operator, a first linear operator, and a self-attention operator. The feature extraction operator, the first normalization operator, and the first linear operator in the first data processing layer are scheduled to the first acceleration chip 111, the self-attention operator is scheduled to the second acceleration chip 112, and the second linear operator and the second normalization operator are scheduled to the general-purpose processing chip 110a. The partitioning strategy for each operator in the remaining data processing layers is similar to that in the first data processing layer, and will not be repeated here. The first data processing layer can be any one of the multiple data processing layers. Alternatively, the first data processing layer may be the first data processing layer executed among multiple data processing layers, such as processing layer 1 in Figure 8. This application does not limit the specific form of the data processing layer in its embodiments.

[0348] In one alternative implementation, the first data processing layer is any one of a plurality of data processing layers, and the first input data of the first data processing layer is the output of the preceding data processing layer. Alternatively, the first input data of the first data processing layer is the data corresponding to the output of the preceding data processing layer and the input sequence.

[0349] In another alternative implementation, the first data processing layer is the first data processing layer executed among multiple data processing layers, and the first input data of the first data processing layer is data related to the inference request.

[0350] The following is an example illustrating the data corresponding to the input sequence.

[0351] In a first alternative approach, the entire large model can be run in computing device 110. For example, the model shown in Figure 1 above can be run in computing device 110. The corresponding "data corresponding to the input sequence" can be the data after processing the input sequence through the input layer in Figure 1 above. The data corresponding to the input sequence can be a matrix, an encoding, or other form of data, which is not limited in this application.

[0352] In the second alternative approach, computing device 110 can run a subset of the processing layers in a large model. For example, taking a large model comprising an encoder and a decoder, the encoder is run by an external device (e.g., a server or other computing device), while the decoder is run by computing device 110. Both the decoder and encoder include multiple transformer layers. Thus, computationally intensive feature extraction operations are performed by the external device, while computationally less intensive feature inference operations are implemented by computing device 110. The external device and computing device 110 collaboratively provide the computational power required for model inference, reducing the computational power consumption of computing device 110 and improving the model's inference speed. Furthermore, during the decoding process on computing device 110, for each transformer layer in the decoder, the first and second operators of that transformer layer are scheduled to the acceleration chip and the general-purpose processing chip 110a, respectively, further enhancing the model's inference speed.

[0353] Taking an external device as the server as an example, the server can be a cloud server or other server, server cluster, or high-performance computing device with feature extraction capabilities. This application does not limit the specific form of the server.

[0354] Accordingly, when the computing device 110 is running the decoder, the "data corresponding to the input sequence" is the feature sequence corresponding to the input sequence. This feature sequence can be used to indicate the contextual information of the input sequence.

[0355] For example, the input sequence is an image, and the corresponding feature sequence can include, but is not limited to, visual features, image semantic features, and style features. Visual features represent the overall texture information of the input sequence, image semantic features represent the semantic information of the input sequence, and style features represent the image style information of the input sequence. The image can be a single image, an image sequence, or a video image.

[0356] For example, if the input sequence is text, the corresponding feature sequence can include, but is not limited to, text semantic features and text style features. Text semantic features represent the contextual semantic information of the input sequence. Text style features represent the tone and style of the input sequence. For instance, taking image-to-text conversion, the server generates features based on the input sequence to obtain the text semantic features in the image. Similarly, taking text summarization generation, the server generates features based on the input sequence to obtain both text semantic features and text style features from the text data.

[0357] For example, the input sequence may consist of text data and image data, and the corresponding feature sequence may include image-text fusion features. These image-text fusion features are used to characterize one or more of the following: overall texture information of the image, semantic information of the image, image style information, contextual semantic information of the text, and text style features.

[0358] The two options above are simply different ways of displaying the data corresponding to the input sequence. In other embodiments, the data corresponding to the input sequence may have other forms, which this application does not limit.

[0359] S820, the general-purpose processing chip 110a controls the acceleration chip to sequentially execute the first operator of multiple data processing layers, and controls the general-purpose processing chip 110a to sequentially execute the second operator of multiple data processing layers to obtain the inference result of the inference request.

[0360] In one alternative approach, "sequentially executing the first operators of multiple data processing layers" can be the accelerator chip sequentially executing the first operators of the first data processing layer and the first operators of the second data processing layer. Here, the first and second data processing layers are two adjacent data processing layers. The first data processing layer can be any one of the multiple data processing layers. For example, taking L data processing layers as an example, the accelerator chip executes them in the order of first operator L1 → first operator L2 → ... → first operator LL according to the execution order of the multiple data processing layers. Here, first operator L1 refers to the first operator of the first data processing layer, such as the feature extraction operator, first normalization operator, first linear operator, and self-attention operator of transformer layer 1 in Figure 5 above. Correspondingly, first operator L2 is the first operator of the second data processing layer, such as the first operator of transformer layer 2 in Figure 5 above. First operator LL is the first operator of the Lth data processing layer, such as the first operator of transformer layer L in Figure 5 above.

[0361] Similarly, "sequentially executing the second operators of multiple data processing layers" can mean that the general-purpose processing chip 110a sequentially executes the second operators of the first data processing layer and the second operators of the second data processing layer. For example, taking L data processing layers as an example, the general-purpose processing chip 110a executes the multiple data processing layers in the order of second operator L1 → second operator L2 → ... → second operator LL. Here, second operator L1 refers to the second operator of the first data processing layer, such as the second linear operator and the second normalization operator of transformer layer 1 in Figure 5 above. Correspondingly, second operator L2 is the second operator of the second data processing layer, such as the second operator of transformer layer 2 in Figure 5 above. Second operator LL is the second operator of the Lth data processing layer, such as the second operator of transformer layer L in Figure 5 above.

[0362] The accelerator chip sequentially executes the first operator of multiple data processing layers, and the general-purpose processing chip 110a sequentially executes the second operator of multiple data processing layers. This can be achieved by the accelerator chip and the general-purpose processing chip 110a executing each data processing layer in the order of the multiple data processing layers, and within each data processing layer, executing in the order of first operator → second operator. That is, the accelerator chip and the general-purpose processing chip 110a execute in the order of first operator L1 → second operator L1 → first operator L2 → second operator L2 → ... → first operator LL → second operator LL.

[0363] For example, taking the structure of the transformer layer provided in Figure 5 above as an example, as shown in Figure 9A, the data 1 output by transformer layer 1 is used as the input of transformer layer 2, the data 2 output by transformer layer 2 is used as the input of the next transformer layer, and the input of transformer layer L is the data L-1 output by transformer layer L-1. Taking transformer layer 1 as an example, the intermediate data 1 output by the first operator L1 is used as the input of the second operator L1, and the data 1 output by the second operator L1 is used as the input of the first operator L1 of the next transformer layer.

[0364] In one alternative implementation, the general-purpose processing chip 110a takes the data corresponding to the input sequence as the input to the acceleration chip, controls the acceleration chip to execute the first operator of multiple data processing layers sequentially, and controls the general-purpose processing chip 110a to execute the second operator of multiple data processing layers sequentially, so as to process the data corresponding to the input sequence and obtain the inference result of the inference request.

[0365] In the first alternative approach, "control" can be that the general-purpose processing chip 110a sends an instruction to the general-purpose processing chip 110a or the accelerator chip to start operator execution, the general-purpose processing chip 110a or the accelerator chip executes the corresponding instruction to start operator execution, and the general-purpose processing chip 110a or the accelerator chip sends an instruction to stop operator execution, the general-purpose processing chip 110a or the accelerator chip executes the corresponding instruction to stop operator execution.

[0366] In the second alternative approach, "control" can also be that the general-purpose processing chip 110a sends data or the address of the data in shared memory to the general-purpose processing chip 110a or the accelerator chip, so that the general-purpose processing chip 110a or the accelerator chip starts operator execution after reading the data.

[0367] The two optional methods mentioned above are only different ways in which the general-purpose processing chip 110a controls the general-purpose processing chip 110a and the acceleration chip. In other embodiments, the general-purpose processing chip 110a can control the general-purpose processing chip 110a and the acceleration chip in other ways, which are not limited in this application.

[0368] Taking the "data corresponding to the input sequence" as an example, the general-purpose processing chip 110a receives an inference request. Based on the input sequence carried in the inference request, it sends a first request to the server. The server responds to the first request by calling the encoder to process the input sequence and generate the feature sequence corresponding to the input sequence. The server quantizes the feature sequence corresponding to the input sequence to obtain a quantized feature sequence and sends the quantized feature sequence to the computing device 110. The general-purpose processing chip 110a receives the quantized feature sequence, dequantizes it, and obtains the feature sequence corresponding to the input sequence. The general-purpose processing chip 110a uses the feature sequence corresponding to the input sequence as input to the decoder, controls the acceleration chip to sequentially execute the first operator of multiple data processing layers, and controls the general-purpose processing chip 110a to sequentially execute the second operator of multiple data processing layers to generate the inference result corresponding to the feature sequence.

[0369] Compared to the server-side model inference method, the server sends the feature sequence corresponding to the input sequence to the computing device 110, instead of the inference result corresponding to the input sequence. Compared to the inference result corresponding to the input sequence, the feature sequence corresponding to the input sequence has a smaller data volume and requires less bandwidth, improving transmission speed and thus improving the display rate of the inference result. Compared to the method where the entire inference process is implemented by the computing device 110, having the server perform the more computationally demanding feature extraction operation reduces the computing power consumption of the computing device 110, thereby reducing the load on the computing device 110 during the inference process and improving inference speed. Furthermore, the computing device 110 schedules multiple operators to different execution chips within the same computing device at the granularity of data processing layer operators, which can fully utilize the computing resources within the computing device 110, reduce the load on individual chips, thereby reducing the execution time required by individual chips, and thus reducing the total time required for the computing device 110 to perform model inference, thereby improving inference speed.

[0370] Taking the example that "data corresponding to the input sequence" can be the processed data of the input sequence, in the first optional implementation, the general-purpose processing chip 110a encodes the input sequence included in the inference request to obtain the data corresponding to the input sequence. The general-purpose processing chip 110a inputs the data corresponding to the input sequence into the first transformer layer in the model. Multiple transformer layers process the data corresponding to the input sequence in a pipeline manner to obtain the inference result of the input sequence. Taking the first data processing layer as the first data processing layer executed among multiple data processing layers as an example, as shown in Figure 9B(a), the general-purpose processing chip 110a inputs the data corresponding to the input sequence into the first data processing layer, and the first data processing layer processes the data corresponding to the input sequence. The output of the first data processing layer is input into the second data processing layer, and the second data processing layer processes the output of the first data processing layer. The second data processing layer is the next data processing layer adjacent to the first data processing layer.

[0371] In the second optional implementation, the general-purpose processing chip 110a uses the data corresponding to the input sequence as the input to each transformer layer in the model, and uses multiple transformer layers to process the data corresponding to the input sequence to obtain the inference result of the input sequence.

[0372] For example, as shown in Figure 9B(b), the general-purpose processing chip 110a inputs the data corresponding to the input sequence into a first data processing layer and a second data processing layer. The first data processing layer processes the data corresponding to the input sequence. The output of the first data processing layer is input into the second data processing layer. The second data processing layer processes both the output of the first data processing layer and the data corresponding to the input sequence.

[0373] The two optional implementation methods mentioned above are only different ways for the general-purpose processing chip 110a to input the data corresponding to the input sequence into the model. In other embodiments, the general-purpose processing chip 110a may also use other implementation methods to input the data corresponding to the input sequence into the model. This application embodiment does not limit this.

[0374] During model inference, for each transformer layer, the general-purpose processing chip 110a schedules the operators in that transformer layer to the corresponding execution chip (general-purpose processing chip 110a or accelerator chip) according to a strategy of partitioning multiple operators. The general-purpose processing chip 110a and the accelerator chip run the operators in that transformer layer in a pipelined parallel manner, thereby processing the input to that transformer layer. Regarding the pipelined parallel method, please refer to the embodiments provided in Figures 11A to 13 below; the embodiments of this application will not be elaborated upon here.

[0375] The acceleration chip includes one or more of the first acceleration chip 111, the second acceleration chip 112, the third acceleration chip 113, and the fourth acceleration chip 114 mentioned above. For example, as shown in Figure 8, the acceleration chip includes the first acceleration chip 111 and the second acceleration chip 112.

[0376] Based on the embodiment provided in Figure 8, during model inference, data is transferred between chips within the same device, without involving data transfer between different devices. This reduces communication overhead, thereby reducing inference time and improving inference speed. Furthermore, this application schedules multiple operators to different execution chips within the same computing device at the operator level, fully utilizing the computing resources within the computing device 110, reducing the load on individual chips, and thus reducing the execution time required by individual chips. This reduces the total time required for the computing device 110 to perform model inference, thereby improving inference speed.

[0377] Figure 8 above illustrates the inference method of the model provided in this application using the general-purpose processing chip 110a as the execution entity. In some optional implementations, the inference method of the model provided in this application can also be executed by other components. For example, the component can be the aforementioned acceleration chip, or a chip within the acceleration chip, or other components or computing cards with operator scheduling or instruction issuance functions.

[0378] Furthermore, Figure 8 above illustrates the example of a first operator including a feature extraction operator, a first normalization operator, a first linear operator, and a self-attention operator, and a second operator including a second linear operator and a second normalization operator. In practical applications, the first and second operators may have other forms, which are not limited in this application.

[0379] Furthermore, Figure 8 above uses an acceleration chip including a first acceleration chip 111 and a second acceleration chip 112 as an example. In practical applications, the computing device 110 may include more or fewer acceleration chips than those in Figure 8 above.

[0380] Furthermore, Figure 9B above illustrates a model with two data processing layers. In other embodiments, the model may include more processing layers, and different processing layers may include different operators, with different partitioning strategies for the operators in different processing layers. For example, for each processing layer, different operators in that layer can be scheduled to different chips to reduce the load on a single chip. Alternatively, for another example, for each processing layer, different operators in that layer can be scheduled to different chips on different chips to reduce the load on a single chip.

[0381] Based on Figures 8, 9A, and 9B, in order to improve the inference efficiency of the model during the inference stage, for each data processing layer, the general-purpose processing chip 110a and the acceleration chip run the operators in that data processing layer in a pipeline manner to make full use of the chips in the computing device 110, thereby accelerating the inference process.

[0382] Regarding the pipeline approach, the following example illustrates the process using the first data processing layer, which is executed first among multiple data processing layers.

[0383] In the first optional implementation, pipeline parallelism is achieved using the data corresponding to the entire input sequence as the data granularity. For example, the general-purpose processing chip 110a uses the data corresponding to the entire input sequence as the input to the first data processing layer, and the general-purpose processing chip 110a and the acceleration chip process the data corresponding to the entire input sequence in a pipeline manner of general-purpose processing chip 110a → acceleration chip → general-purpose processing chip 110a (or acceleration chip → general-purpose processing chip 110a → acceleration chip).

[0384] In the second optional implementation, pipeline parallelism is achieved using sub-data within the data corresponding to the input sequence as the data granularity. For example, the general-purpose processing chip 110a segments the data corresponding to the input sequence into multiple sub-data. The general-purpose processing chip 110a uses these multiple sub-data as input to the first data processing layer. The general-purpose processing chip 110a and the acceleration chip process each sub-data in a pipelined manner of general-purpose processing chip 110a → acceleration chip → general-purpose processing chip 110a (or acceleration chip → general-purpose processing chip 110a → acceleration chip).

[0385] In some alternative approaches, sub-data may also be referred to as data, data segment, token, sub-sequence, sub-feature, feature sub-sequence, or other names, and the embodiments of this application do not limit this. For example, taking a sub-sequence as an example, the data corresponding to the input sequence is segmented to obtain multiple sub-sequences, such as the first sub-sequence and the second sub-sequence. As another example, taking data as an example, the data corresponding to the input sequence is segmented to obtain multiple data, such as the first data and the second data.

[0386] The two optional implementation methods mentioned above are merely different options for pipeline parallelism. In other embodiments, pipeline parallelism can also be implemented in other ways, which are not limited in this application.

[0387] For example, when the chip's computing power meets the requirements, the general-purpose processing chip 110a and the accelerator chip implement pipelined parallelism using the data corresponding to the entire input sequence as the data granularity. When the chip's computing power does not meet the requirements, the general-purpose processing chip 110a and the accelerator chip implement pipelined parallelism using the sub-data of the data corresponding to the input sequence as the data granularity. The computing power requirement can be a computing power threshold, such as the bandwidth threshold, frequency threshold, or floating-point arithmetic threshold mentioned above.

[0388] For example, when the length of the data corresponding to the input sequence is greater than or equal to a length threshold, the general-purpose processing chip 110a and the accelerator chip implement pipelined parallelism using sub-data as the data granularity. When the length of the data corresponding to the input sequence is less than the length threshold, the general-purpose processing chip 110a and the accelerator chip implement pipelined parallelism using the entire data corresponding to the input sequence as the data granularity.

[0389] For example, when the chip's computing power meets the requirements and the length of the data corresponding to the input sequence is less than a length threshold, or when the chip's computing power meets the requirements and the length of the data corresponding to the input sequence is greater than or equal to the length threshold, the general-purpose processing chip 110a and the accelerator chip implement pipelined parallelism with the data corresponding to the entire input sequence as the data granularity. When the chip's computing power does not meet the requirements and the length of the data corresponding to the input sequence is greater than or equal to the length threshold, the general-purpose processing chip 110a and the accelerator chip implement pipelined parallelism with the sub-data of the data corresponding to the input sequence as the data granularity. This application does not limit this aspect.

[0390] The following examples, in conjunction with Figures 10 to 13, illustrate two approaches to pipeline parallelism: one using the data granularity of the entire input sequence as the data granularity, and the other using the data granularity of sub-data within the input sequence as the data granularity.

[0391] First, we will use the data corresponding to the entire input sequence as the data granularity to illustrate the pipeline parallelism method.

[0392] Please refer to Figure 10, which is a schematic diagram of the model inference stage provided in this application. Taking the first data processing layer, which is executed first among multiple data processing layers, as an example, the processing procedure of this model inference stage includes S821A to S824A.

[0393] The S821A general-purpose processing chip 110a schedules multiple operators of the first data processing layer to the corresponding execution chips according to the division strategy of multiple operators.

[0394] Taking the execution chip as an example, which includes a general-purpose processing chip 110a, a first acceleration chip 111, and a second acceleration chip 112. As shown in Figure 10, the feature extraction operator, the first normalization operator, and the first linear operator in the first data processing layer are scheduled to the first acceleration chip 111, the self-attention operator is scheduled to the second acceleration chip 112, and the second linear operator and the second normalization operator are scheduled to the general-purpose processing chip 110a.

[0395] In a first optional implementation, after determining the execution chip corresponding to each operator in the first data processing layer, the general-purpose processing chip 110a can send the scripts or code files of each operator in the first data processing layer to the corresponding execution chips, so as to schedule multiple operators of the first data processing layer to the corresponding execution chips. For example, the general-purpose processing chip 110a sends the scripts or code files of the feature extraction operator, the first normalization operator, and the first linear operator to the first acceleration chip 111.

[0396] In the second optional implementation, the general-purpose processing chip 110a loads the scripts or code files of each operator in the first data processing layer from memory into shared memory, and sends scheduling instructions to the corresponding execution chips based on the execution chips corresponding to each operator in the first data processing layer. These scheduling instructions indicate which operators the execution chips need to execute. The execution chips respond to the scheduling instructions and execute the corresponding operators in the shared memory.

[0397] For example, the general-purpose processing chip 110a loads the script or code files of the feature extraction operator, the first normalization operator, the first linear operator, the self-attention operator, the second linear operator, and the second normalization operator from the memory into the shared memory, and sends a scheduling instruction to the first acceleration chip 111 according to the partitioning strategy of multiple operators. In response to the scheduling instruction, the first acceleration chip 111 loads the script or code files of the feature extraction operator, the first normalization operator, and the first linear operator from the shared memory into the memory of the first acceleration chip 111.

[0398] The two optional implementation methods mentioned above are merely different ways for the general-purpose processing chip 110a to schedule multiple operators of the first data processing layer to the corresponding execution chips. In other embodiments, the general-purpose processing chip 110a may also use other implementation methods to schedule multiple operators of the first data processing layer to the corresponding execution chips. This application does not limit this.

[0399] The S822A general-purpose processing chip 110a inputs the data corresponding to the input sequence into the first execution chip to obtain the first processing result of the input sequence.

[0400] The first execution chip is the execution chip corresponding to the first operator executed among the multiple operators in the first data processing layer. In some optional embodiments, the data corresponding to the input sequence may also be referred to as the first input data of the first data processing layer or the input data of the first data processing layer; this embodiment of the application does not limit this.

[0401] As shown in Figure 10, the first operator executed among the multiple operators in the first data processing layer is the first normalization operator, and the first execution chip is the first acceleration chip 111.

[0402] In some alternative approaches, the general-purpose processing chip 110a writes the data corresponding to the input sequence into shared memory and sends a first instruction to the first acceleration chip 111. In response to the first instruction, the first acceleration chip 111 reads the data corresponding to the input sequence from the shared memory into its own memory and uses this data as input to the feature extraction operator.

[0403] In some alternative embodiments, the general-purpose processing chip 110a sends a second instruction to the first acceleration chip 111 based on the data corresponding to the input sequence. In response to the second instruction, the first acceleration chip 111 acquires the data corresponding to the input sequence carried in the second instruction and uses this data as input to the feature extraction operator.

[0404] The two optional methods mentioned above are merely different implementations of the general-purpose processing chip 110a inputting the data corresponding to the input sequence into the first execution chip. In other embodiments, the general-purpose processing chip 110a may also use other implementations to input the data corresponding to the input sequence into the first execution chip, and this application embodiment does not limit this.

[0405] In one optional implementation, after receiving the data corresponding to the input sequence, the first execution chip processes the input sequence sequentially using operators in the order of execution to obtain a first processing result. For example, as shown in Figure 10, the first normalization operator and the first linear operator are executed in a continuous order. The first acceleration chip 111 executes the first normalization operator and the first linear operator sequentially to process the data corresponding to the input sequence and obtain the first processing result. Alternatively, if the first normalization operator and the second linear operator are scheduled to the first acceleration chip 111, and their execution order is not continuous, then the first acceleration chip executes the first normalization operator to process the input sequence and obtain the first processing result.

[0406] In an alternative approach, where the first execution chip includes a core, the core can process the input sequence sequentially by executing operators in a sequential order.

[0407] In another alternative approach, where the first execution chip comprises multiple cores, the first execution chip can schedule different operators to different cores. The first execution chip processes the input sequence sequentially through different cores according to the execution order of the operators. Taking the first acceleration chip 111 as an example, core 1 deploys a feature extraction operator and a first linear operator, while core 2 deploys a first normalization operator. Then, the first acceleration chip 111 processes the data corresponding to the input sequence sequentially through the first normalization operator and the first linear operator, following the order of core 1 → core 2.

[0408] In one alternative implementation, to facilitate data pass-through between different chips and reduce communication overhead, the first execution chip writes the first processing result into shared memory after obtaining the first processing result.

[0409] In some alternative approaches, after the first processing result is written to the shared memory, the general-purpose processing chip 110a sends a second instruction to the second execution chip, which instructs the second execution chip to read the second processing result from the shared memory and execute S823A as described below.

[0410] The second execution chip can refer to the execution chip corresponding to the operator executed subsequently among the multiple operators in the first data processing layer. "The operator executed subsequently among the multiple operators" can be the operator whose execution order is adjacent to the operator on the first execution chip. For example, taking Figure 10 as an example, if the operator executed subsequently among the multiple operators is a self-attention operator, then the second execution chip refers to the second acceleration chip 112.

[0411] Alternatively, "the operator executed subsequently among multiple operators" can also be the operator executed after the operator on the first execution chip. For example, taking Figure 10 as an example, the operators executed subsequently among multiple operators include the self-attention operator, the second linear operator, and the second normalization operator. Correspondingly, the second execution chip may include the second acceleration chip 112 and the general processing chip 110a.

[0412] The following explanation uses the example of "the operator executed subsequently among multiple operators" to illustrate this point.

[0413] For example, after the first accelerator chip 111 writes the first processing result to the shared memory, the first accelerator chip 111 sends a first response to the general-purpose processing chip 110a. In response to the first response, the general-purpose processing chip 110a sends a second instruction to the second accelerator chip 112. The first response indicates that the first accelerator chip 111 has completed the operator processing and written the processing result to the shared memory.

[0414] In some alternative embodiments, after the first processing result is written to shared memory, the first execution chip sends a first response to the second execution chip that processes subsequent operators. In response to the first response, the second execution chip reads the second processing result from shared memory and executes S823A as described below.

[0415] The two optional methods described above are merely different implementations of triggering the second execution chip to execute subsequent operators. In other embodiments, there may be other implementations for triggering the second execution chip to execute operators. For example, the second execution chip periodically detects the shared memory. When it detects that the first execution chip is writing data to the shared memory, the second execution chip reads the second processing result from the shared memory and executes S823A as described below. This application does not limit this aspect.

[0416] S823A, the second execution chip takes the first processing result as input to the second execution chip and obtains the second processing result of the input sequence.

[0417] In one alternative implementation, the second execution chip can refer to the processing procedure of the first execution chip described above to obtain a second processing result of the input sequence. Further details are omitted here.

[0418] In one alternative implementation, similar to the processing method of the first processing result described above, the second execution chip writes the second processing result into shared memory. As shown in Figure 10, the second acceleration chip 112 processes the first processing result using a self-attention operator to obtain the second processing result of the input sequence, and writes the second processing result of the input sequence into shared memory.

[0419] S824A obtains the output of the first data processing layer based on the second processing result of the input sequence.

[0420] In a first optional implementation, when the execution chips for executing the first data processing layer include only a first execution chip and a second execution chip, the second execution chip outputs the output result of the first data processing layer based on the second processing result of the input sequence. For example, the second execution chip uses the second processing result as the output result of the first data processing layer so that subsequent data processing layers can process the output result of the first data processing layer. Another example is that the second execution chip normalizes the second and first processing results to obtain the initial output result of the first data processing layer. Then, it performs feature extraction on the initial output result of the first data processing layer using the feature extraction operator in the first execution chip to obtain the final output result of the first data processing layer, and writes the final output result of the first data processing layer into shared memory.

[0421] In the second optional implementation, if the execution chip executing the first data processing layer also includes a third execution chip, the second execution chip writes the second processing result of the input sequence into the shared memory, and the third execution chip triggers the operator to run, reads the second processing result from the shared memory, processes the second processing result, and obtains the output result of the first data processing layer.

[0422] The third execution chip can refer to the chip used to execute the subsequent operator of the operator executed by the second execution chip. The third execution chip can be the first execution chip mentioned above, or the third execution chip can be a chip in the computing device 110 other than the first execution chip and the second execution chip mentioned above, as shown in FIG10, the third execution chip is the general processing chip 110a mentioned above.

[0423] For example, as shown in Figure 10, after the second processing result of the input sequence is written to the shared memory, the general-purpose processing chip 110a triggers the operator operation to read the second processing result from the shared memory (S11). The general-purpose processing chip 110a processes the second processing result sequentially through the second linear operator and the second normalization operator to obtain the third processing result of the input sequence (S12). The general-purpose processing chip 110a writes the third result of the input sequence into the shared memory (S13). The first execution chip reads the third processing result from the shared memory (S14). The first execution chip processes the third result through the feature extraction operator to obtain the output result of the first data processing layer (S15).

[0424] The above two methods are merely different implementations of obtaining the output results of the data processing layer. In other embodiments, other implementation methods can also be used to obtain the output results of the data processing layer, and this application does not limit these methods.

[0425] Based on the embodiment shown in Figure 10, for each processing layer, the general-purpose processing chip 110a and the acceleration chip schedule multiple operators in the processing to different chips through a heterogeneous computing power scheduling method. During the inference process of the processing layer, the general-purpose processing chip 110a and the acceleration chip process the input sequence in a pipeline manner. This reduces the load on a single chip during the model inference process, thereby shortening the execution time required for a single processing layer, thus reducing the total time required for the computing device 110 to perform inference through the model, and thereby improving the inference speed.

[0426] In one alternative implementation, to facilitate recording the processing results of different operators, the first execution chip writes the intermediate result of an operator to shared memory after executing it. Similarly, the second execution chip writes the intermediate result of an operator to shared memory after executing it, and the third execution chip writes the intermediate result of an operator to shared memory after executing it.

[0427] Alternatively, after executing a specific operator, the execution chip (first execution chip, second execution chip, or third execution chip) writes the processing result of the specific operator into shared memory. The specific operator can be an operator preceding the normalization operator (first normalization operator or second normalization operator). Alternatively, the specific operator can be a second linear operator. This application does not limit the specific form of the specific operator in its embodiments.

[0428] For example, taking the second linear operator as an example, the data transfer method between processing layers during model inference is illustrated below. As shown in Figure 11A, Figure 11A is a schematic diagram of data transfer between processing layers provided in this application. The network structure of processing layer 1 and processing layer 2 is the same as that of transformer layer 1 in Figure 5 above. The partitioning strategy of multiple operators in processing layer 1 and processing layer 2 is the same, and the partitioning strategy of multiple operators is the same as that of multiple operators in the data processing layer shown in Figure 10 above.

[0429] For processing layer 1, the first execution chip, the second execution chip, and the third execution chip, according to the embodiment provided in FIG10 above, obtain the output result 1 of processing layer 1, write the output result 1 of processing layer 1 into shared memory, and the third execution chip writes the intermediate result of the second linear operator into shared memory. After the processing of processing layer 2 is started, the first execution chip reads the output result 1 of processing layer 1 and the intermediate result of the second linear operator from shared memory, and uses the output result 1 of processing layer 1 and the intermediate result of the second linear operator as the input of the first execution chip. The subsequent processing process of processing layer 2 can refer to the processing process of the data processing layer in FIG10 above, which will not be described in detail here.

[0430] Furthermore, during the self-attention operator processing of the second execution chip in processing layer 2, the second execution chip reads the output result 1 of processing layer 1 and the intermediate result of the second linear operator from the shared memory, and uses the output result 1 of processing layer 1, the intermediate result of the second linear operator, and the output of the first execution chip as the input of the first execution chip.

[0431] As can be seen from the embodiment provided in Figure 11A, during model inference, data pass-through is achieved between processing layers through shared memory, reducing communication overhead and thus shortening inference time. Furthermore, during inference, intermediate results of operators and output results of processing layers are transmitted between them. With a single device processing a single inference request, the amount of data transmitted is small, and the memory footprint is minimal, enabling the application of large model inference on memory-constrained devices.

[0432] Exemplarily, the description takes an acceleration chip including a first acceleration chip 111 as an example. As shown in Figure 11B, Figure 11B is a second schematic diagram of data transfer between processing layers provided in this application. Unlike the data transfer process shown in Figure 11A above, in Figure 11B, the feature extraction operator, the first normalization operator, the first linear operator, and the self-attention operator are executed by the first acceleration chip 111, and the second linear operator and the second normalization operator are executed by the general processing chip 110a. The data corresponding to the input sequence is input to the first acceleration chip 111. The first acceleration chip 111 sequentially executes the first normalization operator, the first linear operator, and the self-attention operator of processing layer 1 to process the data corresponding to the input sequence and obtain a first processing result. The first processing result is input to the general processing chip 110a. The general processing chip 110a sequentially executes the second linear operator and the second normalization operator of processing layer 1 to process the first processing result and obtain a second processing result. The second processing result is input to the first acceleration chip 111. The first accelerator chip 111 sequentially executes the feature extraction operator of processing layer 1, the first normalization operator, the first linear operator, and the self-attention operator of processing layer 2 to process the second processing result and obtain a third processing result. The third processing result is input into the general processing chip 110a. The general processing chip 110a sequentially executes the second linear operator and the second normalization operator of processing layer 2 to process the third processing result and obtain a fourth processing result. The fourth processing result is input into the first accelerator chip 111. The first accelerator chip 111 sequentially executes the feature extraction operator of processing layer 2, the first normalization operator, the first linear operator, and the self-attention operator of processing layer 3 to process the fourth processing result.

[0433] As can be seen from the embodiment provided in Figure 11B, during model inference, data is transferred between operators at the operator level, rather than between processing layers at the processing layer level. Compared to the transfer method between processing layers, the amount of data transferred between operators is smaller, which can reduce the amount of data transferred between different chips. Furthermore, each chip (general-purpose processing chip 110a or first acceleration chip 111) independently executes a portion of the operators in the processing layer. The execution complexity of the operators is low, resulting in a low load on a single chip, thereby accelerating the processing speed of a single chip and improving the inference speed of the model. Next, the pipeline parallelism method is illustrated by taking the sub-data in the data corresponding to the input sequence as the data granularity.

[0434] Taking sub-data as a sub-sequence as an example, during the processing of the data input model corresponding to the input sequence, the general-purpose processing chip 110a divides the data corresponding to the input sequence into multiple sub-sequences using the segmentation length. Each sub-sequence is used as the input to the model, allowing the model to perform inference on each sub-sequence separately, as shown in Figure 6A above. In this way, for each chip, the amount of data processed by each chip at one time is reduced, and the processing time of each chip is shortened, thereby shortening the overall inference time of the model and further improving the inference efficiency of the model.

[0435] For example, taking an input sequence whose corresponding data includes a first subsequence and a second subsequence as an example. For a first data processing layer, the acceleration chip runs the first operator of the first data processing layer to sequentially process the first data and the second data, obtaining a first result corresponding to the first data and a second result corresponding to the second data. Here, the first data can be the first subsequence. Alternatively, the first data can be the output result of the previous data processing layer on the first subsequence. Correspondingly, the second data can be the second subsequence, or the second data can also be the output result of the previous data processing layer on the first subsequence. This application does not limit this.

[0436] The acceleration chip transmits the first result corresponding to the first data and the second result corresponding to the second data to the general processing chip 110a in the order of first result → second result. The general processing chip 110a runs the second operator of the first data processing layer to process the first result and the second result in sequence, and obtains the third result corresponding to the first result and the third result of the second result.

[0437] For example, consider a first operator comprising a first normalization operator, a first linear operator, and a self-attention operator, and a second operator comprising a second linear operator and a second normalization operator. When the accelerator chip is a first accelerator chip 111, the first accelerator chip 111 sequentially executes the first normalization operator, the first linear operator, and the self-attention operator of the first data processing layer to process the first data, obtaining a first result corresponding to the first data. The first accelerator chip 111 then continues to sequentially execute the first normalization operator, the first linear operator, and the self-attention operator of the first data processing layer to process the second data, obtaining a second result corresponding to the second data.

[0438] After obtaining the first result, the general-purpose processing chip 110a sequentially executes the second linear operator and the second normalization operator of the first data processing layer to process the first result, thereby obtaining the third result corresponding to the first result. After obtaining the second result, the general-purpose processing chip 110a sequentially executes the second linear operator and the second normalization operator of the first data processing layer to process the second result, thereby obtaining the fourth result corresponding to the second result.

[0439] Thus, for each data processing layer, the first acceleration chip 111 (or the general-purpose processing chip 11a) processes each input in the order of first data → second data (first result → second result), reducing the amount of data processed by each chip each time, accelerating the chip's data processing speed, and thereby improving the model's inference speed. Furthermore, each data processing layer operates in the order of first operator → second operator, with the first and second operators executed by different chips. Each operator (either the first or the second operator) runs independently, reducing the load on each chip and minimizing idle time, thereby improving chip resource utilization and ultimately increasing the model's inference speed.

[0440] For example, consider a scenario where the first operator includes a feature extraction operator, a first normalization operator, a first linear operator, and a self-attention operator, and the second operator includes a second linear operator and a second normalization operator. Referring to the embodiment provided in Figure 6A above, the acceleration chip and the general-purpose processing chip 110a process the first and second data sequentially.

[0441] In some alternative approaches, the segment length may also be referred to as the length of the sub-data, the length of the data segment, the length of the token, or other names, which are not limited in this application.

[0442] The following is an exemplary description of the implementation process of the model inference stage based on length segmentation.

[0443] First, an example is given to illustrate the method of segmenting the data corresponding to the input sequence.

[0444] In one alternative implementation, the general-purpose processing chip 110a can divide the data corresponding to the input sequence into multiple subsequences according to a pre-set segmentation length.

[0445] The preset segmentation length can be set during the model training phase or preset by the computing device 110; this embodiment does not limit this.

[0446] In some alternative methods, the different subsequences have the same length. The general-purpose processing chip 110a sets the sliding step size based on a pre-set segmentation length. Moving along the input sequence according to the sliding step size, the input sequence is segmented into multiple subsequences.

[0447] The sliding step size can be the segmentation length, or it can be different from the segmentation length. For example, if the segmentation length can evenly segment the data corresponding to the input sequence, the segmentation length is used as the sliding step size. If the segmentation length cannot evenly segment the data corresponding to the input sequence, the segmentation length is used as the base length of the sliding step size, and the sliding step size is adjusted so that the adjusted sliding step size can evenly segment the data corresponding to the input sequence. This adjustment can involve decreasing or increasing the sliding step size.

[0448] In other alternative methods, the different subsequences have different lengths. The general-purpose processing chip 110a can use a pre-set segmentation length as a base length to randomly segment the data corresponding to the input sequence, resulting in multiple subsequences of different lengths.

[0449] The two optional methods described above are merely different implementations of the general-purpose processing chip 110a segmenting the data corresponding to the input sequence with a fixed segmentation length. In other embodiments, different models have different numbers of processing layers. If the input sequences of different models are segmented with the same segmentation length, it may cause a mismatch between the number of subsequences and the number of processing layers in the model, affecting the final inference result. Furthermore, different inference requests may correspond to data with input sequences of different lengths. If the same segmentation length is used to segment the data corresponding to input sequences of different lengths, the data corresponding to the longer input sequence will be segmented into more subsequences, increasing the number of inferences and intermediate results, thereby increasing the consumption of computing and memory resources. Therefore, to reduce the consumption of computing and memory resources, the general-purpose processing chip 110a can determine the segmentation length based on the length of the data corresponding to the input sequence and the number of processing layers in the model. In this way, by determining the degree of fit between the segmentation length and the length of the data corresponding to the input sequence and the number of processing layers in the model, the rationality of the segmentation length is ensured, thereby reducing the consumption of computing and memory resources.

[0450] The following example illustrates how to determine the segmentation length.

[0451] In the first optional implementation, the general-purpose processing chip 110a queries the first mapping relationship based on the number of processing layers in the model and the length of the data corresponding to the input sequence to obtain the segmentation length.

[0452] The first mapping relationship is used to indicate the correspondence between different segmentation lengths and the number of corresponding processing layers, as well as the length of the data corresponding to the input sequence.

[0453] In the second optional implementation, the general-purpose processing chip 110a obtains the segmentation coefficients based on the number of processing layers in the model and the length of the data corresponding to the input sequence. The segmentation length is then obtained by querying the second mapping relationship based on the segmentation coefficients.

[0454] The second mapping relationship is used to indicate the correspondence between the segmentation coefficient and the corresponding segmentation length.

[0455] In some alternative approaches, the number of processing layers can be the sum of the lengths of the data corresponding to the input sequence, or a weighted sum, or a weighted average, as the splitting coefficient.

[0456] In the third optional implementation, the general-purpose processing chip 110a inputs the number of processing layers in the model and the length of the data corresponding to the input sequence into the length prediction model to obtain the segmentation length.

[0457] The length prediction model can be a neural network-based model. Alternatively, it can be a machine learning-based model; this application does not limit the specific form of the length prediction model. Furthermore, the length prediction model can also be called a length predictor, a segmentation length predictor, or other names; this application does not limit these names.

[0458] The above three optional implementation methods are merely different ways for the general-purpose processing chip 110a to determine the segmentation length based on the number of processing layers in the model and the length of the data corresponding to the input sequence. In other embodiments, the general-purpose processing chip 110a may also use other implementation methods to determine the segmentation length. This application does not limit this.

[0459] The following example illustrates how to determine the segmentation length using a length prediction model.

[0460] First, the construction method of the length prediction model will be explained.

[0461] In the first alternative implementation, the length prediction model can be built by training device 120, and after the length prediction model is built, it can be deployed to computing device 110.

[0462] In the first example, the training device 120 acquires multiple sample input lengths, multiple samples of the number of processing layers, sample hardware configuration parameters, and an initial model. The training device 120 inputs the multiple sample input lengths and multiple samples of the number of processing layers into the initial model, obtaining multiple test lengths output by the initial model. Using different test lengths as examples, the training device 120 simulates the prediction inference time of the initial model after the sample data is segmented according to different test lengths using devices with different sample hardware configuration parameters. If the difference between the prediction inference time and the time threshold is less than a difference threshold, the initial model is used as the length prediction model. If the difference between the prediction inference time and the time threshold is greater than a difference threshold, the parameters of the initial model are adjusted according to the difference between the prediction inference time and the time threshold, and prediction is performed again based on the adjusted initial model until the difference between the prediction inference time corresponding to the prediction length output by the latest adjusted initial model and the time threshold is less than the difference threshold; at this point, the latest adjusted initial model is used as the length prediction model.

[0463] The parameters of the initial model include its weights, biases, etc. The initial model can be based on a neural network or a machine learning model. This application does not limit the specific form of the initial model.

[0464] The sample hardware configuration parameters include: chip computing power, memory access bandwidth, and memory space.

[0465] In the second example, training device 120 simulates multiple different test lengths and, after simulating the hardware configuration parameters of the sample data according to different test lengths, the prediction inference time of the initial model. The prediction length with the shortest prediction inference time is fitted with multiple sample input lengths and multiple processing layer numbers of samples to establish a fitting relationship between the prediction length and multiple sample input lengths and multiple processing layer numbers of samples. This fitting relationship is used as the length prediction model.

[0466] The above two examples are merely different implementations of the training device 120 in constructing the length prediction model. In other embodiments, the training device 120 may also use other implementations to construct the length prediction model, and this application does not limit this.

[0467] In the second alternative implementation, the length prediction model can be constructed by computing device 110.

[0468] For example, computing device 110 acquires an initial length prediction model and predicts the predicted segmentation length for different input lengths based on the initial length prediction model. Then, according to the hardware configuration parameters of computing device 110, it simulates the prediction inference time of segmenting test data with different predicted segmentation lengths. If the difference between the predicted inference time and a time threshold is less than or equal to a preset difference threshold, the initial length prediction model is used as the length prediction model. If the difference between the predicted inference time and the time threshold is greater than the preset difference threshold, the parameters in the initial length prediction model are adjusted to obtain the final length prediction model.

[0469] The initial length prediction model can be constructed by the training device 120 with reference to the first optional implementation method described above, or the initial length prediction model can be constructed by other devices. This application embodiment does not limit the source of the initial length prediction model. The parameters in the initial length prediction model include the weights and biases of the initial length prediction model.

[0470] The two optional implementation methods mentioned above are merely different ways to construct the length prediction model. In other embodiments, there may be other ways to construct the length prediction model, and this application does not limit these methods.

[0471] Second, an exemplary description is given of how to determine the segmentation length using a length prediction model.

[0472] In the first optional implementation, the general-purpose processing chip 110a can input the number of processing layers in the model and the length of the data corresponding to the input sequence into the length prediction model to obtain the prediction probabilities of multiple candidate segment lengths. The candidate segment length with the highest prediction probability is determined as the segment length of the input sequence.

[0473] In the second optional implementation, the general-purpose processing chip 110a can input the number of processing layers in the model and the length of the data corresponding to the input sequence into the length prediction model to obtain multiple different candidate segment lengths and the prediction inference time corresponding to different candidate segment lengths. The candidate segment length with the shortest inference time is determined as the segment length of the data corresponding to the input sequence.

[0474] In a third optional implementation, the general-purpose processing chip 110a can input the number of processing layers in the model and the length of the data corresponding to the input sequence into the length prediction model to obtain the confidence scores of multiple candidate segment lengths. The candidate segment length with the highest confidence score is determined as the segment length of the data corresponding to the input sequence. The confidence score indicates the degree of acceptance of the candidate segment length.

[0475] The above three optional implementation methods are only different ways of using the length prediction model to determine the segmentation length. In other embodiments, there may be other ways of using the length prediction model to determine the segmentation length, and this application does not limit these methods.

[0476] The above primarily uses the number of processing layers in the model and the length of the data corresponding to the input sequence as inputs to determine the segmentation length of the data corresponding to the input sequence. In other embodiments, due to the influence of the computing power of the chip in the computing device 110, the model inference time may increase if the computing power of the execution chip is not matched with the length of the subsequence. For example, for an execution chip with high computing power, using a smaller segmentation length will increase the number of subsequences that the execution chip needs to process, thereby increasing the model inference time. Conversely, for an execution chip with low computing power, using a longer segmentation length will increase the time that the execution chip takes to process subsequences, thus increasing the model inference time. Therefore, to ensure the compatibility between the segmentation length and the chip's computing power and improve the rationality of the segmentation length, the general-purpose processing chip 110a can determine the segmentation length based on the chip's computing power, the number of processing layers in the model, and the length of the data corresponding to the input sequence during the process of determining the segmentation length.

[0477] The following provides an example of how to determine the segmentation length based on the chip's computing power, the number of processing layers in the model, and the length of the data corresponding to the input sequence.

[0478] In the first optional implementation, the general-purpose processing chip 110a queries the third mapping relationship based on the chip's computing power, the number of processing layers in the model, and the length of the data corresponding to the input sequence to obtain the segmentation length.

[0479] In some alternative approaches, a third mapping is used to indicate the correspondence between the segmentation length and the chip's computing power, the number of processing layers in the model, and the length of the data corresponding to the input sequence.

[0480] In some alternative approaches, a third mapping is used to indicate the correspondence between the segmentation length and the segmentation coefficient. This segmentation coefficient is determined by the chip's computing power, the number of processing layers in the model, and the length of the data corresponding to the input sequence. For example, the segmentation coefficient can be determined by the sum, weighted sum, weighted average, or arithmetic mean of the chip's computing power, the number of processing layers in the model, and the length of the data corresponding to the input sequence.

[0481] In the first optional scenario, the computing power of a chip can be a statistical characteristic of the computing power of all chips in computing device 110. This statistical characteristic may include, but is not limited to, the total computing power of all chips in computing device 110, the arithmetic mean, the weighted average, the median, the maximum value, and the minimum value.

[0482] For example, taking the computing device 110 provided in Figure 4 above as an example, the computing power of the chip can be the statistical characteristics of the computing power of the first acceleration chip 111, the second acceleration chip 112, the third acceleration chip 113, the fourth acceleration chip 114 and the general-purpose processing chip 110a.

[0483] In the second alternative scenario, the acceleration chip in computing device 110 executes most of the operators in the model, so the computing power of the chip can be a statistical characteristic of the computing power of the acceleration chip in computing device 110.

[0484] For example, taking the computing device 110 provided in Figure 4 above as an example, the computing power of the chip can be the statistical characteristics of the computing power of the first acceleration chip 111, the second acceleration chip 112, the third acceleration chip 113, and the fourth acceleration chip 114.

[0485] In the third optional scenario, the operators in the model are mainly executed by the NPU, so the computing power of the chip can be determined based on the computing power of the NPU-type acceleration chip in the computing device 110.

[0486] For example, if the computing device 110 includes an NPU, the computing power of the NPU is determined as the computing power of the chip.

[0487] For example, if the computing device 110 includes multiple NPUs, the arithmetic mean, weighted average, maximum or minimum value of the computing power of the multiple NPUs is determined as the computing power of the chip.

[0488] In the fourth optional scenario, the computing power of the chip can be a statistical characteristic of the computing power of the execution chip. For example, the computing power statistical characteristics of the first and second execution chips mentioned above can be determined as the computing power of the chip.

[0489] The above four scenarios are merely optional methods for determining the computing power of a chip. In other embodiments, other optional methods may also be used to determine the computing power of a chip, and this application does not limit these methods.

[0490] In the second alternative implementation, the general-purpose processing chip 110a uses the chip's computing power, the number of processing layers in the model, and the length of the data corresponding to the input sequence to input the length prediction model and obtain the segmentation length.

[0491] In one optional scenario, as can be seen from the method of creating the length prediction model described above, the length prediction model records the chip's computing power. However, due to the usage time of the computing device 110 or adjustments to the chip's frequency, there may be a difference between the actual computing power of the chip in the computing device 111 and the baseline computing power of the chip corresponding to the length prediction model. To ensure the compatibility between the segmentation length and the chip's actual computing power, the general-purpose processing chip 110a compares the actual computing power of the chip in the computing device 110 with the baseline computing power of the chip corresponding to the length prediction model. If the chip's actual computing power does not match its baseline computing power, the general-purpose processing chip 110a inputs the chip's computing power, the number of processing layers in the model, and the length of the data corresponding to the input sequence into the length prediction model to obtain the segmentation length. If the chip's actual computing power matches its baseline computing power, the general-purpose processing chip 110a inputs the number of processing layers in the model and the length of the input sequence into the length prediction model to obtain the segmentation length. The baseline computing power can refer to the computing power used in the creation phase of the length prediction model, or it can be the initial actual computing power of the chip in the computing device 110 when the length prediction model is deployed to the computing device 110.

[0492] In some alternative approaches, since the chip's frequency can indicate its computing power, the general-purpose processing chip 110a can compare the chip's actual frequency with its reference frequency. If the frequency difference between the chip's actual frequency and its reference frequency is greater than a preset difference threshold, it is determined that the chip's actual computing power does not match its reference computing power. If the frequency difference between the chip's actual frequency and its reference frequency is less than the preset difference threshold, it is determined that the chip's actual computing power matches its reference computing power.

[0493] For example, taking the computing power of the accelerator chip as an example and the preset difference threshold as 0.5Hz, when the accelerator chip is the first accelerator chip 111 and the base frequency of the first accelerator chip 111 is 2.90GHz, and the actual frequency of the first accelerator chip 111 is 3.70GHz, the general-purpose processing chip 110a inputs the computing power of the first accelerator chip 111, the number of processing layers in the model, and the length of the data corresponding to the input sequence into the length prediction model to obtain the segmentation length. When the actual frequency of the first accelerator chip 111 is 2.78GHz, the general-purpose processing chip 110a inputs the number of processing layers in the model and the length of the data corresponding to the input sequence into the length prediction model to obtain the segmentation length.

[0494] In the third optional implementation, to ensure the compatibility between the segmentation length and the chip's actual computing power, the general-purpose processing chip 110a can determine the initial segmentation length based on the number of processing layers in the model and the length of the data corresponding to the input sequence. If the chip's actual computing power does not match its baseline computing power, the general-purpose processing chip 110a adjusts the initial segmentation length based on the difference between the two. If the chip's actual computing power matches its baseline computing power, the general-purpose processing chip 110a determines the initial segmentation length as the segmentation length of the input sequence.

[0495] The "difference between the chip's actual computing power and its baseline computing power" can be either the difference between the chip's actual computing power and its baseline computing power, or the ratio between the chip's actual computing power and its baseline computing power.

[0496] In the first alternative approach, when the difference is a differential value, the general-purpose processing chip 110a determines the length change based on the difference between the chip's actual computing power and the chip's reference computing power. The initial segmentation length is then adjusted according to the length change to obtain the segmentation length.

[0497] In the second alternative approach, where the difference is a ratio, the general-purpose processing chip 110a obtains an adjustment ratio based on the ratio between the chip's actual computing power and its baseline computing power. The initial slicing length is then adjusted according to this adjustment ratio to obtain the slicing length.

[0498] Taking the computing power of the accelerator chip as an example, after determining the initial slicing length, the general-purpose processing chip 110a obtains the reference frequency and the actual frequency of the accelerator chip. If the frequency difference between the actual frequency and the reference frequency is greater than a preset difference threshold, the general-purpose processing chip 110a increases or decreases the slicing length according to the frequency ratio between the actual frequency and the reference frequency of the accelerator chip.

[0499] The above three optional implementation methods are only different implementation methods for determining the segmentation length based on the chip's computing power, the number of processing layers in the model, and the length of the data corresponding to the input sequence. In other embodiments, the general-purpose processing chip 110a may also use other implementation methods to determine the segmentation length, which is not limited in this application embodiment.

[0500] Secondly, based on the above-mentioned implementation of the data segmentation method corresponding to the input sequence, the model inference stage under the sequence segmentation method is illustrated by example.

[0501] As shown in Figure 12, Figure 12 is a schematic diagram of the second process of the model inference stage provided in this application. The processing of the model inference stage includes S821B to S823B.

[0502] The S821B general-purpose processing chip 110a obtains the segmentation length of the data corresponding to the input sequence.

[0503] In an alternative implementation, the general-purpose processing chip 110a can refer to the above-described embodiment for determining the segmentation length to obtain the segmentation length of the data corresponding to the input sequence, which will not be elaborated further.

[0504] The S822B general-purpose processing chip 110a uses a reference segmentation length to segment the data corresponding to the input sequence into multiple subsequences.

[0505] In some alternative implementations, the general-purpose processing chip 110a can refer to the above-described embodiment of segmenting the input sequence to segment the data corresponding to the input sequence into multiple sub-sequences, which will not be elaborated upon in this application.

[0506] As shown in Figure 12, the general-purpose processing chip 110a uses the reference segmentation length to segment the data corresponding to the input sequence into four subsequences: subsequence 1, subsequence 2, subsequence 3 and subsequence 4.

[0507] The S823B general-purpose processing chip 110a inputs each of the multiple subsequences into the model, controls the accelerator chip to execute the first operator of the multiple data processing layers sequentially, and controls the general-purpose processing chip 110a to execute the second operator of the multiple data processing layers sequentially, so as to process each of the multiple subsequences in a pipelined parallel manner.

[0508] In an alternative implementation, the general-purpose processing chip 110a can refer to the embodiment provided in FIG10 above, scheduling multiple operators of each data processing layer in the model to the corresponding execution chip. During the inference process of each data processing layer, the general-purpose processing chip 110a and the acceleration chip process each sub-sequence in a pipeline manner. The pipeline processing of each sub-sequence can be referred to the embodiment provided in FIG6A above, and will not be described in detail here.

[0509] As shown in Figure 12, the general-purpose processing chip 110a inputs subsequence j into the first acceleration chip 111 in the order of subsequence 1 to subsequence 4. The first acceleration chip 111, the second acceleration chip 112, and the general-purpose processing chip 110a execute multiple data processing layers in the order of processing layer 1 → processing layer 2 → ... → processing layer L to process subsequence j. Within each data processing layer (e.g., processing layer 1), the first acceleration chip 111, the second acceleration chip 112, and the general-purpose processing chip 110a execute in the order of first acceleration chip 111 → second acceleration chip 112 → general-purpose processing chip 110a → first acceleration chip 111 to process subsequence j. After the first acceleration chip 111 processes subsequence j, it continues to process subsequence j+1, so that the pipeline consisting of first acceleration chip 111 → second acceleration chip 112 → general-purpose processing chip 110a → first acceleration chip 111 can process at least two subsequences in parallel. Here, j is a positive integer greater than 0 and less than 4.

[0510] Based on the embodiment shown in Figure 12, the general-purpose processing chip 110a segments the data corresponding to the input sequence to obtain multiple sub-sequences. Each sub-sequence is used as input to the model, enabling the model to perform inference on each sub-sequence separately. For each chip, this reduces the amount of data processed per chip and shortens the processing time of each chip, thereby shortening the overall inference time of the model and further improving the inference efficiency of the model.

[0511] In one alternative implementation, the execution order of the operator scheduling stage and the model inference stage described above is not limited in this embodiment.

[0512] In the first optional approach, the general-purpose processing chip 110a executes in the order of operator scheduling phase → model inference phase. For example, the general-purpose processing chip 110a executes in the order of receiving inference request → determining the partitioning strategy of multiple operators → scheduling operators to the corresponding execution chips → using the data corresponding to the input sequence as the input of the first execution chip. Another example is that the general-purpose processing chip 110a executes in the order of determining the partitioning strategy of multiple operators → receiving inference request → scheduling operators to the corresponding execution chips according to the partitioning strategy of multiple operators → using the data corresponding to the input sequence as the input of the first execution chip.

[0513] In the second alternative approach, the general-purpose processing chip 110a executes in the order of model inference phase → operator scheduling phase. For example, the general-purpose processing chip 110a receives an inference request → uses the data corresponding to the input sequence as the model input → determines the partitioning strategy for multiple operators → schedules the operators to the corresponding execution chips for execution in that order. Another example is that the general-purpose processing chip 110a executes in the following order: determines the partitioning strategy for multiple operators → receives an inference request → uses the data corresponding to the input sequence as the model input → schedules the operators to the corresponding execution chips according to the partitioning strategy for multiple operators.

[0514] The two optional methods described above are merely different implementations of executing the operator scheduling stage and the model inference stage in different orders. In other embodiments, other orders of execution for the operator scheduling stage and the model inference stage may also be used. This application does not limit this. For example, the general-purpose processing chip 110a executes the operator scheduling stage and the model inference stage alternately. Exemplarily, the general-purpose processing chip 110a receives an inference request → determines the segmentation length → segments the data corresponding to the input sequence according to the segmentation length → determines the partitioning strategy of multiple operators → schedules the operators to the corresponding execution chips → inputs each subsequence into multiple execution chips → controls multiple execution chips to process the sequential execution of each subsequence in parallel in a pipeline manner.

[0515] The following example, using the CPU, GPU, and NPU as execution chips and the inference process of the first data processing layer, illustrates the implementation of the alternating execution of the operator scheduling stage and the model inference stage. The network structure of this first data processing layer is shown in Figure 5 above. Figure 13 is a flowchart illustrating the inference method of a model provided in this application. The inference method of this model includes stages ① to ⑤, wherein:

[0516] Phase ①: The CPU responds to the inference request and determines the segmentation length of the data corresponding to the input sequence as X based on the length of the data corresponding to the input sequence, the number of data processing layers in the model, and the hardware configuration parameters of the NPU.

[0517] In one alternative implementation, the CPU can refer to the above implementation example of determining the segmentation length to obtain the segmentation length X of the data corresponding to the input sequence, which will not be elaborated upon in this application.

[0518] Phase 2: The CPU divides the data S corresponding to the input sequence into multiple subsequences: S1, S2, S3 and S4 according to the segmentation length X.

[0519] In one optional implementation, the CPU, referring to the method described above for segmenting the data corresponding to the input sequence, divides the data S corresponding to the input sequence into multiple subsequences. This embodiment of the application will not elaborate further on this.

[0520] Phase 3: The CPU determines the partitioning strategy for multiple operators in the first data processing layer.

[0521] In one alternative implementation, the CPU can refer to the above-described operator scheduling phase embodiment to determine the partitioning strategy of multiple operators in the first data processing layer. This application embodiment does not limit this.

[0522] Phase 4: The CPU schedules multiple operators in the first data processing layer to the CPU, GPU, and NPU according to the multi-operator partitioning strategy.

[0523] As shown in Figure 13, the feature extraction operator, the first normalization operator, and the first linear operator are scheduled to the NPU, the self-attention operator is scheduled to the GPU, and the second linear operator and the second normalization operator are scheduled to the CPU.

[0524] Phase 5: The CPU inputs S1, S2, S3 and S4 into the NPU, and the NPU, CPU and GPU process each subsequence in parallel in a pipeline manner.

[0525] As shown in Figure 13, the NPU processes each subsequence in the order of subsequence S1 → subsequence S2 → subsequence S3 → subsequence S4. For subsequence Si, it processes subsequence Si in the order of execution of the first normalization operator → the first linear operator, and obtains the first processing result (S131) ​​of subsequence Si.

[0526] After processing the subsequence Si, the NPU passes the first processing result of the subsequence Si to the GPU (S132).

[0527] The GPU processes the first processing result of subsequence Si sequentially using the self-attention operator in the order of subsequence S1→subsequence S2→subsequence S3→subsequence S4 to obtain the second processing result (S133) of subsequence Si.

[0528] After processing a subsequence Si, the GPU passes the second processing result of the subsequence Si to the CPU (S134).

[0529] The CPU processes the second processing result of subsequence S1 in the order of subsequence S2, subsequence S3, and subsequence S4. Then, for the second processing result of subsequence S1, it processes the second processing result of subsequence S1 in the order of execution of the second linear operator, and obtains the third processing result of subsequence S1 (S135).

[0530] After processing a subsequence, the CPU will pass the third processing result of the subsequence Si to the NPU (S136).

[0531] The NPU processes the third processing result of the subsequence Si in the order of subsequence S1→subsequence S2→subsequence S3→subsequence S4 through the feature extraction operator, and obtains the output result of the first data processing layer for the subsequence Si (S137).

[0532] In an alternative implementation, similar to the embodiment shown in Figure 6A above, data transfer between the NPU, CPU, and GPU is achieved through shared memory. After obtaining the first processing result of subsequence S4, the NPU processes the third processing result of subsequence S1 using feature extraction operators. After the NPU obtains the output result of the first data processing layer for subsequence S1, the NPU processes subsequence S1 according to the execution order of the first normalization operator → the first linear operator, thus initiating the processing of the next processing layer.

[0533] Based on the embodiment shown in Figure 13, during the model inference process, the computing device 110 utilizes the chips within the computing device 110 to schedule multiple operators of a processing layer onto different chips using a heterogeneous computing power scheduling approach. By running the processing layer in a pipeline manner through multiple chips, the load on a single chip is reduced. Furthermore, after completing a single processing iteration, each chip (NPU, CPU, or GPU) continues to execute the next processing iteration, reducing the idle time of the chip (NPU, CPU, or GPU) and fully utilizing the computing power of the chip (NPU, CPU, or GPU), thereby improving the model's inference efficiency. Additionally, the appropriate segmentation length is configured using NPU and hardware configuration parameters, the length of the data corresponding to the input sequence, and the number of processing layers in the model. The input sequence is segmented using this segmentation length, reducing the number of operations per chip in each inference process while fully utilizing the computing power of the chips within the computing device 110, thus accelerating model inference.

[0534] For example, taking the operator scheduling method provided in Figure 13 above as an example, with the length of the input sequence corresponding to the data being 1024 bits, the segmentation length being 7 bits, the number of floating-point operations for the NPU being 9, the number of floating-point operations for the GPU being 6, the number of floating-point operations for the CPU being 9, and the bandwidth of the shared memory being 51.2 GB / s, for the same model, when executing the same inference request, using the inference method of the model provided in this application, the data transmission latency is 0.00008 seconds, the computation latency is 1.05 seconds, and the total latency of model inference is approximately 1.05 seconds. However, when only the NPU is used, the computation latency is 1.2 seconds. When only the CPU is used, the computation latency is 1.24 seconds. When only the GPU is used, the computation latency is 1.8 seconds.

[0535] It is evident that the inference method using the model provided in this application can reduce communication overhead, making data transmission latency nearly zero, without increasing the total inference latency of the model. Furthermore, compared to using only a single chip for model inference, this application coordinates multiple chips to achieve heterogeneous scheduling of computing power, adopts pipelined parallelism to reduce chip idle time, and fully utilizes the computing power of multiple chips, thereby reducing the computational latency of model inference.

[0536] The above embodiments are illustrated using model inference as an example, but the technical solutions provided in this application can also be used in model training scenarios or the model training stage. For example, in a model training scenario, scheduling different operators of each data processing layer onto different chips or chipsets, and executing multiple data processing layers sequentially through different chips or chipsets, is beneficial for improving training efficiency. Furthermore, during training, the sample sequence is segmented, and multiple sample subsequences are used to reduce the sequence length of each training input, thereby increasing the training rate. In some optional cases, not only can the above-mentioned computing device 110 execute the model training process, but the above-mentioned training device 120 or other devices can also execute the model training process; this application does not limit this.

[0537] To achieve the functions described in the above embodiments, the computing device 110 includes hardware components and / or software modules for executing each function. Those skilled in the art will readily recognize that, based on the units and method steps of the examples described in conjunction with the embodiments disclosed in this application, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application scenario and design constraints of the technical solution.

[0538] The computing device, acceleration chip, and general-purpose processing chip provided in this application can be referred to the description of the foregoing embodiments, and will not be repeated here.

[0539] In an optional implementation, this application also provides a model inference apparatus. This model inference apparatus can implement the above-described method embodiments through software units, such as the model inference apparatus being applicable to the aforementioned computing device 110. In an optional case, the model inference apparatus may include an acquisition module and a scheduling module.

[0540] The acquisition module receives inference requests. The scheduling module, in response to the inference request, schedules the model's first operator to the acceleration chip, schedules the model's second operator to the general-purpose processing chip, controls the acceleration chip to sequentially execute the first operators of multiple data processing layers, and controls the general-purpose processing chip to sequentially execute the second operators of multiple data processing layers.

[0541] It should be understood that the inference apparatus of the model provided in the embodiments of this application can be implemented by a processing chip. The inference apparatus of the model according to the embodiments of this application can correspond to the execution of the methods described in the embodiments of this application, and the above and other operations and / or functions of each unit and module in the model inference apparatus are respectively to implement the corresponding processes of the various methods in the foregoing figures. For the sake of brevity, they will not be described again here.

[0542] For example, this processing chip can be used to implement the functions of the general-purpose processing chip in the above method embodiments, and thus can also achieve the beneficial effects of the above method embodiments. In this embodiment, the processing chip can be the general-purpose processing chip 110a shown in FIG3, or it can be a general-purpose processing chip or acceleration chip in subsequent embodiments, or a module (such as a chip) applied to a general-purpose processing chip.

[0543] As shown in Figure 14, which is a schematic diagram of the processing chip provided in this application, the processing chip 1400 may include a processor 1420. Optionally, the processing chip 1400 may also include a memory 1430 and / or a transceiver 1410. The processor 1420 is coupled to the memory 1430 and the transceiver 1410, for example, through a communication bus. This communication bus may include, but is not limited to, a PCIe bus, or an extended industry standard architecture (EISA) bus, a unified bus (Ubus or UB), a compute express link (CXL), a cache coherent interconnect for accelerators (CCIX), etc.

[0544] The following section, with reference to Figure 14, provides a detailed description of each component of the processing chip 1400:

[0545] The processor 1420 is the control center of the processing chip 1400. It can be a single processor or a collective term for multiple processing elements. For example, the processor 1420 can be one or more CPUs, an ASIC, or one or more integrated circuits configured to implement the embodiments of this application, such as one or more digital signal processors (DSPs), or one or more field-programmable gate arrays (FPGAs).

[0546] Optionally, the processor 1420 can perform various functions of the processing chip 1400 by running or executing software programs stored in the memory 1430 and by calling data stored in the memory 1430. In a specific implementation, as one embodiment, the processor 1420 may include one or more CPUs.

[0547] Optionally, the processing chip 1400 may also include multiple processors. Each of these processors may be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). Here, a processor may refer to one or more devices, circuits, and / or processing cores used to process data (e.g., computer program instructions).

[0548] The memory 1430 stores the software program for executing the operator scheduling and the software program for input sequence segmentation in this application, and is controlled by the processor 1420 for execution. Specific implementation details can be found in the operator scheduling stage and model inference stage of the above method embodiments. Further details are omitted here. For example, the memory 1430 can be a ROM or other type of static storage device capable of storing static information and instructions, RAM or other type of dynamic storage device capable of storing information and instructions, or it can be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), a magnetic disk storage medium, or other magnetic storage device, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures that can be accessed by a computer, but is not limited thereto. The memory 1430 can be integrated with the processor 1420 or exist independently and is coupled to the processor 1420 through the interface circuit of the processing chip 1400 (not shown in FIG14). This application embodiment does not specifically limit this.

[0549] Transceiver 1410 is used for communication with other devices. For example, if processing chip 1400 is a user terminal (such as a client) or application server, transceiver 1410 can be used to communicate with an accelerator chip or with another processing chip. As another example, if processing chip 1400 is a multi-core chip, transceiver 1410 can be used to communicate with another multi-core chip.

[0550] Optionally, transceiver 1410 may include a receiver and a transmitter (not shown separately in FIG14). The receiver is used to implement the receiving function, and the transmitter is used to implement the transmitting function. Optionally, transceiver 1410 may be integrated with processor 1420, or it may exist independently and be coupled to processor 1420 through the interface circuit of processing chip 1400 (not shown in FIG14). This embodiment of the application does not specifically limit this aspect.

[0551] In this embodiment, transceiver 1410 is used to receive inference requests. In response to the inference request, processor 1420 schedules the first operator of the model to the acceleration chip, schedules the second operator of the model to the general-purpose processing chip, controls the acceleration chip to sequentially execute the first operators of multiple data processing layers, and controls the general-purpose processing chip to sequentially execute the second operators of multiple data processing layers. This processor 1420 can be used to implement the functions and corresponding beneficial effects of the general-purpose processing chip 110a in the aforementioned embodiment, which will not be elaborated upon here.

[0552] The method steps in this embodiment can be implemented in hardware or by a processor executing software instructions. The software instructions can consist of corresponding software modules, which can be stored in RAM, flash memory, ROM, PROM, EPROM, EEPROM, registers, hard disk, portable hard disk, CD-ROM, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor, enabling the processor to read information from and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and storage medium can reside in an ASIC. Alternatively, the ASIC can reside in a computing device. Of course, the processor and storage medium can also exist as discrete components in a network device or terminal device.

[0553] This application also provides a terminal device including the aforementioned processing chip 1400, which further includes one or more acceleration chips and a memory. The memory stores a model and a set of computer instructions. The processing chip 1400 invokes the set of computer instructions, schedules multiple operators in the model to the processing chip 1400 and the acceleration chip respectively, receives inference requests, and coordinates with the acceleration chip and memory to process the inference requests to obtain the inference results.

[0554] In some alternative embodiments, the structure of the terminal device can refer to the structure of the computing device 110 shown in Figure 4 above, which will not be described in detail here.

[0555] This application also provides a computer program product containing instructions. The computer program product may be a software or program product containing instructions, capable of running on a computing device or stored on any usable medium. When the computer program product is run on at least one computing device, it causes the at least one computing device to execute the reasoning method of the above-described model.

[0556] For example, when a computer program product is run on at least one computing device, the at least one computing device performs the reasoning method of the model shown in Figure 8.

[0557] For example, when a computer program product is run on at least one computing device, the at least one computing device performs the reasoning method of the model shown in Figure 13.

[0558] This application also provides a computer-readable storage medium. All or part of the processes in the above method embodiments can be implemented by a computer program instructing related hardware. This program can be stored in the computer-readable storage medium, and when executed, it can include the processes of the above method embodiments. The computer-readable storage medium can be a terminal of any of the foregoing embodiments, such as an internal storage unit including a data transmission end and / or a data receiving end, like a hard disk or memory of the terminal. The computer-readable storage medium can also be an external storage device of the terminal, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the terminal. Further, the computer-readable storage medium can include both the internal storage unit and the external storage device of the terminal. The computer-readable storage medium is used to store the computer program and other programs and data required by the terminal. The computer-readable storage medium can also be used to temporarily store data that has been output or will be output.

[0559] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. A computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed on a computer, the processes or functions described in the embodiments of this application are performed entirely or partially. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a network device, a user equipment, or other programmable device. The computer program or instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another. For example, the computer program or instructions can be transferred from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium, such as a floppy disk, hard disk, or magnetic tape; it can also be an optical medium, such as a digital video disc (DVD); or it can be a semiconductor medium, such as a solid-state drive (SSD).

[0560] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A model inference method, comprising: The method is applied to a terminal device, which is equipped with a general-purpose processing chip and an acceleration chip; the method includes: Receive an inference request; the inference request is used to request the terminal device to invoke a model for inference, the model includes multiple data processing layers executed sequentially, each of the data processing layers includes a first operator and a second operator executed sequentially; The acceleration chip is controlled to sequentially execute the first operators of the plurality of data processing layers, and the general-purpose processing chip is controlled to sequentially execute the second operators of the plurality of data processing layers, wherein the input of the second operator of the first data processing layer in the plurality of data processing layers is associated with the output of the first operator of the first data processing layer.

2. The method of claim 1, wherein, The method further includes: A partitioning strategy is created; the partitioning strategy is used to indicate the execution chip corresponding to each of the multiple operators in the data processing layer when the model is inference, the execution chip including the general processing chip or the acceleration chip, wherein the multiple operators include the first operator and the second operator.

3. The method of claim 2, wherein, The partitioning strategy is related to the computing power parameters and resource utilization of the general-purpose processing chip or the acceleration chip. The computing power parameters include one or more of the following: processing accuracy, network bandwidth, or floating-point operations.

4. The method according to any one of claims 1 to 3, characterized in that, The terminal device is also equipped with shared memory, which supports access by the general-purpose processing chip and the acceleration chip.

5. The method according to any one of claims 1 to 4, characterized in that, The data processing layer is a transformer layer, which includes a first normalization operator, a first linear operator, a self-attention operator, a second linear operator, and a second normalization operator. The first operator includes at least one of the following operators: the first normalization operator, the first linear operator, or the self-attention operator; the second operator includes: the second linear operator or the second normalization operator.

6. The method according to claim 5, characterized in that, The acceleration chip includes a graphics processing chip (GPU) and a neural network processing chip (NPU); the execution chip for the first normalization operator or the first linear operator is the NPU; and the execution chip for the self-attention operator is the GPU.

7. The method according to any one of claims 1 to 6, characterized in that, The acceleration chip sequentially executes the first operator of the plurality of data processing layers, including: For the first data processing layer among the plurality of data processing layers, the first operator of the first data processing layer sequentially processes the first data and the second data to obtain the first result corresponding to the first data and the second result corresponding to the second data, respectively; wherein, the first data and the second data are data related to the inference request.

8. The method according to claim 7, characterized in that, The general-purpose processing chip sequentially executes the second operator of the plurality of data processing layers, including: The second operator of the first data processing layer processes the first result and the second result sequentially to obtain the third result corresponding to the first result and the fourth result corresponding to the second result, respectively.

9. The method according to claim 7, characterized in that, The first data and the second data have the same length, and the length of the first data or the second data is related to the computing power of the general-purpose processing chip or the computing power of the acceleration chip.

10. The method according to any one of claims 1 to 9, characterized in that, Controlling the acceleration chip to sequentially execute the first operators of the plurality of data processing layers includes: controlling the acceleration chip to sequentially execute the first operator of the first data processing layer and the first operator of the second data processing layer; or, Controlling the general-purpose processing chip to sequentially execute the second operators of the plurality of data processing layers includes: controlling the general-purpose processing chip to sequentially execute the second operators of the first data processing layer and the second operators of the second data processing layer; The first data processing layer and the second data processing layer are adjacent data processing layers among the plurality of data processing layers.

11. A terminal device, characterized in that, The terminal device includes a general-purpose processing chip and an acceleration chip; The terminal device is configured to: receive an inference request; the inference request is configured to request the terminal device to invoke a model for inference, the model comprising multiple data processing layers executed sequentially, each of the data processing layers comprising a first operator and a second operator executed sequentially; The acceleration chip is used to: sequentially execute the first operator of the plurality of data processing layers; The general-purpose processing chip is also used to: sequentially execute the second operators of the plurality of data processing layers; the input of the second operator of the first data processing layer in the plurality of data processing layers is associated with the output of the first operator of the first data processing layer.

12. The terminal device according to claim 11, characterized in that, The general-purpose processing chip is further used to: create a partitioning strategy; the partitioning strategy is used to indicate the execution chip corresponding to each of the multiple operators in the data processing layer when the model is inference, the execution chip including the general-purpose processing chip or the acceleration chip, wherein the multiple operators include the first operator and the second operator; the partitioning strategy is related to the computing power parameters and resource utilization of the general-purpose processing chip or the acceleration chip.

13. The terminal device according to claim 11 or 12, characterized in that, The terminal device is also equipped with shared memory, which supports access by the general-purpose processing chip and the acceleration chip.

14. The terminal device according to any one of claims 11 to 13, characterized in that, The acceleration chip is specifically used for: For the first data processing layer among the plurality of data processing layers, the first data and the second data are processed sequentially using the first operator of the first data processing layer to obtain the first result corresponding to the first data and the second result corresponding to the second data, respectively; wherein, the first data and the second data are data related to the inference request; The general-purpose processing chip is specifically used for: The first result and the second result are processed sequentially using the second operator of the first data processing layer to obtain the third result corresponding to the first result and the fourth result corresponding to the second result, respectively.

15. The terminal device according to claim 14, characterized in that, The first data and the second data have the same length, and the length of the first data or the second data is related to the computing power of the general-purpose processing chip or the computing power of the acceleration chip.

16. The terminal device according to any one of claims 11 to 15, characterized in that, The acceleration chip is specifically used to: sequentially execute the first operator of the first data processing layer and the first operator of the second data processing layer; or... The general-purpose processing chip is specifically used to: sequentially execute the first operator of the first data processing layer and the first operator of the second data processing layer; The first data processing layer and the second data processing layer are adjacent data processing layers among the plurality of data processing layers.

17. The terminal device according to any one of claims 11 to 16, characterized in that, The inference request carries an input sequence; the general-purpose processing chip is also used to: send a first request to the server based on the input sequence, receive a feature sequence sent by the server based on the first request, and input the feature sequence into the acceleration chip.

18. The terminal device according to any one of claims 11 to 17, characterized in that, The data processing layer is a transformer layer, which includes a first normalization operator, a first linear operator, a self-attention operator, a second linear operator, and a second normalization operator. The first operator includes at least one of the following operators: the first normalization operator, the first linear operator, or the self-attention operator; the second operator includes: the second linear operator or the second normalization operator.

19. A computer program product, characterized in that, When the computer program product is run on a terminal device, the terminal device performs the method according to any one of claims 1 to 10.