Data processing method, data processor, electronic device, storage medium

By adopting a multi-execution group parallel strategy in neural network computing, using hardware computing unit grouping and calculation graph execution order, the problem of taking into account both model inference latency and throughput is solved, and efficient model inference operations are achieved.

CN118313458BActive Publication Date: 2025-07-08SHANGHAI BIREN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410338784.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-03-22
Publication Date
2025-07-08
Estimated Expiration
2044-03-22

AI Technical Summary

Technical Problem

The prior art is difficult to take into account the latency and throughput of model in neural network computing, especially when the batch size is small, and model parallelism strategies lead to waste of computing resources and low throughput.

Method used

Multiple execution groups are used to perform inference tasks in parallel. Each execution group includes multiple hardware computing units. Each execution group receives different input data and processes it independently. It uses the execution sequence relationship of the calculation graph to perform inference operations, and group and allocate the hardware computing units to meet the maximum delay requirements and improve throughput.

Benefits of technology

While taking into account latency and throughput, it improves the efficiency of model inference, reduces waste of computing resources, and improves throughput, especially when the batch size is small, it significantly improves processing power.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118313458B_ABST
    Figure CN118313458B_ABST
Patent Text Reader

Abstract

A data processing method, a data processor, a storage medium, and an electronic device. The data processing method includes: obtaining an inference model corresponding to an inference task and at least one input data for the inference task; using multiple execution groups to execute the inference task in parallel, where each execution group includes multiple hardware computing units, and different execution groups receive different input data to perform inference operations on the received input data using the inference model. This data processing method provides a parallel strategy for model inference that takes into account both throughput and latency. By using multiple execution groups to execute the inference task in parallel, and each execution group receiving different input data, inference operations can be performed on multiple input data simultaneously, thereby improving the throughput of model inference. In addition, each execution group in this data processing method includes multiple hardware computing units, and compared with using 1 hardware computing unit to execute the complete model inference, the latency of model inference can be taken into account.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to the field of artificial intelligence technology, and in particular, to a data processing method, a data processor, an electronic device, and a non-transitory computer-readable storage medium. Background Art

[0002] To improve the computing efficiency of a neural network, usually, multiple adjacent operators in the neural network that meet certain conditions or rules are fused to form a fused operator. Generally, a fused operator or a single operator that cannot be fused can be referred to as a fused operator, also known as a fusion layer. The computing process of the neural network is carried out layer-by-layer with the fusion layer as the unit. Usually, the output of the previous layer (or the previous several layers) serves as the input of the next layer (or the next several layers), so there is data dependency between the fusion layers.

[0003] When specifically performing an inference task, it is implemented by running the kernel functions corresponding to the respective fused operators. The kernel function is the source program compiled from the fused operator. Before running the kernel function, an execution stream needs to be created. Each execution stream is an operation queue, and the kernel functions within the execution stream are executed in sequence according to the execution order between the operators in the computation graph. Summary of the Invention

[0004] At least one embodiment of the present disclosure provides a data processing method, including: obtaining an inference model corresponding to an inference task and at least one input data for the inference task; and performing the inference task in parallel by using a plurality of execution groups, where each execution group includes a plurality of hardware computing units, and different execution groups receive different input data to perform an inference operation on the received input data by using the inference model.

[0005] For example, in the data processing method provided by at least one embodiment of the present disclosure, the input data input into the same execution group is serially performed the inference operation in sequence, and the plurality of hardware computing units included in each execution group are configured to jointly perform the inference operation on the input data input into the execution group.

[0006] For example, in the data processing method provided by at least one embodiment of the present disclosure, the inference model includes at least one computation graph, and the at least one computation graph represents the data dependency relationship and execution order among multiple operators included in the inference model. Using multiple execution groups to execute the inference task in parallel includes: for each execution group: receiving an input data; taking the computation graph as a unit, and sequentially processing the at least one computation graph according to the execution order relationship among the at least one computation graph, so as to perform the inference operation on the one input data and obtain the inference result of the one input data, wherein each hardware computing unit included in the execution group is configured to process a part of the currently executed computation graph.

[0007] For example, in the data processing method provided by at least one embodiment of the present disclosure, taking the computation graph as a unit, and sequentially processing the at least one computation graph according to the execution order relationship among the at least one computation graph, so as to perform the inference operation on the one input data and obtain the inference result of the one input data, includes: according to the number P of hardware computing units included in the execution group, dividing the one input data into Q sub-data, where P and Q are positive integers, and Q is less than or equal to P; respectively inputting the Q sub-data into Q of the P hardware computing units included in the execution group, so as to sequentially process the at least one computation graph according to the execution order relationship among the at least one computation graph, and obtain Q inference intermediate results respectively output by the Q hardware computing units; obtaining the inference result based on the Q inference intermediate results.

[0008] For example, in the data processing method provided by at least one embodiment of the present disclosure, dividing the one input data into Q sub-data according to the number P of hardware computing units included in the execution group includes: dividing the one input data into P intermediate data according to the number P of hardware computing units included in the execution group; in response to the size of each intermediate data being greater than or equal to the minimum execution size requirement of a single hardware computing unit, determining the P intermediate data as the Q sub-data; in response to the size of each intermediate data being less than the minimum execution size requirement, dividing the input data into the Q sub-data according to the minimum execution size requirement.

[0009] For example, in the data processing method provided by at least one embodiment of the present disclosure, the input data input to different execution groups performs the inference operation in parallel, and each execution group is independent of each other when performing the inference operation on the received input data.

[0010] For example, in the data processing method provided by at least one embodiment of the present disclosure, the inference task is executed in parallel by a plurality of execution groups, including: creating corresponding execution streams for the plurality of execution groups respectively, where each execution stream is an operation queue formed by kernel functions respectively corresponding to a plurality of operators included in the inference model arranged according to the data dependency relationship and execution order between the plurality of operators, and the execution streams respectively corresponding to the plurality of execution groups are independent of each other; using a plurality of hardware computing units included in each execution group to run the execution stream corresponding to the execution group, so as to execute the inference operation on the input data input to the execution group.

[0011] For example, in the data processing method provided by at least one embodiment of the present disclosure, the plurality of execution groups are obtained by grouping M hardware computing units included in a data processor that executes the inference task, M is a positive integer greater than 1, and the number of the plurality of execution groups is related to the maximum latency requirement for performing the inference operation on a single input data.

[0012] For example, in the data processing method provided by at least one embodiment of the present disclosure, the number of hardware computing units included in each execution group is the same.

[0013] For example, in the data processing method provided by at least one embodiment of the present disclosure, when grouping the M hardware computing units, determine the minimum number of hardware computing units required to meet the maximum latency requirement when performing the inference operation on the single input data, and determine the number of the plurality of execution groups according to the minimum number, where the number of hardware computing units included in each execution group is greater than or equal to the minimum number.

[0014] At least one embodiment of the present disclosure provides a data processor, including an acquisition module and an execution module, the execution module includes M hardware computing units, the M hardware computing units are divided into a plurality of execution groups, M is a positive integer greater than 1, the acquisition module is configured to acquire an inference model corresponding to the inference task and at least one input data for the inference task; the execution module is configured to execute the inference task in parallel by using the plurality of execution groups, where each execution group includes a plurality of hardware computing units, and different execution groups receive different input data to execute the inference operation on the received input data by using the inference model.

[0015] For example, in the data processor provided by at least one embodiment of the present disclosure, the input data input to different execution groups performs the inference operation in parallel, and each execution group is independent of each other when executing the inference operation on the received input data.

[0016] For example, in the data processor provided by at least one embodiment of the present disclosure, when the execution module executes the inference task in parallel using the multiple execution groups, it is configured to run multiple execution streams corresponding one-to-one to the multiple execution groups to execute the inference task in parallel. Among them, the execution stream corresponding to each execution group is an operation queue formed by the kernel functions corresponding to the multiple operators included in the inference model arranged according to the data dependency relationship and execution order between the multiple operators, and the multiple execution streams operate independently of each other.

[0017] For example, in the data processor provided by at least one embodiment of the present disclosure, the input data input to the same execution group undergoes the inference operation in sequence serially, and the multiple hardware computing units included in each execution group are configured to jointly execute the inference operation on the input data input to the execution group.

[0018] For example, in the data processor provided by at least one embodiment of the present disclosure, the inference model includes at least one computational graph, and the at least one computational graph characterizes the data dependency relationship and execution order between the multiple operators included in the inference model. The execution group is configured to run the execution stream corresponding to the execution group, and process the at least one computational graph in sequence according to the execution order relationship between the at least one computational graph as a unit to perform the inference operation on the input data input to the execution group. Among them, each hardware computing unit is configured to process a part of the currently executed computational graph.

[0019] For example, in the data processor provided by at least one embodiment of the present disclosure, the number of the multiple execution groups is related to the maximum latency requirement for performing the inference operation on a single input data.

[0020] For example, in the data processor provided by at least one embodiment of the present disclosure, the number of hardware computing units included in each execution group is the same.

[0021] For example, in the data processor provided by at least one embodiment of the present disclosure, the data processor is a general-purpose graphics processor or a graphics processor, and the hardware computing unit is a streaming multiprocessor or a streaming processor cluster.

[0022] At least one embodiment of the present disclosure provides an electronic device, including: a memory that stores computer-executable instructions non-transiently; a processor configured to run the computer-executable instructions, wherein when the computer-executable instructions are run by the processor, the data processing method according to any one of the embodiments of the present disclosure is implemented.

[0023] At least one embodiment of the present disclosure provides a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, a data processing method according to any embodiment of the present disclosure is implemented. Description of the Drawings

[0024] To more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings of the embodiments will be briefly introduced below. Obviously, the drawings in the following description only relate to some embodiments of the present disclosure and do not limit the present disclosure.

[0025] Figure 1 A schematic flowchart of a data processing method provided by at least one embodiment of the present disclosure;

[0026] Figure 2 Shows a schematic structural diagram of a general graphics processor;

[0027] Figure 3A Shows a schematic diagram of a directed acyclic graph of a neural network;

[0028] Figure 3B Shows according to Figure 3A The implementation process of obtaining a computational graph by operator fusion according to the shown directed acyclic graph;

[0029] Figure 3C Shows according to Figure 3B The execution sequence obtained by serializing the shown computational graph;

[0030] Figure 4 A schematic diagram of the processing process of the data processing method provided by an embodiment of the present disclosure;

[0031] Figure 5 A schematic block diagram of a data processor provided by at least one embodiment of the present disclosure;

[0032] Figure 6 A schematic block diagram of an electronic device provided by an embodiment of the present disclosure;

[0033] Figure 7 A schematic diagram of a non-transitory computer-readable storage medium provided by at least one embodiment of the present disclosure. Detailed Embodiments

[0034] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present disclosure. Obviously, the described embodiments are only a part rather than all of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the described embodiments of the present disclosure without creative efforts shall fall within the scope of protection of the present disclosure.

[0035] Unless otherwise defined, the technical terms or scientific terms used in the present disclosure shall have the ordinary meanings understood by those of ordinary skill in the art to which the present disclosure pertains. The terms "first", "second", and similar terms used in the present disclosure do not denote any order, quantity, or importance, but are merely used to distinguish different components. The terms such as "including" or "comprising" mean that the elements or items appearing before this word cover the elements or items listed after this word and their equivalents, without excluding other elements or items. The terms such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left", "right", etc. are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly. To keep the following description of the embodiments of the present disclosure clear and concise, detailed descriptions of some known functions and known components are omitted in the present disclosure.

[0036] The processor can execute machine learning tasks. The machine learning tasks can be inference tasks. For example, the machine learning tasks can be learning tasks that perform inference on input data using a neural network as an inference model. The inference model can be an artificial neural network (Artificial Neural Networks). Artificial neural networks are also simply referred to as neural networks and are algorithmic mathematical models that mimic the behavioral characteristics of animal neural networks and perform distributed parallel information processing. Such a network relies on the complexity of the system and adjusts the relationships between a large number of internal nodes to achieve the purpose of processing information. Regardless of the type of artificial neural network, their common characteristics include large-scale parallel processing, distributed storage, elastic topology, high redundancy, and non-linear operations, etc., and they have capabilities in aspects such as computing speed, associative ability, adaptability, fault tolerance ability, and self-organization ability. These characteristics and capabilities constitute the technical basis for artificial neural networks to simulate intelligent activities and have obtained important applications in various technical fields. For example, artificial neural networks can be used in data compression, data processing, video coding, signal processing, etc.

[0037] Before the processor runs the execution flow to perform an inference task, it is necessary to specify the number of hardware computing units for running the execution flow. The hardware computing units can refer to the Stream Processor Cluster (SPC) in General-purpose computing on graphics processing units (GPGPU), or the Stream Processor Cluster in Graphics Processing Unit (GPU), etc.

[0038] The process of model inference is to obtain the final output result of the model based on the input data (samples). To complete the inference task, the inference model needs to have strong generalization ability, that is, it can apply the knowledge learned during the training process to new and unseen data. In addition, factors such as computational efficiency and real-time performance also need to be considered during the inference process to meet the requirements of practical applications. Latency and throughput are the two most important metrics for measuring model inference performance. Latency refers to the time required for the inference model to complete a full inference on an input data, and throughput refers to the number of samples that the inference model can process per unit time.

[0039] The number of samples input to the model can be one or more. When the number of input data is different, different parallel strategies can be adopted. Usually, an execution flow can run on only one hardware computing unit or on all hardware computing units, corresponding to two different parallel strategies.

[0040] Currently, in one parallel strategy, the inference model in each hardware computing unit is complete, that is, 1 hardware computing unit performs a complete inference operation on 1 input data, and different hardware computing units process different input data. This strategy is called the data parallel strategy. The data parallel strategy has a relatively large throughput and can process multiple input data simultaneously, but has a high latency.

[0041] In another parallel strategy, at a certain moment, one input data is processed, and the inference model is split so that each hardware computing unit only processes a part of the inference model. This strategy is called the model parallel strategy. The inference latency of the model parallel strategy is relatively low, but the data throughput is relatively low.

[0042] During model inference, inference can be performed in batches. The number of data samples input to the model at one time is called the batch size. The batch size is an important hyperparameter in deep learning and other machine learning applications, which affects the performance, memory usage, and computational efficiency of the model.

[0043] When the batch size is 1, during the inference process, only a single input data needs to be input into the model and the output is calculated. This can reduce the time for the processor to load data and lower the video memory occupancy rate, thereby improving the inference efficiency. However, setting a higher batch size can increase the parallelism, thus utilizing more computing resources and improving the computing efficiency. The larger the batch size, the more data needs to be loaded into the memory simultaneously, which may lead to the problem of insufficient memory. For parallel computing devices such as graphics processors, using a larger batch size can more effectively utilize the computing resources because the graphics processor can process multiple samples in parallel. However, an overly large batch size may result in an excess of computing resources and no further improvement in the computing efficiency.

[0044] During the model inference stage, the batch size is usually selected to be consistent with that in the training stage to ensure consistency and comparability. However, sometimes, in order to accelerate the inference speed or reduce the memory usage, a smaller batch size may be selected.

[0045] In the above two parallel strategies, especially when the batch size is relatively small, in order to maximize the utilization of the hardware, the above model parallel strategy is usually adopted. Although the model parallel strategy can reduce the latency of model inference, since the kernel functions within an execution stream are all executed sequentially, during model inference, the input data is processed serially, and the throughput of the model is very low. Moreover, each hardware computing unit has a minimum execution size requirement for the input data. If the input data is below this minimum execution size requirement, the calculation cannot be performed. For example, assume that the size of the input data is 100 * 100 and the minimum execution size requirement of each hardware computing unit is 20. Then, only 5 hardware computing units can be used to perform model inference, and the remaining hardware computing units are idle, resulting in a waste of computing resources.

[0046] At least one embodiment of the present disclosure provides a data processing method, a data processor, an electronic device, and a non-transitory storage medium. The data processing method includes: obtaining an inference model corresponding to an inference task and at least one input data for the inference task; and using a plurality of execution groups to execute the inference task in parallel, where each execution group includes a plurality of hardware computing units, and different execution groups receive different input data to perform inference operations on the received input data by using the inference model.

[0047] The data processing method provided by at least one embodiment of the present disclosure provides a model inference parallel strategy that takes into account both throughput and latency. The data processing method uses multiple execution groups to execute inference tasks in parallel, and each execution group receives different input data. Thus, inference operations can be performed on multiple input data simultaneously, thereby improving the throughput of model inference. In addition, each execution group in the data processing method includes multiple hardware computing units, that is, the inference task can be executed by multiple hardware computing units. Compared with using 1 hardware computing unit to execute the complete model inference, the data processing method can also take into account the latency of model inference, thereby providing a model inference parallel strategy that takes into account both throughput and latency.

[0048] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings, but the present disclosure is not limited to these specific embodiments.

[0049] Figure 1 It is a schematic flowchart of a data processing method provided by at least one embodiment of the present disclosure. As Figure 1 shown, the data processing method provided by at least one embodiment of the present disclosure includes step S10 and step S20.

[0050] For example, the data processing method can be used in a data processor including M hardware computing units, where M is a positive integer greater than 1.

[0051] For example, the data processor can be a graphics processor or a general-purpose graphics processor, and the hardware computing unit can be a streaming processor cluster.

[0052] Figure 2 Shows a schematic structural diagram of a general-purpose graphics processor.

[0053] As Figure 2 shown, the general-purpose graphics processor is actually an array of streaming processor clusters. For example, it includes Figure 1 the streaming processor cluster 1 shown, …, the streaming processor cluster M, where M is a positive integer greater than 1. In the graphics processor, 1 streaming processor cluster processes one computing task, or multiple streaming processor clusters process one computing task. Data sharing is performed between multiple streaming processor clusters through a global cache or global memory.

[0054] As Figure 2 shown, taking the streaming processor cluster 1 as an example, 1 streaming processor cluster includes multiple computing units. For example, Figure 1The computing units 1, 2, ..., N in it, where N is a positive integer. Each computing unit (abbreviated as CU) is used to perform operations such as accumulation, reduction, conventional addition, subtraction, multiplication, division, etc. A computing unit includes multiple cores (also known as computing cores or computing nuclei), and each core includes an arithmetic logic unit (ALU), a floating-point computing unit, etc. The cores are used to perform specific computing tasks. In addition, the computing unit also includes registers and shared caches, which are used to hierarchically store the source data and destination data related to the computing tasks. The shared cache in a computing unit is used to share data among the cores of the computing unit.

[0055] In parallel computing, computing tasks are generally executed by multiple threads. Before these threads are executed in a general-purpose graphics processing unit (or called a parallel computing processor), they are divided into multiple thread blocks, and then Figure 1 (not shown in the figure) distributes multiple thread blocks to each computing unit through a thread block distribution module. All the threads in a thread block must be assigned to the same computing unit for execution. At the same time, the thread block will be split into the smallest execution thread bundles (or simply called thread bundles, warps), and each thread bundle contains a fixed number (or less than this fixed number) of threads. For example, 32 threads. Multiple thread blocks can be executed in the same computing unit or in different computing units.

[0056] In each computing unit, a thread bundle scheduling / distribution module Figure 1 (not shown in the figure) schedules and allocates the thread bundles so that multiple cores in the computing unit can run the thread bundles. According to the number of cores in the computing unit, multiple thread bundles in a thread block can be executed simultaneously or time-divisionally. Multiple threads in each thread bundle will execute the same instructions. Memory execution instructions will be sent to the shared cache in the computing unit or further sent to the mid-level cache or global cache or global memory for read / write operations, etc.

[0057] In step S10, obtain the inference model corresponding to the inference task and at least one input data for the inference task.

[0058] For example, an inference task can refer to the process of using a trained model to predict or classify new data. In other words, inference is the process by which a model, after receiving new, unlabeled data, classifies, predicts, or makes decisions on this new data based on the knowledge it learned from the training data before.

[0059] For example, the inference task can be an image recognition task, the inference model can be a pre-trained image category recognition model, and the input data for the inference task can be one or more new pictures received. Based on the features learned by the inference model during training, the category to which this picture belongs (such as a dog, a cat, a car, etc.) can be inferred as the inference result.

[0060] For example, the inference task can be a speech recognition task, the inference model can be a pre-trained speech recognition model, and the input data for the inference task can be a segment of speech received. Based on the features learned by the inference model during training, it can be converted into text as the inference result, thus realizing the function of speech-to-text conversion.

[0061] According to different usage scenarios, the inference task can also be other inference tasks, and the corresponding inference models and input data can also be different according to different usage scenarios, etc. The present disclosure does not make specific limitations on this.

[0062] For example, the input data can be part or all of all input data sets for the inference task. For example, the input data can belong to the same batch of input data (data samples input into the model for processing at one time) or different batches of input data. The present disclosure does not make specific limitations on this.

[0063] In step S20, the inference task is executed in parallel by multiple execution groups.

[0064] For example, each execution group includes multiple hardware computing units.

[0065] For example, multiple execution groups are obtained by grouping the M hardware computing units included in the data processor for executing the inference task, where M is a positive integer greater than 1.

[0066] For example, the hardware computing unit can be Figure 2 the shown streaming processor cluster as Figure 2 shown. The general graphics processor includes M streaming processor clusters. The M streaming processor clusters can be divided into multiple execution groups, and each execution group includes multiple streaming processor clusters.

[0067] For example, the number of multiple execution groups is related to the maximum latency requirement for performing an inference operation on a single input data. For example, the maximum latency requirement is the maximum latency value stipulated in advance for obtaining an inference result by performing a complete inference operation on one input data.

[0068] For example, when grouping the M hardware computing units, determine the minimum number of hardware computing units required to meet the maximum latency requirement when performing an inference operation on a single input data, and determine the number of multiple execution groups according to this minimum number, where the number of hardware computing units included in each execution group is greater than or equal to this minimum number.

[0069] Since the more hardware computing units are used to perform reasoning operations on one input data at the same time, the smaller the delay is, the minimum number of hardware computing units required to perform reasoning operations on a single input data in order to meet the maximum delay requirement can be determined. For example, based on the maximum delay requirement, it is determined that t hardware computing units can be used to meet the maximum delay requirement, so it can be stipulated that the number of hardware computing units included in each execution group must be greater than or equal to t, so that each execution group can at least meet the maximum delay requirement, and t is a positive integer. On the premise of meeting the maximum delay requirement, the number of execution groups is set as much as possible, that is, the number of hardware computing units included in each execution group is as close to t as possible, for example, all t, so that setting more groups can obtain higher throughput, achieving both delay and throughput improvement.

[0070] For example, according to the maximum latency requirement, the number of hardware computing units included in each execution group can be set to be the same, for example, all are t. This allows each execution group to have consistent performance when performing inference operations on input data, without fluctuations in model inference latency, thereby maximizing throughput and reducing latency.

[0071] For example, in some embodiments, the number of execution groups and the number of hardware computing units included in each execution group can be dynamically adjusted according to the maximum delay requirements of different reasoning tasks. For example, when executing reasoning task 1, it can be stipulated that each execution group includes two hardware computing units according to the maximum delay requirements of reasoning task 1. When executing reasoning task 2, it can be stipulated that each execution group includes four hardware computing units according to the maximum delay requirements of reasoning task 2.

[0072] For example, in other embodiments, the number of execution groups and the number of hardware computing units included in each execution group may be predefined based on the relevant information of the executed reasoning tasks, and the same execution group configuration may be used for model reasoning when the executed reasoning tasks and batch sizes are different. For example, assuming that the data processor includes 9 hardware computing units, it may be specified that every 3 hardware computing units are an execution group, and the 9 hardware computing units are divided into 3 execution groups. When the batch sizes of the input data and the reasoning tasks are different, the above configuration is used for model reasoning.

[0073] For example, different execution groups receive different input data to perform reasoning operations on the received input data using the reasoning model.

[0074] For example, the reasoning operation refers to using a complete reasoning model to reason about received input data to obtain a reasoning result of the input data.

[0075] For example, the input data of different execution groups are input in parallel for inference operations, and each execution group is independent of each other when performing inference operations on the received input data. That is, each execution group independently performs inference operations on the input data received by this execution group, and there is no data dependency or data constraint between execution groups.

[0076] For example, different execution groups receive different input data, so that inference operations can be performed on multiple different input data in parallel. Multiple samples can be processed in parallel at one operation moment, thereby improving the throughput of model inference. The throughput of the model can be increased exponentially, and the specific increase multiple depends on the number of execution groups. For example, when the number of execution groups is 3, the throughput is increased by 3 times compared with the model parallel strategy.

[0077] For example, when specifically performing inference operations, an independent execution flow can be created for each execution group. For example, each execution flow is an operation queue formed by kernel functions respectively corresponding to multiple operators included in the inference model arranged according to the data dependency relationship and execution order between the multiple operators. For example, the execution flows corresponding to multiple execution groups are independent of each other.

[0078] For example, taking the inference model as a neural network, the neural network can be regarded as a directed acyclic graph (DAG) composed of many computing nodes, and each node corresponds to an operator. Specifically, Figure 3A shows a schematic diagram of the directed acyclic graph of the neural network. As Figure 3A shown, a directed acyclic graph composed of 13 computing nodes is shown, that is, the corresponding neural network includes 13 operators. In addition, the nodes are connected by connection lines, and the connection lines represent the data dependency relationship and data flow direction between the computing nodes. For example, the output data of node 1 flows to node 2, and the output data of node 2 flows to node 3 and node 6, and so on. Thus, it can be determined that there is a data dependency relationship between node 1 and node 2, and there is a data dependency relationship between node 2 and node 6 as well as node 3. It can be understood that Figure 3A the network structure shown in is only schematic, and the method according to the embodiments of the present disclosure can be applied to various types of neural network structures.

[0079] Next, according to Figure 3A the shown directed acyclic graph, operator fusion can be performed to obtain a fused computation graph. Figure 3B shows the implementation process of obtaining a computation graph by performing operator fusion according to Figure 3A the shown directed acyclic graph.

[0080] In the related art, for the calculation of a neural network, in order to more efficiently execute the calculation process of operators in the network, operator fusion can be performed according to some rules and patterns, that is, multiple operators are fused into one fused operator, such asFigure 3B As shown, operator 1 and operator 2 can be fused into a fused operator A. Similarly, operator 3, operator 4, and operator 5 can be fused into a fused operator C, and operator 6, operator 7, operator 8, and operator 9 can be fused into a fused operator B, and so on. In addition, Figure 3B the computational graph in also includes individual operators that do not fuse with other operators. For example, operator D and operator E. According to the above rules of operator fusion, a computational graph composed of operators (including individual operators that cannot be fused and fused operators) and the connections between the operators can be obtained. It can be understood that the embodiments of the present disclosure do not limit the specific rules for performing operator fusion, and conventional methods in related technologies can be adopted.

[0081] Figure 3C shows an execution sequence obtained by serializing the computational graph shown in Figure 3B According to the embodiments of the present disclosure, serialization refers to converting a non-linear structured graph into a time-ordered linear execution sequence according to the data dependency relationships and graph traversal rules on the above computational graph, and retaining the data dependency relationships on the graph, while ensuring that all data dependency relationships are unidirectional on the time axis, that is, having a structure as shown in Figure 3C This execution sequence is in units of operators, and each operator in the execution sequence starts computing timely and in sequence.

[0082] For example, an execution stream is an operation queue obtained by arranging the kernel functions compiled from the operators according to the execution order and data dependency relationships shown in the execution sequence. The kernel functions in the execution stream are executed in sequence one by one, thereby realizing the sequential startup of computing for each operator. That is to say, in the execution stream, each element is the kernel function corresponding to the operator.

[0083] For example, according to the inference model, an independent execution stream can be created for each execution group. When using multiple execution groups to execute inference tasks in parallel, the multiple hardware computing units included in each execution group are used to run the execution stream corresponding to the execution group to perform inference operations on the input data input to the execution group. For example, when an input data is ready, an execution group can be started to run the execution stream to perform inference operations on the input data.

[0084] Of course, since the inference models used for different input data in the inference task are the same, the kernel functions in the execution streams run by each execution group essentially have the same execution order and data dependency relationships. However, the execution streams run by each execution group are independent of each other and there is no data dependency or data constraint between them. That is to say, although the kernel functions in each execution stream have the same execution order relationship and data dependency relationship, different execution streams are completely parallel, and different input data input to the same execution stream are executed in sequence.

[0085] For example, the input data of the same execution group is sequentially and serially subjected to inference operations. Each of the multiple hardware computing units included in each execution group is configured to jointly perform an inference operation on the input data input to the execution group. For example, each hardware computing unit is configured to execute a part of the inference model.

[0086] For example, an inference model can generally be divided into at least one computation graph according to a fusion strategy or the support of operators by a processor. Different computation graphs can have different computation scales. For example, an inference model includes at least one computation graph, and the at least one computation graph represents the data dependency relationship and execution order among multiple operators included in the inference model.

[0087] For example, an inference model can include 1 computation graph, such as Figure 3B the computation graph after operator fusion shown in

[0088] For example, an inference model can include multiple computation graphs. For example, the inference model can be divided into multiple computation graphs of different sizes according to a fusion strategy or the support of operators by a processor. Each computation graph includes 1 or more operators. For example, Figure 3B the computation graph after operator fusion can be used as the inference model and further divided into multiple computation graphs. For example, it can be divided into three computation graphs. Computation Figure 1 includes operators A, B, and C. Computation Figure 2 includes operator E. Computation graph 3 includes operators D and F. For example, in this case, there are also execution order relationships and data dependency relationships among the computation graphs. For example, computation Figure 2 needs to wait until the output results of operators B and C in computation Figure 1 are ready before starting to execute. Computation graph 3 needs to wait until the output results of operator C in computation Figure 1 and operator E in computation Figure 2 are ready before starting to execute. Therefore, the 3 computation graphs need to be executed in sequence according to the execution order of computation Figure 1 ->computation Figure 2 ->computation graph 3.

[0089] For example, step S20 can include: for each execution group: receiving an input data; taking the computation graph as a unit, and processing at least one computation graph in accordance with the execution order relationship among at least one computation graph to perform an inference operation on an input data and obtain an inference result of the input data, where each hardware computing unit included in the execution group is configured to process a part of the currently executed computation graph.

[0090] For example, taking a computational graph as a unit, according to the execution order relationship between at least one computational graph, using an execution group to sequentially process at least one computational graph to perform an inference operation on an input data and obtain an inference result of the input data, which may include: dividing an input data into Q sub-data according to the number P of hardware computing units included in the execution group, where P and Q are positive integers and Q is less than or equal to P; inputting the Q sub-data into Q of the P hardware computing units included in the execution group respectively to sequentially process at least one computational graph according to the execution order relationship between at least one computational graph and obtain Q intermediate inference results respectively; and obtaining an inference result based on the Q intermediate inference results.

[0091] For example, dividing an input data into Q sub-data according to the number P of hardware computing units included in the execution group may include: dividing an input data into P intermediate data according to the number P of hardware computing units included in the execution group; determining the P intermediate data as Q sub-data in response to the size of each intermediate data being greater than or equal to the minimum execution size requirement of a single hardware computing unit; and dividing the input data into Q sub-data according to the minimum execution size requirement in response to the size of each intermediate data being less than the minimum execution size requirement.

[0092] For example, if the size of the intermediate data is greater than or equal to the minimum execution size requirement of a single hardware computing unit, then P = Q; if the size of the intermediate data is less than the minimum execution size requirement, divide the input data into Q sub-data according to the minimum execution size requirement, and at this time Q is less than P.

[0093] For example, in some embodiments, assuming the number of hardware computing units is M, dividing the hardware computing units into a execution groups according to the maximum latency requirement, where each execution group includes P hardware computing units, that is, M = a * P, and M, a, and P are positive integers.

[0094] For example, the a execution groups can process a input data simultaneously at most, and each execution group independently performs an inference operation on the received input data.

[0095] In an execution group, the P hardware computing units jointly complete an inference operation on the received input data. For example, each hardware computing unit calculates 1 / P of the inference model.

[0096] In the present disclosure, each hardware computing unit is configured to compute a part of an inference model, such as 1 / P of the inference model (or compute 1 / Q of the inference model when Q is less than P). Since the hardware computing units are grouped in the present disclosure, P is less than M. For example, limited by the minimum execution size requirement of a single hardware computing unit, assuming that an input data requires P hardware computing units for inference operations, in the model parallel strategy, M - P hardware computing units are in an idle state, while in the data processing method provided by the present disclosure, all hardware computing units are in a working state, reducing the resource waste caused by the idle hardware computing units.

[0097] For example, it is assumed that the inference model includes c computational graphs, where c is a positive integer. The c computational graphs are arranged according to the data dependency relationship and execution order among multiple operators included in the inference model. Obtain the execution order relationship among the c computational graphs, and the c computational graphs are sequentially executed in units of computational graphs according to this execution order relationship. For example, the created execution flow embodies this execution order relationship, and specifically executes each kernel function according to this execution order relationship. Figure 1 Computation Figure 2 Computation

[0098] For execution group 1 in a execution groups, after receiving input data 1, it starts the execution flow and performs inference operations on input data 1 in the order of Computation Figure 1 Computation Figure 2 Computation

[0099] ... Computational graph c to obtain the inference result of input data 1. For example, for a currently executed computational graph, each hardware computing unit executes a part of the computations in the computational graph.

[0100] For example, the input data can be evenly divided into P pieces of intermediate data. If the size of each piece of intermediate data is greater than or equal to the minimum execution size requirement of a single hardware computing unit, these intermediate data can be used as sub-data and input into P hardware computing units respectively for inference operations.

[0101] For example, if the size of each piece of intermediate data is smaller than the minimum execution size requirement of a single hardware computing unit, the input data is divided into Q sub-data according to the minimum execution size requirement, the size of each sub-data is the minimum execution size requirement, and the Q sub-data are respectively input into Q hardware computing units for inference operations.

[0102] For example, each hardware computing unit sequentially executes the c computational graphs on the input sub-data for inference operations. For example, in the order of computational Figure 1 , computational Figure 2 ,... computational graph c, the c computational graphs are sequentially processed to obtain the inference intermediate result corresponding to the sub-data. Finally, according to the positional relationship of the sub-data input to each hardware computing unit, the inference intermediate results obtained by each hardware computing unit are re-stitched to obtain the inference result of the input data.

[0103] Specifically, at the kernel function level, when using multiple hardware computing units included in each execution group to run the execution stream corresponding to the execution group to perform inference operations on the input data of the execution group, the multiple hardware computing units within an execution group can calculate each computational graph sequentially in units of computational graphs by running the execution stream created for the execution group, that is, sequentially execute each kernel function located in the execution stream. Each hardware computing unit executes a part of the computational graph, that is, processes the sub-data input to the hardware computing unit to obtain the inference intermediate result.

[0104] The model inference strategy that takes into account both latency and throughput provided by at least one embodiment of the present disclosure is applicable to scenarios with different batch sizes, especially applicable to scenarios where the batch size of the input data set for the inference task is relatively small compared to the number of hardware computing units, such as the batch size is less than or equal to 2 times the number of hardware computing units. For example, when the batch size is relatively small, such as less than or equal to 2 times the number of hardware computing units, the hardware execution units can be grouped and multiple execution groups can be used for inference operations. For example, when the batch size is relatively large, such as greater than 2 times the number of hardware computing units, the inference operations can be performed according to the data parallel strategy.

[0105] The data processing method provided by at least one embodiment of the present disclosure can flexibly configure the number of execution groups according to the specific scenario. On the one hand, the throughput is improved by multiple execution groups performing inferences on different input data in parallel. On the other hand, the latency of a single execution group is reduced by multiple hardware computing units in an execution group jointly performing inferences on the input data. The provided operator parallel strategy can take into account both latency and throughput.

[0106] Figure 4 It is a schematic diagram of the processing process of the data processing method provided by an embodiment of the present disclosure. The following combines Figure 4 , and specifically describes the data processing method provided by at least one embodiment of the present disclosure.

[0107] As Figure 4 shown in, the input data pool (sample pool or input data set) includes multiple input data, such as Figure 4 the boxes shown in, which respectively represent input data 1, input data 2,..., input data 16,....

[0108] For example, the data processor for performing inference tasks includes 9 hardware computing units, namely Figure 4 the hardware computing unit 1, hardware computing unit 2,..., hardware computing unit 9 in

[0109] The 9 hardware computing units are divided into 3 execution groups, and each execution group includes 3 hardware computing units. For example, as Figure 4 shown, execution group 1 includes hardware execution unit 1, hardware execution unit 2, and hardware execution unit 3, execution group 2 includes hardware execution unit 4, hardware execution unit 5, and hardware execution unit 6, and execution group 3 includes hardware execution unit 7, hardware execution unit 8, and hardware execution unit 9.

[0110] For example, the maximum latency of the inference operation on a single input data expected by the user can be obtained in advance, and it is determined that using 3 hardware computing units can meet this latency requirement. Therefore, it can be set that the 9 computing units are divided into 3 groups, and each group includes 3 hardware computing units. For example, the number of hardware computing units included in each execution group is the same, so that the performance of each execution group is consistent when performing the inference operation on the input data, and there will be no fluctuations in the model inference latency, and the throughput can be improved as much as possible and the latency can be reduced.

[0111] As Figure 4 shown, when input data 1 is ready, start execution group 1 to run the execution flow corresponding to execution group 1, and use hardware computing unit 1, hardware computing unit 2, and hardware computing unit 3 to perform inference operations on input data 1. After obtaining the inference result of input data 1, input data 4 is input into execution group 1, and hardware computing unit 1, hardware computing unit 2, and hardware computing unit 3 are used to perform inference operations on input data 4. After obtaining the inference result of input data 4, input data 7 is input into execution group 1, and hardware computing unit 1, hardware computing unit 2, and hardware computing unit 3 are used to perform inference operations on input data 7, and so on.

[0112] Similarly, execution group 2 sequentially performs inference operations on input data 2, input data 5, input data 8, etc. in series, and execution group 3 sequentially performs inference operations on input data 3, input data 6, input data 9, etc. in series.

[0113] Thus, three hardware computing units can be utilized to complete the inference operation on an input data, ensuring that the latency of the inference operation meets the requirements and is as low as possible. For example, compared with the data parallel strategy, the latency is reduced by three times.

[0114] For example, as Figure 4 shown, input data 1, input data 2, and input data 3 can be used in parallel by three execution groups to perform inference operations in parallel; after completing the inference operations on input data 1, input data 2, and input data 3, input data 4, input data 5, and input data 6 can be input into the three execution groups and the hardware computing units within each execution group can be used to perform inference operations in parallel; after completing the inference operations on input data 4, input data 5, and input data 6, input data 7, input data 8, and input data 9 can be input into the three execution groups and the hardware computing units within each execution group can be used to perform inference operations in parallel, and so on. Thus, the inference operations on three input data can be executed simultaneously, and the throughput is increased by three times compared with the model parallel strategy.

[0115] In this embodiment, each input data is not restricted as to whether it belongs to the same batch. For example, input data 1, 2, and 3 can be the same batch of sample data or different batches of sample data. Similarly, input data 1, 4, and 7 can be the same batch of sample data or different batches of sample data. In the present disclosure, when an input data is ready, it can be input into an idle execution group for inference operation.

[0116] In this embodiment, on the one hand, the throughput is improved by performing inferences on different input data in parallel through multiple execution groups, and on the other hand, the latency of a single execution group is reduced by jointly performing inferences on the input data by multiple hardware computing units in one execution group.

[0117] At least one embodiment of the present disclosure further provides a data processor, Figure 5 which is a schematic block diagram of a data processor provided by at least one embodiment of the present disclosure.

[0118] As Figure 5 shown, the data processor 100 includes an acquisition module 101 and an execution module 102.

[0119] As Figure 5 shown, the execution module 102 includes M hardware computing units 103, where M is a positive integer greater than 1. For example Figure 5 the hardware computing units 103_1, hardware computing units 103_2,..., hardware computing units 103_M shown.

[0120] For example, the M hardware computing units 103 can be divided into multiple execution groups, and each execution group includes multiple hardware computing units.

[0121] For example, the number of multiple execution groups is related to the maximum latency requirement for performing inference operations on a single input data. For the determination of the number of execution groups and the grouping strategy of the hardware computing units, reference can be made to the relevant content in the foregoing data processing method, and the repeated parts will not be elaborated herein.

[0122] For example, the number of hardware computing units included in each execution group is the same. This can ensure that the performance of each execution group is consistent when performing inference operations on the input data, avoid fluctuations in model inference latency, and improve throughput and reduce latency as much as possible.

[0123] The obtaining module 101 is configured to obtain an inference model corresponding to an inference task and at least one input data for the inference task.

[0124] For the relevant descriptions of the inference model and the input data, reference can be made to the relevant descriptions in step S10 of the foregoing data processing method, and the repeated parts will not be elaborated herein.

[0125] The execution module 102 is configured to perform the inference task in parallel by using multiple execution groups.

[0126] For example, different execution groups receive different input data to perform inference operations on the received input data by using the inference model.

[0127] For example, the input data input to different execution groups are subjected to inference operations in parallel, and each execution group is independent of each other when performing inference operations on the received input data.

[0128] For example, different execution groups receive different input data, so that inference operations can be performed on multiple different input data in parallel, and multiple samples can be processed in parallel at one operation moment, thereby improving the throughput of model inference. Moreover, the throughput of the model can be increased exponentially, and the specific increase multiple depends on the number of execution groups. For example, when the number of execution groups is 3, the throughput is increased by 3 times compared with the model parallel strategy.

[0129] For example, when the execution module performs the inference task in parallel by using multiple execution groups, it is configured to run multiple execution streams corresponding one by one to the multiple execution groups to perform the inference task in parallel.

[0130] For example, an independent execution stream can be created for each execution group. For example, each execution stream is an operation queue formed by arranging kernel functions corresponding to multiple operators included in the inference model according to the data dependency relationship and execution order between the multiple operators. For example, the execution streams corresponding to the multiple execution groups are independent of each other.

[0131] For the relevant content of the execution stream, reference can be made to the relevant content in the foregoing data processing method, and the repeated parts will not be elaborated herein.

[0132] For example, according to the inference model, an independent execution flow can be created for each execution group. When using multiple execution groups to execute inference tasks in parallel, the multiple hardware computing units included in each execution group are used to run the execution flow corresponding to the execution group, so as to perform inference operations on the input data input to the execution group. For example, when an input data is ready, an execution group can be started to run the execution flow to perform inference operations on the input data. Thus, inference operations can be performed on multiple input data in parallel at the same time, and the inference results of these input data can be obtained, improving the throughput of the inference process.

[0133] For example, the input data input to the same execution group are serially and sequentially subjected to inference operations, and the multiple hardware computing units included in each execution group are configured to jointly perform inference operations on the input data input to the execution group.

[0134] For example, the inference model includes at least one computation graph, and at least one computation graph characterizes the data dependency relationship and execution order among multiple operators included in the inference model. The execution group is configured to run the execution flow corresponding to the execution group, and sequentially process at least one computation graph in accordance with the execution order relationship among at least one computation graph in units of the computation graph, so as to perform inference operations on the input data input to the execution group. For example, each hardware computing unit is configured to process a part of the currently executed computation graph.

[0135] For example, each hardware computing unit is configured to run the execution flow to perform inference operations on the received sub-data. For example, at least one computation graph is sequentially processed in accordance with the execution order relationship among at least one computation graph in units of the computation graph to obtain the intermediate inference result corresponding to the sub-data. Then, based on the intermediate inference results output by each hardware computing unit within the execution group, the inference result of the input data is obtained. Since each hardware computing unit is used to execute a part of the inference model, that is, the hardware computing unit processes the sub-data obtained by dividing the input data, and the amount of input data is small, therefore, by jointly performing inference operations on the input data by multiple hardware computing units, the latency can be reduced.

[0136] For the specific content of the execution group performing inference operations on the input data input to the execution group, reference can be made to the relevant descriptions in the foregoing data processing method, and the repeated parts will not be elaborated.

[0137] For example, the data processor can be a general-purpose graphics processor or a graphics processor, and the hardware computing unit can be a streaming multiprocessor or a streaming processor cluster.

[0138] For example, the hardware computing unit can be Figure 2 the shown streaming processor cluster as Figure 2As shown, the general graphics processor includes M streaming processor clusters. The M streaming processor clusters can be divided into multiple execution groups, and each execution group includes multiple streaming processor clusters.

[0139] It should be noted that Figure 5 The components and structures of the data processor 100 shown are only exemplary and not restrictive. According to needs, the data processor 100 may also have other components and structures.

[0140] For example, these modules can be implemented by hardware (such as circuits) modules, software modules, or any combination of the two. The same applies to the following embodiments and will not be elaborated further. For example, these units can be implemented by a central processing unit (CPU), a graphics processor, a tensor processing unit (TPU), a field programmable gate array (FPGA), or other forms of processing units with data processing capabilities and / or instruction execution capabilities, along with corresponding computer instructions.

[0141] For example, the data processor can be a processor that needs to perform inference operations. For example, the data processor can be a general-purpose processor such as a central processing unit, a multi-core processor, a graphics processor, a general graphics processor, or a digital signal processor, or a dedicated processor such as an AI (Artificial Intelligence) processor, a tensor processing unit, or a field programmable gate array. For example, the data processor can be a component specifically provided for optimized computing in some graphics processors or AI accelerators, such as a deep learning accelerator. Of course, the present disclosure is not limited thereto, and the data processor can be any device that needs to execute a data processing method to perform corresponding inference operations.

[0142] It should be noted that the acquisition module 101 can be used to implement Figure 1 the step S10 shown, and the execution module 102 can be used to implement Figure 1 the step S20 shown. Therefore, for a specific description of the functions that the acquisition module 101 can implement, reference can be made to the relevant description of step S10 in the embodiment of the above data processing method. For a specific description of the functions that the execution module 102 can implement, reference can be made to the relevant description of step S20 in the embodiment of the above data processing method. Repeated parts will not be elaborated further.

[0143] The data processor provided by at least one embodiment of the present disclosure can flexibly configure the number of execution groups according to specific scenarios. On the one hand, by multiple execution groups reasoning on different input data in parallel to improve throughput, and on the other hand, by multiple hardware computing units in one execution group jointly reasoning on the input data to reduce the latency of a single execution group. The provided operator parallel strategy can take both latency and throughput into account. The data processor 100 can achieve similar technical effects as the aforementioned data processing method, which will not be elaborated here.

[0144] It should be noted that, in the embodiments of the present disclosure, the data processor 100 may include more or fewer circuits or units, and the connection relationships between the various circuits or units are not limited and can be determined according to actual needs. The specific composition manners of the various circuits or units are not limited and can be composed of analog devices according to circuit principles, or composed of digital chips, or in other applicable manners.

[0145] Figure 6 It is a schematic block diagram of an electronic device provided by an embodiment of the present disclosure. As Figure 6 shown, the electronic device 200 is, for example, suitable for implementing the data processing method provided by the embodiments of the present disclosure. It should be noted that, Figure 6 the components of the electronic device 200 shown are exemplary and not restrictive. According to actual application needs, the electronic device 200 may also have other components.

[0146] As Figure 6 shown, the electronic device 200 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 201, which can perform various appropriate actions and processes according to non-transitory computer-readable instructions stored in the memory to achieve various functions.

[0147] For example, when the computer-readable instructions are run by the processing device 201, one or more steps in the data processing method according to any of the above embodiments can be executed. It should be noted that the detailed description of the processing process of the data processing method can refer to the relevant descriptions in the embodiments of the above data processing method, and the repeated parts will not be elaborated.

[0148] For example, the memory may include any combination of one or more computer program products, and the computer program products may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) 203 and / or cache memory, etc. For example, computer-readable instructions may be loaded from the storage device 208 into the random access memory (RAM) 203 to run the computer-readable instructions. Non-volatile memory may include, for example, read-only memory (ROM) 202, hard disk, erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), USB memory, flash memory, etc. Various application programs and various data may also be stored in the computer-readable storage media, such as style images, and various data used and / or generated by the application programs, etc.

[0149] For example, the processing device 201, read-only memory 202, and random access memory 203 are connected to each other via the bus 204. The input / output (I / O) interface 205 is also connected to the bus 204.

[0150] Generally, the following devices may be connected to the input / output interface 205: an input device 206 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 207 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 208 including, for example, a magnetic tape, a hard disk, a flash memory, etc.; and a communication device 209. The communication device 209 may allow the electronic device 200 to communicate with other electronic devices wirelessly or wiredly to exchange data. Although Figure 6 the electronic device 200 with various devices is shown, it should be understood that it is not required to implement or have all the shown devices, and the electronic device 200 may alternatively implement or have more or fewer devices. For example, the processor 201 may control other components in the electronic device 200 to perform the desired functions. The processor 201 may be a central processing unit, a tensor processing unit, or a graphics processing unit, etc., which has data processing capabilities and / or program execution capabilities. The central processing unit may be of the X86 or ARM architecture, etc. The graphics processing unit may be directly integrated onto the motherboard alone or built into the northbridge chip of the motherboard. The graphics processing unit may also be built into the central processing unit

[0151] Figure 7 Schematic diagram of a non-transitory computer-readable storage medium provided by at least one embodiment of the present disclosure. For example, as Figure 7As shown, the storage medium 300 can be a non-transitory computer-readable storage medium, and one or more computer-readable instructions 301 can be non-temporarily stored on the storage medium 300. For example, when the computer-readable instructions 301 are executed by a processor, the data processing method described above can be performed.

[0152] For example, the storage medium 300 can be applied to the above-mentioned electronic device. For example, the storage medium 300 can include the memory in the electronic device 200.

[0153] For example, the storage medium can include a memory card of a smart phone, a storage component of a tablet computer, a hard disk of a personal computer, a random access memory, a read-only memory, an erasable programmable read-only memory, a portable compact disc read-only memory, a flash memory, or any combination of the above storage media, and can also be other applicable storage media.

[0154] For example, the description of the storage medium 300 can refer to the description of the memory in the embodiment of the electronic device, and the repeated parts will not be described again.

[0155] Those skilled in the art can understand that the content disclosed in the present disclosure can have various variations and improvements. For example, the various devices or components described above can be implemented by hardware, or can be implemented by software, firmware, or some or all of the combinations of the three.

[0156] In addition, although the present disclosure makes various references to certain units in the system according to the embodiments of the present disclosure, however, any number of different units can be used and run on the client and / or server. The units are only illustrative, and different aspects of the system and method can use different units.

[0157] Flowcharts are used in the present disclosure to illustrate the steps of the methods according to the embodiments of the present disclosure. It should be understood that the steps before or after do not necessarily have to be carried out precisely in order. On the contrary, they can be carried out in reverse order or various steps can be processed simultaneously. At the same time, other operations can also be added to these processes.

[0158] Those of ordinary skill in the art can understand that all or part of the steps in the above methods can be completed by instructing relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium, such as a read-only memory, a magnetic disk, or an optical disc, etc. Optionally, all or part of the steps of the above embodiments can also be implemented using one or more integrated circuits. Correspondingly, the modules / units in the above embodiments can be implemented in the form of hardware or in the form of software function modules. The present disclosure is not limited to any specific form of the combination of hardware and software.

[0159] Unless otherwise defined, all terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. It should also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning that is consistent with their meaning in the context of the relevant art and should not be interpreted in an idealized or overly formal sense unless expressly so defined herein.

[0160] The foregoing is a description of the disclosure and should not be construed as a limitation thereof. Although several exemplary embodiments of the disclosure have been described, those skilled in the art will readily appreciate that many modifications can be made to the exemplary embodiments without departing from the novel teachings and advantages of the disclosure. Accordingly, all such modifications are intended to be included within the scope of the disclosure as defined by the claims. It should be understood that the foregoing is a description of the disclosure and should not be considered limited to the particular embodiments disclosed, and modifications to the disclosed embodiments as well as other embodiments are intended to be included within the scope of the appended claims. The disclosure is defined by the claims and their equivalents.

Claims

1. A data processing method, applied to a processor, wherein, The processor includes multiple execution groups, and each execution group includes multiple hardware computing units. The data processing method includes: Obtaining an inference model corresponding to an inference task and multiple input data for the inference task, where the multiple input data are input into the inference model for the inference model to perform an inference operation on the multiple input data to obtain inference results of the multiple input data; Using the multiple execution groups to execute the inference task in parallel, where each execution group deploys the inference model, and different execution groups receive different input data to execute the inference operation on the received input data in parallel using the inference models deployed by themselves; Wherein, the inference model includes at least one computational graph, and the at least one computational graph represents the data dependency relationship and execution order between multiple operators included in the inference model; Using multiple execution groups to execute the inference task in parallel includes: For each execution group: Receiving one input data; Taking the computational graph as a unit, and sequentially processing the at least one computational graph according to the execution order relationship between the at least one computational graph by using the execution group to perform the inference operation on the one input data to obtain the inference result of the one input data, where each hardware computing unit included in the execution group is configured to process a part of the currently executed computational graph.

2. The data processing method according to claim 1, wherein The input data input into the same execution group are sequentially serially performed with the inference operation. The multiple hardware computing units included in each execution group are configured to jointly execute the inference operation on the input data input into the execution group.

3. The data processing method according to claim 1, wherein Taking the computational graph as a unit, and sequentially processing the at least one computational graph according to the execution order relationship between the at least one computational graph by using the execution group to perform the inference operation on the one input data to obtain the inference result of the one input data, includes: Dividing the one input data into Q sub-data according to the number P of hardware computing units included in the execution group, where P and Q are positive integers, and Q is less than or equal to P; Inputting the Q sub-data into Q of the P hardware computing units included in the execution group respectively to sequentially process the at least one computational graph according to the execution order relationship between the at least one computational graph to obtain Q inference intermediate results respectively output by the Q hardware computing units; Obtaining the inference result based on the Q inference intermediate results.

4. The data processing method according to claim 3, wherein, Dividing the one input data into Q sub-data according to the number P of hardware computing units included in the execution group, includes: Dividing the one input data into P intermediate data according to the number P of hardware computing units included in the execution group; In response to the size of each intermediate data being greater than or equal to the minimum execution size requirement of a single hardware computing unit, determining the P intermediate data as the Q sub-data; In response to the size of each intermediate data being less than the minimum execution size requirement, dividing the input data into the Q sub-data according to the minimum execution size requirement.

5. The data processing method according to claim 1, wherein, The input data of different execution groups are used to perform the inference operation in parallel, and each execution group is independent of each other when performing the inference operation on the received input data.

6. The data processing method according to claim 5, wherein, Using multiple execution groups to perform the inference task in parallel includes: Creating corresponding execution streams for the multiple execution groups respectively, where each execution stream is an operation queue formed by kernel functions corresponding to multiple operators included in the inference model arranged according to the data dependency relationship and execution order between the multiple operators, and the execution streams corresponding to the multiple execution groups are independent of each other; Using multiple hardware computing units included in each execution group to run the execution stream corresponding to the execution group to perform the inference operation on the input data input to the execution group.

7. The data processing method according to any one of claims 1-6, wherein, The multiple execution groups are obtained by grouping M hardware computing units included in a data processor for performing the inference task, where M is a positive integer greater than 1. The number of the multiple execution groups is related to the maximum latency requirement for performing the inference operation on a single input data.

8. The data processing method according to claim 7, wherein, The number of hardware computing units included in each execution group is the same.

9. The data processing method according to claim 7, wherein When grouping the M hardware computing units, determine the minimum number of hardware computing units required to meet the maximum latency requirement when performing the inference operation on the single input data, and determine the number of the multiple execution groups according to the minimum number, where the number of hardware computing units included in each execution group is greater than or equal to the minimum number.

10. A data processor, including an acquisition module and an execution module. The execution module includes M hardware computing units, and the M hardware computing units are divided into multiple execution groups, each execution group including multiple hardware computing units, where M is a positive integer greater than 1. The obtaining module is configured to obtain an inference model corresponding to an inference task and a plurality of input data for the inference task, wherein, The multiple input data are input to the inference model for the inference model to perform an inference operation on the multiple input data to obtain inference results of the multiple input data. The execution module is configured to use the multiple execution groups to perform the inference task in parallel, where each execution group deploys the inference model respectively, and different execution groups receive different input data to perform the inference operation on the received input data in parallel by using the inference model deployed by each of them. The inference model includes at least one computation graph, and the at least one computation graph represents the data dependency relationship and execution order between multiple operators included in the inference model. The execution group is configured to run the execution stream corresponding to the execution group, and process the at least one computation graph in sequence according to the execution order relationship between the at least one computation graph in units of the computation graph to perform the inference operation on the input data input to the execution group, where each hardware computing unit is configured to process a part of the currently executed computation graph.

11. The data processor according to claim 10, wherein, The input data of different execution groups are used to perform the inference operation in parallel, and each execution group is independent of each other when performing the inference operation on the received input data.

12. The data processor according to claim 11, wherein, When the execution module performs the inference task using the multiple execution groups in parallel, it is configured to run multiple execution streams corresponding to the multiple execution groups one by one to perform the inference task in parallel. Among them, the execution flow corresponding to each execution group is an operation queue formed by arranging kernel functions corresponding to multiple operators included in the inference model according to the data dependency relationship and execution order among the multiple operators, and the multiple execution flows operate independently of each other.

13. The data processor according to claim 10, wherein The input data input into the same execution group are sequentially and serially subjected to the inference operation. The multiple hardware computing units included in each execution group are configured to jointly execute the inference operation on the input data input into the execution group.

14. The data processor according to any one of claims 10-13, wherein, The number of the multiple execution groups is related to the maximum latency requirement for performing the inference operation on a single input data.

15. The data processor according to any one of claims 10-13, wherein, The number of hardware computing units included in each of the execution groups is the same.

16. The data processor according to any one of claims 10-13, wherein, The data processor is a general-purpose graphics processor or a graphics processor, and the hardware computing unit is a streaming processor cluster.

17. An electronic device, comprising: a memory that stores computer-executable instructions non-transiently; a processor configured to run the computer-executable instructions, wherein, when the computer-executable instructions are run by the processor, the data processing method according to any one of claims 1-9 is implemented.

18. A non-transitory computer-readable storage medium, wherein, The non-transient computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, the data processing method according to any one of claims 1-9 is implemented.

Citation Information

Patent Citations

  • Disparity map processing method and device, storage medium and electronic equipment

    CN116957901A

  • Model reasoning method and device based on calculation unit deployment, equipment and medium

    CN117494816A