Similar calculation optimization method and device for tensor programs

By identifying and eliminating redundant calculations in tensor programs in the artificial intelligence programming framework, optimizing calculation charts, and generating efficient task execution solutions, the problem of redundant calculations in the existing technology is solved, and computing efficiency and resource utilization are improved.

CN118605931BActive Publication Date: 2025-07-08TSINGHUA UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410533655.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-29
Publication Date
2025-07-08
Estimated Expiration
2044-04-29

AI Technical Summary

Technical Problem

The existing automatic optimization technology fails to effectively identify and utilize the characteristics and properties of input and output tensors in the artificial intelligence programming framework, making it difficult to detect and optimize redundant calculations.

Method used

By obtaining shared inputs in user requests, segmenting the task processing calculation diagram as a subgraph, and specialized calculation diagram optimization is performed on each subgraph, identifying and eliminating redundant calculations, and using batch optimization irregular operators to generate efficient task execution plans.

Benefits of technology

It realizes automatic identification and elimination of redundant computing in the artificial intelligence programming framework, improving computing efficiency and resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118605931B_ABST
    Figure CN118605931B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a similarity calculation optimization method and apparatus for tensor programs. A plurality of user requests each including at least one initial data are obtained, and the initial data included in at least two user requests is determined as a shared input. At least one task execution plan is determined according to the shared input, a plurality of sub-graphs obtained by splitting a task processing computational graph, and a plurality of specialized computational graphs corresponding to each sub-graph. Each task execution plan is used to execute at least one user request, and includes a target specialized computational graph corresponding to each sub-graph. Then, according to each task execution plan, similarity optimization is performed on the task processing computational graph to obtain an optimized processing computational graph for executing the corresponding user request. The present disclosure automatically identifies the redundancy of the shared input in the data information after receiving the user request, and then selects specialized computational graphs corresponding to a plurality of sub-graphs in the task computational graph through redundancy, and formulates an adaptive and efficient task execution plan to automatically optimize the task computational graph.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular, to a method and apparatus for optimizing similarity calculation of tensor programs. Background Art

[0002] In order to generate efficient executable code on general-purpose processors or domain-specific artificial intelligence chips, typical artificial intelligence programming frameworks represent artificial intelligence applications as tensor programs, and their calculation processes are represented as calculation graphs composed of tensors and operators. In the compilation optimization stage, artificial intelligence applications represented by calculation graphs are optimized through multiple layers, and more efficient calculation graphs are found through equivalent transformations of the calculation graphs. In terms of graph optimization, existing work is mainly divided into two categories: manual optimization and automatic optimization. Manual optimization, that is, domain experts artificially discover optimization rules and apply them to the compilation optimization process, is widely used in artificial intelligence programming frameworks including TensorFlow, TensorRT, TVM, MetaFlow, etc. However, this type of optimization requires system development engineers to have high domain knowledge and spend a lot of time finding optimization opportunities. Therefore, automatic optimization has become a new research direction in academia and industry in recent years. And existing automatic optimization technologies all focus on the calculation itself and its equivalence, ignoring the characteristics and properties of the input and output tensors involved in the calculation. Therefore, there are a large number of optimization opportunities such as redundant calculations in increasingly complex artificial intelligence applications that are difficult to be automatically discovered and utilized. Summary of the Invention

[0003] In view of this, the present disclosure proposes a method and apparatus for optimizing similarity calculation of tensor programs, aiming to automatically formulate an efficient execution plan using redundant calculations and perform similarity optimization of calculation graphs.

[0004] According to a first aspect of the present disclosure, there is provided a method for optimizing similarity calculation of tensor programs, the method including:

[0005] A method for optimizing similarity calculation of tensor programs, characterized in that the method includes:

[0006] Obtain a plurality of user requests, each user request including at least one initial data, and each initial data having a corresponding position in the user request;

[0007] Determine the initial data located at the same position in at least two of the user requests as shared inputs;

[0008] Determine a plurality of subgraphs obtained by splitting a task processing calculation graph, and a plurality of specialized calculation graphs corresponding to each of the subgraphs;

[0009] Determine at least one task execution plan according to at least one of the shared inputs and the multiple specialized computational graphs corresponding to each subgraph, where each task execution plan is used to execute at least one user request, including the target specialized computational graph corresponding to each subgraph;

[0010] Perform similarity optimization on the task processing computational graph according to each task execution plan to obtain an optimized processing computational graph for executing the corresponding user request.

[0011] In a possible implementation, the task processing computational graph includes multiple operators, as well as the input tensors and output tensors of each operator. The input tensors include at least one input data that needs to be processed by the operator, and the output tensors include the output data obtained by calculating each input data in the input tensors through the operator. Determining the multiple subgraphs obtained by splitting the task processing computational graph, and the multiple specialized computational graphs corresponding to each subgraph includes:

[0012] Perform consistency analysis on the task computational graph, and cluster at least one operator with the same optimization enabling condition to obtain the corresponding subgraph;

[0013] Perform subgraph specialization according to the input tensors of each subgraph to obtain at least one corresponding candidate computational graph, and each candidate computational graph has a corresponding attribute signature, where the attribute signature is used to characterize whether there is redundancy in each input tensor corresponding to the candidate computational graph;

[0014] Perform format conversion on the input tensors and output tensors of the operators in each candidate computational graph according to the corresponding attribute signature to obtain the corresponding specialized computational graph.

[0015] In a possible implementation, the performing subgraph specialization according to the input tensors of each subgraph to obtain at least one corresponding candidate computational graph includes:

[0016] Determine multiple attribute signatures according to the number of input tensors of each subgraph;

[0017] Perform subgraph specialization on each subgraph to obtain a computational graph to be screened corresponding to each attribute signature;

[0018] Screen the computational graphs to be screened corresponding to each subgraph according to a preset computational graph screening rule to obtain at least one corresponding candidate computational graph.

[0019] In a possible implementation, the computational graph screening rule includes at least one of the following:

[0020] Delete the computational graphs to be screened that conflict with the attribute signature;

[0021] In response to the inclusion relationship existing between two computational graphs to be screened, delete the included computational graph to be screened;

[0022] Delete the computational graphs to be screened where all input tensors have redundancies.

[0023] In a possible implementation, the formatting the input tensors and output tensors of the operators in each of the candidate computational graphs according to the corresponding attribute signature to obtain the corresponding specialized computational graph includes:

[0024] For each of the candidate computational graphs, perform redundancy propagation calculation on each of the included operators according to the corresponding attribute signature to obtain the redundancy attributes of the input tensors and output tensors of each operator;

[0025] Format the input tensors and output tensors of the operator according to the redundancy attributes of the input tensors and output tensors of each operator to obtain the corresponding specialized computational graph.

[0026] In a possible implementation, the for each of the candidate computational graphs, performing redundancy propagation calculation on each of the included operators according to the corresponding attribute signature to obtain the redundancy attributes of the input tensors and output tensors of each operator includes:

[0027] Determine the redundancy propagation function of each operator, where the redundancy propagation function is used to determine the redundancy attributes of the output tensors after calculation according to the redundancy attributes of the input tensors before calculation;

[0028] For each of the candidate computational graphs, determine the redundancy attributes of the input tensors of the first operator therein according to the corresponding attribute signature;

[0029] Starting from the first operator in the candidate computational graph, calculate the redundancy attributes of the input tensors and output tensors of each of the operators therein in sequence according to the redundancy attributes of the input tensors and the redundancy propagation function.

[0030] In a possible implementation, the formatting conversion includes:

[0031] In response to the redundancy attributes of the input tensors and output tensors of the operator being redundant, perform dimensionality reduction processing on the input tensors and the output tensors.

[0032] In a possible implementation, the task processing computational graph includes multiple operators and the input tensors of each operator, and the determining the multiple subgraphs obtained by splitting the task processing computational graph and the multiple specialized computational graphs corresponding to each subgraph further includes:

[0033] For each of the specialized computation graphs, when the number of input tensors of an operator is greater than 1 and there are at least two input tensors with different shapes, determine that the operator is an irregular operator;

[0034] Perform batch processing optimization on the irregular operator to obtain an optimized specialized computation graph.

[0035] In a possible implementation, the performing batch processing optimization on the irregular operator to obtain an optimized specialized computation graph includes:

[0036] Determine the type of each irregular operator;

[0037] Perform corresponding batch processing optimization according to the type of the irregular operator to obtain an optimized specialized computation graph.

[0038] In a possible implementation, the performing corresponding batch processing optimization according to the type of the irregular operator includes:

[0039] In response to the type of the irregular operator being a data sharing operator, add corresponding shape conversion operators before and after the irregular operator according to the shapes of the input tensors, for converting the input tensors of the irregular operator into regular input tensors and performing reverse conversion on the output tensors of the irregular operator;

[0040] In response to the type of the irregular operator being a data independent operator, split the input tensors to parallel process each input data in the input tensors through the data independent operator.

[0041] In a possible implementation, the determining at least one task execution plan according to at least one of the shared inputs and multiple specialized computation graphs corresponding to each subgraph includes:

[0042] Group multiple user requests according to at least one of the shared inputs to obtain at least one task splitting plan, where each task splitting plan includes at least one user request group, and the user requests in the user request group with the number of user requests greater than 1 have shared inputs;

[0043] For each task splitting plan, determine the input tensors of the plan with the redundancy attribute being redundant according to the shared inputs in each user request group, and determine the input tensors of the plan with the redundancy attribute being non - redundant for other initial data;

[0044] According to the attribute signatures of the specialized computation graphs corresponding to each subgraph and the redundancy attributes of the input tensors of the plan, sequentially select the target specialized computation graphs corresponding to each subgraph in the execution order to generate corresponding candidate execution plans;

[0045] Based on multiple candidate execution plans corresponding to each of the task splitting plans, select the candidate execution plans corresponding to the optimal task splitting plan as the task execution plan.

[0046] In a possible implementation manner, the similarity optimization of the task processing computation graph according to each of the task execution plans to obtain an optimized processing computation graph for executing the corresponding user request includes:

[0047] For each of the task execution plans, replace the subgraphs in the task processing computation graph with the target specialization computation graphs corresponding to each of the subgraphs therein to obtain an optimized processing computation graph for executing the corresponding user request.

[0048] According to a second aspect of the present disclosure, there is provided a similarity calculation optimization device for a tensor program, the device including:

[0049] A request acquisition module, configured to acquire multiple user requests each including at least one initial data with a corresponding position;

[0050] A shared input determination module, configured to determine the initial data located at the same position in at least two of the user requests as shared inputs;

[0051] A subgraph specialization module, configured to determine multiple subgraphs obtained by splitting the task processing computation graph, and multiple specialization computation graphs corresponding to each of the subgraphs;

[0052] A plan generation module, configured to determine at least one task execution plan according to at least one of the shared inputs and the multiple specialization computation graphs corresponding to each of the subgraphs, each of the task execution plans being used to execute at least one user request, and including the target specialization computation graph corresponding to each of the subgraphs;

[0053] A computation graph optimization module, configured to perform similarity optimization on the task processing computation graph according to each of the task execution plans to obtain an optimized processing computation graph for executing the corresponding user request.

[0054] In a possible implementation manner, the task processing computation graph includes multiple operators, as well as the input tensors and output tensors of each of the operators, the input tensors including at least one input data to be processed by the operator, and the output tensors including the output data obtained by computing each of the input data in the input tensors through the operator. The subgraph specialization module is further configured to:

[0055] By performing consistency analysis on the task computation graph, cluster at least one operator with the same optimization enabling condition to obtain a corresponding subgraph;

[0056] Perform sub-graph specialization based on the input tensors of each said sub-graph to obtain at least one corresponding candidate computation graph, and each said candidate computation graph has a corresponding attribute signature, and the attribute signature is used to characterize whether there is redundancy in each said input tensor corresponding to the candidate computation graph;

[0057] Perform format conversion on the input tensors and output tensors of the operators in each said candidate computation graph according to the corresponding attribute signature to obtain the corresponding specialized computation graph.

[0058] In one possible implementation, the sub-graph specialization module is further configured to:

[0059] Determine a plurality of attribute signatures according to the number of input tensors of each said sub-graph input tensor;

[0060] Perform sub-graph specialization on each said sub-graph to obtain a computation graph to be screened corresponding to each said attribute signature;

[0061] Screen the computation graphs to be screened corresponding to each said sub-graph according to a preset computation graph screening rule to obtain at least one corresponding candidate computation graph.

[0062] In one possible implementation, the computation graph screening rule includes at least one of the following:

[0063] Delete the computation graph to be screened that conflicts with the attribute signature;

[0064] In response to the existence of an inclusion relationship between two computation graphs to be screened, delete the included computation graph to be screened;

[0065] Delete the computation graph to be screened in which all input tensors have redundancy.

[0066] In one possible implementation, the sub-graph specialization module is further configured to:

[0067] For each said candidate computation graph, perform redundancy propagation calculation on each operator included therein according to the corresponding attribute signature to obtain the redundancy attributes of the input tensors and output tensors of each operator;

[0068] Perform format conversion on the input tensors and output tensors of the operator according to the redundancy attributes of the input tensors and output tensors of each said operator to obtain the corresponding specialized computation graph.

[0069] In one possible implementation, the sub-graph specialization module is further configured to:

[0070] Determine the redundancy propagation function of each operator, and the redundancy propagation function is used to determine the redundancy attribute of the output tensor after calculation according to the redundancy attribute of the input tensor before calculation;

[0071] For each of the candidate computational graphs, determine the redundant attributes of the input tensor of the first operator therein according to the corresponding attribute signature;

[0072] Starting from the first operator in the candidate computational graph, calculate the redundant attributes of the input tensor and output tensor of each operator therein in sequence according to the redundant attributes of the input tensor and the redundant propagation function.

[0073] In a possible implementation manner, the format conversion includes:

[0074] In response to the redundant attributes of the input tensor and output tensor of the operator being redundant, perform dimensionality reduction processing on the input tensor and the output tensor.

[0075] In a possible implementation manner, the task processing computational graph includes multiple operators and the input tensors of each operator. The sub-graph specialization module is further configured to:

[0076] For each of the specialized computational graphs, when the number of input tensors of the operator is greater than 1 and there are at least two input tensors with different shapes, determine that the operator is an irregular operator;

[0077] Perform batch processing optimization on the irregular operator to obtain an optimized specialized computational graph.

[0078] In a possible implementation manner, the sub-graph specialization module is further configured to:

[0079] Determine the type of each irregular operator;

[0080] Perform corresponding batch processing optimization according to the type of the irregular operator to obtain an optimized specialized computational graph.

[0081] In a possible implementation manner, the sub-graph specialization module is further configured to:

[0082] In response to the type of the irregular operator being a data sharing operator, add corresponding shape conversion operators before and after the irregular operator according to the shape of the input tensor, for converting the input tensor of the irregular operator into a regular input tensor and performing reverse conversion on the output tensor of the irregular operator;

[0083] In response to the type of the irregular operator being a data-independent operator, split the input tensor to process each input data in the input tensor in parallel through the data-independent operator.

[0084] In a possible implementation manner, the scheme generation module is further configured to:

[0085] Group the multiple user requests according to at least one of the shared inputs to obtain at least one task splitting scheme. Each task splitting scheme includes at least one user request group, and the user requests in the user request groups with more than one user request have shared inputs;

[0086] For each task splitting scheme, determine a scheme input tensor with redundant attributes as redundant based on the shared inputs in each user request group, and determine a scheme input tensor with non-redundant attributes as non-redundant based on other initial data;

[0087] According to the attribute signatures of the specialized computation graphs corresponding to each subgraph and the redundant attributes of the scheme input tensors in the execution order, sequentially select the target specialized computation graphs corresponding to each subgraph to generate corresponding candidate execution schemes;

[0088] According to the multiple candidate execution schemes corresponding to each task splitting scheme, select the multiple candidate execution schemes corresponding to an optimal task splitting scheme as the task execution scheme.

[0089] In a possible implementation manner, the computation graph optimization module is further configured to:

[0090] For each task execution scheme, replace the subgraphs in the task processing computation graph with the target specialized computation graphs corresponding to each subgraph therein to obtain an optimized processing computation graph for executing the corresponding user requests.

[0091] According to a third aspect of the present disclosure, there is provided an electronic device, including: a processor; a memory for storing processor-executable instructions; wherein, the processor is configured to implement the above method when executing the instructions stored in the memory.

[0092] According to a fourth aspect of the present disclosure, there is provided a non-volatile computer-readable storage medium, on which computer program instructions are stored, wherein the computer program instructions implement the above method when executed by a processor.

[0093] According to a fifth aspect of the present disclosure, there is provided a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code. When the computer-readable code runs in a processor of an electronic device, the processor in the electronic device executes the above method.

[0094] In the embodiments of the present disclosure, multiple user requests including at least one initial data are obtained, and the initial data included in at least two user requests is determined as a shared input. At least one task execution plan is determined according to the shared input, multiple subgraphs obtained by splitting a task processing computational graph, and multiple specialized computational graphs corresponding to each subgraph, where each task execution plan is used to execute at least one user request and includes the target specialized computational graph corresponding to each subgraph. Then, the task processing computational graph is optimized for similarity according to each task execution plan to obtain an optimized processing computational graph for executing the corresponding user request. The present disclosure automatically identifies the redundancy of the shared input in the data information after receiving the user request, and then selects the specialized computational graphs corresponding to multiple subgraphs in the task computational graph through redundancy, and formulates an adaptive and efficient task execution plan to automatically optimize the task computational graph.

[0095] Other features and aspects of the present disclosure will become clear from the following detailed description of exemplary embodiments with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0096] The accompanying drawings, which are included in and constitute a part of this specification, illustrate exemplary embodiments, features, and aspects of the present disclosure together with the specification, and are used to explain the principles of the present disclosure.

[0097] Figure 1 The flowchart showing a method for optimizing similarity calculation of a tensor program according to an embodiment of the present disclosure;

[0098] Figure 2 The schematic diagram showing a process for determining a specialized computational graph according to an embodiment of the present disclosure;

[0099] Figure 3 The schematic diagram showing a process for redundant propagation calculation according to an embodiment of the present disclosure;

[0100] Figure 4 The schematic diagram showing a process for batch processing optimization according to an embodiment of the present disclosure;

[0101] Figure 5 The schematic diagram showing an optimized processing computational graph according to an embodiment of the present disclosure;

[0102] Figure 6 The schematic diagram showing a device for optimizing similarity calculation of a tensor program according to an embodiment of the present disclosure;

[0103] Figure 7 The schematic diagram showing an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0104] Various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. Like reference numerals in the drawings denote functionally identical or similar elements. Although various aspects of the embodiments are shown in the drawings, the drawings are not necessarily drawn to scale unless otherwise specified.

[0105] As used herein, the term "exemplary" means "serving as an example, embodiment, or illustration". Any embodiment described herein as "exemplary" is not necessarily to be construed as superior to or better than other embodiments.

[0106] In addition, for a better description of the present disclosure, numerous specific details are given in the following detailed description. Those skilled in the art should understand that the present disclosure can be implemented without some of these specific details. In some instances, methods, means, elements, and circuits well known to those skilled in the art are not described in detail so as to highlight the gist of the present disclosure.

[0107] The similar calculation optimization method of the tensor program according to the embodiments of the present disclosure can be executed by an electronic device such as a terminal device or a server. Among them, the terminal device can be any fixed or mobile terminal such as a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. The server can be a single server or a server cluster composed of multiple servers. Any electronic device can implement the similar calculation optimization method of the tensor program according to the embodiments of the present disclosure by a processor calling computer-readable instructions stored in a memory.

[0108] Figure 1 A flowchart showing a method for optimizing similar calculations of a tensor program according to an embodiment of the present disclosure is shown. As Figure 1 shown, the method for optimizing similar calculations of the tensor program according to the embodiments of the present disclosure may include the following steps S10 - S50.

[0109] Step S10, obtain a plurality of user requests.

[0110] In a possible implementation, an electronic device obtains multiple user requests in a receiving manner or the like. Each user request includes at least one initial data. The multiple initial data included in each user request may have corresponding positions in the user request. The electronic device can execute each user request by calculating the initial data included therein, and the electronic device performs different calculations on the initial data at different positions. The number of initial data included in each user request, and the execution methods corresponding to the same position are the same, that is, when executing each user request, the same calculations need to be performed on the initial data at the same position. Exemplarily, when the electronic device receives three user requests r1(a, x), r2(a, y), and r3(b, x) respectively, the initial data included in user request r1 are a and x, the initial data included in user request r2 are a and y, and the initial data included in user request r3 are b and x. When the electronic device executes user request r1, it needs to calculate based on the initial data a and x. When executing user request r2, it needs to calculate based on the initial data a and y. When executing user request r3, it needs to calculate based on the initial data b and x. Among them, the positions of a in r1, a in r2, and b in r3 are the same, and the positions of x in r1, y in r2, and x in r3 are the same.

[0111] Step S20: Determine the initial data located at the same position in at least two of the user requests as the shared input.

[0112] In a possible implementation, after the electronic device obtains multiple user requests, it compares the initial data included in each user request. When there is an initial data that exists in at least two user requests at the same time and the corresponding positions are the same, it determines that the initial data is the shared input of the multiple user requests including it. Exemplarily, when the electronic device receives three user requests r1(a, x), r2(a, y), and r3(b, x) respectively, the electronic device can determine that the initial data a is the shared input of user request r1 and user request r2, and the initial data x is the shared input of user request r1 and user request r3. When the electronic device receives three user requests r1(a, x), r2(a, y), and r3(b, a) respectively, the electronic device can determine that the initial data a is the shared input of user request r1 and user request r2. Although the initial data a is also included in user request r3, the position of the initial data a in user request r3 is different from the positions in user request r1 and user request r2. Therefore, a is not the shared input of user request r3.

[0113] Optionally, since the electronic device performs the same form of computational processing on each user request, when the electronic device processes different user requests and performs the same calculation on the shared inputs among them, it will cause computational redundancy. For example, when the electronic device receives three user requests r1(a, x), r2(a, y), and r3(b, x) respectively, the electronic device needs to perform calculations based on the initial data a and x when executing user request r1, perform calculations based on the initial data a and y when executing user request r2, and perform calculations based on the initial data b and x when executing user request r3. Among them, for the shared input a of user request r1 and user request r2, the same calculation is performed once when executing user request r1 and r2 respectively. For the shared input x of user request r1 and user request r3, the same calculation is performed once when executing user request r1 and r3 respectively, thus resulting in redundancy in the calculation process.

[0114] Step S30: Determine multiple sub-graphs obtained by splitting the task processing computational graph, and multiple specialized computational graphs corresponding to each of the sub-graphs.

[0115] In a possible implementation, the electronic device determines a task processing computational graph for processing each user request, multiple sub-graphs obtained by splitting the task processing computational graph, and multiple specialized computational graphs corresponding to each sub-graph. Among them, the form of the task processing computational graph can be in the form of a data flow graph, which includes multiple operators, and the input tensors and output tensors of each operator. The task processing computational graph can be split into multiple sub-graphs, and each sub-graph includes at least one operator whose data processing process is executed in sequence. Each sub-graph has a corresponding multiple specialized computational graphs. A specialized computational graph usually refers to a computational graph optimized for a specific application or field. The specialized computational graph in this embodiment is used to optimize the similarity of the task processing computational graph by replacing the corresponding sub-graph. Each specialized computational graph has a corresponding attribute signature, which is used to characterize whether there is redundancy in each input tensor corresponding to the specialized computational graph, that is, whether the same input data is included in each input tensor in the input specialized computational graph. Exemplarily, the redundancy attribute of each input tensor in the attribute signature can be represented by 0 or 1, where 0 represents non-redundant and 1 represents redundant.

[0116] Optionally, the multiple sub-graphs of the task processing computational graph in the embodiments of the present disclosure, and the multiple specialized computational graphs corresponding to each sub-graph can be determined in advance, or determined during the execution of this solution. At the same time, when the multiple sub-graphs corresponding to the task processing computational graph and the specialized computational graphs of the sub-graphs are determined in advance, they can be determined by the electronic device executing the embodiments of the present disclosure, or determined by other devices and sent to the current electronic device.

[0117] Further, the process of splitting the task processing computation graph in the embodiments of the present disclosure to obtain multiple subgraphs and further obtaining multiple specialized computation graphs corresponding to the multiple subgraphs may include performing consistency analysis on the task computation graph, clustering at least one operator with the same optimization enabling condition to obtain a corresponding subgraph. Then, subgraph specialization is performed according to the input tensors of each subgraph to obtain at least one corresponding candidate computation graph, and each candidate computation graph has a corresponding attribute signature, where the attribute signature is used to characterize whether there is redundancy in each input tensor corresponding to the candidate computation graph. Then, format conversion is performed on the input tensors of the operators in each candidate computation graph according to the corresponding attribute signature to obtain the corresponding specialized computation graph.

[0118] Optionally, if subgraph splitting is not performed and an optimized task processing computation graph is directly generated, it may lead to an exponential increase in the optimization scheme, increasing with the increase in the number of inputs. For example, if each input tensor has only one possible redundant attribute, the task processing computation graph with n operators will be specialized into 2 n execution engines, introducing unacceptable optimization overhead for the optimization of a task processing computation graph that may have dozens of operators. To solve this problem, we observe that there are interdependencies among intermediate tensors in terms of their redundant attributes. For example, the input tensors of some operators with adjacent execution orders are the same, and the redundant attributes of their input tensors and output tensors must also be the same. Therefore, many operators in the task processing computation graph share the same optimization enabling condition, and the electronic device can perform consistency analysis on the task computation graph to cluster at least one operator with the same optimization enabling condition to obtain a corresponding subgraph.

[0119] Figure 2 A schematic diagram showing a process of determining a specialized computation graph according to an embodiment of the present disclosure is shown. As Figure 2 shown, in order to cluster (split) the operators in the task processing computation graph into subgraphs, it is necessary to infer when each tensor (including the input tensors and output tensors of each operator) in the task computation graph is redundant, and consider the operators with the same situation of whether the input tensors are redundant as "operators with the same optimization enabling condition" and cluster them into one subgraph. Optionally, this problem can be solved by a preset redundant propagation function. For the input tensors of the initial input task processing computation graph (for example, tensors composed of initial data at the same position in at least two user requests), a conditional expression can be maintained to represent the relationship between the redundant attributes of the current input tensor and the redundant attributes of the input tensors of each operator in the task computation graph. Figure 2(a) shows a simplified scenario where redundant attributes are only for a single dimension. The orange operators a and b are only affected by input A. If input A is redundant, a redundant output is generated. The yellow operators c, e, and d are only affected by input B. If input B is redundant, a redundant output is generated. Therefore, the optimization enabling condition m(A) propagates through operators a and b, and the optimization enabling condition m(B) propagates through operators c, d, and e. Here, m(A) represents the redundant attribute of input tensor A, and m(B) represents the redundant attribute of input tensor B. Furthermore, it can be determined that the orange operators a and b are operators with the same optimization enabling condition and can be clustered to obtain the corresponding subgraph. The yellow operators c, e, and d are operators with the same optimization enabling condition and can be clustered to obtain the corresponding subgraph. For the green operator f, the redundant attribute can only propagate when both of its input tensors are redundant. Therefore, its conditional expression is m(A) ∧ m(B) (∧ represents logical "and"). For subsequent operators involving output tensor C, only whether m(A) and m(B) are satisfied simultaneously needs to be considered, rather than considering the two redundant attributes m(A) and m(B) separately. Therefore, it can also be determined that operator g and operator f are operators with the same optimization enabling condition and can be clustered to obtain the corresponding subgraph.

[0120] After propagating the conditional expressions, as Figure 2 (b) shows, the operators can be clustered into the same subgraph with commonality, provided that they can be optimized simultaneously. When the inputs of the operators have different conditional expressions, such as operator f, if m(a) or m(b) is satisfied, it can be specialized, which is different from its subsequent operators. The electronic device does not create an independent subgraph for it separately during the process of determining the subgraph, which may lead to performance degradation due to the overhead of scheduling during service. Instead, these operators are merged into their subsequent subgraphs, that is, operator f and operator g are merged into the same subgraph.

[0121] Furthermore, during subgraph specialization based on the input tensors of each subgraph, at least one corresponding candidate computation graph is obtained. Each candidate computation graph has a corresponding attribute signature, and the attribute signature is used to characterize whether there is redundancy in each input tensor corresponding to the specialized computation graph. Among them, in the embodiments of the present disclosure, multiple attribute signatures can be determined first according to the number of input tensors of each subgraph. Then, each subgraph is specialized to obtain the computation graphs to be screened corresponding to each attribute signature. The computation graphs to be screened corresponding to each subgraph are screened according to the preset computation graph screening rules to obtain at least one corresponding candidate computation graph.

[0122] Optionally, when the number of subgraph input tensors is n, 2 can be determined nIndividual attribute signatures. For example, when the input tensor is 1, two attribute signatures can be determined to be redundant and non-redundant respectively. When the number of input tensors of the subgraph is 2, four attribute signatures can be determined to be (redundant, redundant), (redundant, non-redundant), (non-redundant, redundant), and (non-redundant, non-redundant) respectively.

[0123] After determining multiple attribute signatures corresponding to each subgraph, subgraph specialization can be performed on the corresponding subgraph based on each attribute signature to obtain candidate computational graphs that match each attribute signature of the subgraph. For example, if the subgraph has one input and two attribute signatures, redundant and non-redundant, then there are two candidate computational graphs, one is the subgraph with the "redundant" attribute signature added, and the other is the subgraph with the "non-redundant" attribute signature added. That is to say, the candidate computational graphs are subgraphs with attribute labels added, and the subgraphs and their different labels form different candidate computational graphs. However, due to limited performance gains or the assumed attribute signatures not being satisfied, some candidate computational graphs obtained through specialization processing may be unnecessary or impossible. Therefore, unnecessary candidate computational graphs can be further deleted through a screening method, and this screening process can be implemented through preset computational graph screening rules. The computational graph screening rules can include three types of screening conditions, namely, deleting candidate computational graphs that conflict with the attribute signatures. For example Figure 2 in, the attribute signatures of the two inputs of operator e should both be redundant or both non-redundant because they have the same source and are both affected by input B. If one is redundant and the other is non-redundant, it is a conflict. When there is an inclusion relationship between two candidate computational graphs, delete the included candidate computational graph. For example, for the candidate computational Figure 1 of the same subgraph, the attribute signature is (i1 ∧ i2), and the candidate computational Figure 2 has an attribute signature of (i1 ∧ i2 ∧ i2), then the candidate computational Figure 2 is included by the candidate computational Figure 1 , and the candidate computational Figure 2 can be deleted. And delete candidate computational graphs where all input tensors are redundant.

[0124] Exemplarily, the computational graph screening rules can delete the corresponding specialized candidate computational graphs through e1 ∧ e2 → False (conflict condition), if e1 → e2 or e2 → e1 (the arrow represents the inclusion relationship), then delete the candidate computational graphs corresponding to e1, e2, and e1 ∧ e2 (duplication condition), and delete the candidate computational graphs corresponding to e1 ∧ e2 → i1 ∧ i2 ∧ ··· ∧ in (complete redundancy). Here, e1 and e2 are the redundant attributes corresponding to the input tensors of two input candidate computational graphs respectively, and i1 and i2, etc. represent the redundant attributes of the inputs of the overall application.

[0125] In a possible implementation, after the electronic device filters out the candidate computational graphs corresponding to each sub-graph through the computational graph screening rules, it converts each candidate computational graph into a specialized computational graph by converting the formats of the operators and output tensors in the candidate computational graph. Among them, for each candidate computational graph, redundant propagation calculations can be performed on each operator included therein according to the corresponding attribute signature to obtain the redundant attributes of the input tensors and output tensors of each operator. Then, according to the redundant attributes of the input tensors and output tensors of each operator, format conversion is performed on the input tensors and output tensors of the operator to obtain the corresponding specialized computational graph.

[0126] Optionally, the operators of tensors can propagate redundancy in various ways. For example, whether the length and width dimensions in the matrix multiplication result are redundant depends on the redundant nature of a single input tensor, and the output batch dimension is redundant only when the batch dimensions of both input tensors are redundant. Therefore, a set of redundant propagation functions for the operators of the operators can be used to perform redundant propagation calculations on each operator to obtain the redundant attributes of the input tensors and output tensors of each operator. This process can be to first determine the redundant propagation function of each operator, and the redundant propagation function is used to determine the redundant attributes of the output tensor after calculation according to the redundant attributes of the input tensor before calculation. Then, for each candidate computational graph, determine the redundant attributes of the input tensor of the first operator therein according to the corresponding attribute signature. Then, starting from the first operator in the candidate computational graph, calculate the redundant attributes of the input tensors and output tensors of each operator therein in turn according to the redundant attributes of the input tensor and the redundant propagation function.

[0127] Figure 3 A schematic diagram showing a redundant propagation calculation process according to an embodiment of the present disclosure. As Figure 3 shown, the above table can be the redundant propagation functions determined for the operators corresponding to different operators, and the following is a schematic diagram of calculating the redundant attributes of the input tensors and output tensors of each operator according to the redundant propagation function. Among them, the redundant attributes can be calculated by the formula OUT b =Φ(IN b ) IN b =∪ p∈pred(b) (OUT p ). Among them, Φ is the redundant propagation function, IN b and OUT b are the redundant attributes of the input tensor and output tensor of the basic block b composed of at least one operator (which can be determined according to the attribute signature corresponding to the candidate computational graph), and pred(b) is the set of predecessor blocks of the basic block b (for example, Figure 3The predecessor block of block2 is block1), and p is the block in this set. For the basic blocks in the candidate computation graph, the redundancy attributes of the input tensors are propagated to each operator within the basic block. For each operator's operator symbol, the corresponding redundancy propagation function is used to infer the redundancy of the output. The output tensors of the basic block are then forwarded to other basic blocks. In the case where the input tensors have multiple sources, the redundancy attributes of each dimension are combined separately.

[0128] Optionally, when there are loop-computing operators in the current candidate computation graph, the embodiments of the present disclosure also perform redundancy propagation calculations in a loop-iteration manner. The previous calculation results can be updated by the subsequent calculation results. When the redundancy attributes of the output tensors obtained after two consecutive rounds of calculations are the same, it can be determined that there is no further update in a single iteration, and the redundancy attribute of the current output tensor is determined as the corresponding redundancy attribute.

[0129] Furthermore, after obtaining the redundancy attributes of the input tensors and output tensors of each operator in the candidate computation graph, the input tensors and output tensors of the operators can be format-converted according to the redundancy attributes of the input tensors and output tensors of each operator to obtain the corresponding specialized computation graph. Due to the complexity of tensor operators, eliminating the redundancy therein is different from classical scalar operators. The optional redundancy elimination rule can be that when the redundancy attributes of the input tensors and output tensors of the operator are redundant, dimensionality reduction processing is performed on the input tensors and output tensors. With the help of this redundancy elimination rule, the embodiments of the present disclosure can eliminate two types of redundancies caused by data. The first type is redundant computation, which generates duplicate outputs with redundant dimensions. The embodiments of the present disclosure can convert the original input tensor calculation into a non-redundant form by dimensionality reduction, compressing the redundant dimensions of the input tensors and output tensors into one dimension. For complex cases, such as redundant input channels in convolution, additional calculations can also be performed through other preset application elimination rules (such as summing along the weight channels) to achieve equivalent results. At the same time, by recording the redundancy attributes, the embodiments of the present disclosure can restore the original shapes of these format-converted tensors by broadcasting data, eliminating as much redundancy as possible only when necessary while maintaining the equivalence of the results. The second type of redundancy eliminated by the embodiments of the present disclosure is redundant memory access, which can be merged through redundant tensor data to avoid the overhead caused by accessing the same data multiple times.

[0130] In a possible implementation, since the input tensors of the specialized computation graph may have different shapes, in order to further improve service efficiency, the embodiments of the present disclosure can also divide each operator in the task processing computation graph into a regular operator and an irregular operator according to whether the shapes of the input tensors are the same, so as to perform regularization processing on the irregular operators. Among them, for each specialized computation graph, when the number of input tensors of the operator is greater than 1 and there are at least two input tensors with different shapes, the operator can be determined as an irregular operator. Further batch processing optimization is performed on the irregular operators to obtain an optimized specialized computation graph.

[0131] Optionally, the process of batch processing optimization can include determining the type of each irregular operator, and then performing corresponding batch processing optimization according to the type of the irregular operator to obtain an optimized specialized computation graph. Among them, the types of irregular operators can include data sharing operators and data independent operators. A data sharing operator is an operator whose input tensors are shared among requests in batch processing, and a data independent operator is an operator whose input tensors are not shared among requests in batch processing. For example, layers involving weights, such as convolutional layers and linear layers, belong to the category of data sharing operators. This is because the weight tensors of these layers are shared among different requests within the batch. On the other hand, data independent operators include operators that do not involve weights, such as transpose and reduction operators.

[0132] Optionally, a technique called irregular operator regularization is introduced in the embodiments of the present disclosure to achieve the reuse of existing highly optimized computation kernels in irregular computations. This technique opportunistically converts irregular operators into standard operators and irregular data independent operators. Since the data independent operators in batch processing can be decomposed into computable task blocks that can be processed in parallel, the kernels of these operators are relatively easy to implement. Therefore, it is possible to benefit from the high efficiency of the kernel library with only a small amount of effort.

[0133] Specifically, for an irregular operator of the data sharing operator type, corresponding shape conversion operators can be added before and after the irregular operator according to the shape of the input tensor, for converting the input tensor of the irregular operator into a regular input tensor, that is, making the input tensor shapes the same, and performing reverse conversion on the output tensor of the irregular operator. That is, the irregular dimension can be converted into a regular dimension by fusing it with the batch dimension.

[0134] Figure 4 A schematic diagram showing a batch processing optimization process according to an embodiment of the present disclosure is as follows. As Figure 4 shown, for example, in the linear layer in 4(a), we can transpose and deform the irregular input tensor to fuse its batch dimension b and the irregular dimension into a regular (non-irregular) dimension. This process is similar toFigure 4 (b) The height connecting the two matrices. Therefore, existing well-optimized matrix multiplication operators can be used to perform irregular calculations. Optionally, embodiments of the present disclosure may employ a set of preset equivalent graph transformation rules to regularize irregular data sharing operators. The rule set stores rules related to dimension fusion and permutation, making it possible to convert irregular dimensions into regular dimensions. Based on the rule set, each available transformation rule can be applied to the irregular data sharing operator, and then it is evaluated whether the irregular dimensions are successfully converted into regular dimensions. For multiple candidate transformations, they can be analyzed using typical input tensors and the best result is selected.

[0135] Since the regularization process changes the data layout of the input tensor to match the existing well-optimized kernel, the layout transformation kernel may introduce a large overhead. Embodiments of the present disclosure combine two adjacent layout transformations together to minimize this overhead. This method can effectively reduce the individual transformation requirements and eliminate the layout transformation after the reverse transformation. This approach alleviates the potential performance impact and enables more efficient computing.

[0136] Furthermore, in batch processing, requests in data-independent operators do not share data, they can be decomposed into separate requests, and can be efficiently executed in an embarrassingly parallel manner. Therefore, for irregular operators of the data-independent operator type, embodiments of the present disclosure can split the input tensor to process each input data in the input tensor in parallel through the data-independent operator. In addition, since each individual request is a regular operator, existing optimization strategies and the computational micro-kernels of the corresponding operators can be directly reused.

[0137] Step S40: Determine at least one task execution plan according to at least one of the shared inputs and the multiple specialized computational graphs corresponding to each subgraph.

[0138] In a possible implementation manner, after the electronic device determines the shared inputs in multiple user requests and the specialized computational graphs of multiple subgraphs in the task processing computational graph, it can determine at least one task execution plan according to at least one shared input and the multiple specialized computational graphs corresponding to each subgraph. Among them, each task execution plan is used to execute at least one user request and includes the target specialized computational graph corresponding to each subgraph.

[0139] Specifically, the process of determining the task execution plan may include grouping multiple user requests according to at least one shared input to obtain at least one task segmentation plan. Each task segmentation plan includes at least one user request group, and the user requests in the user request group with more than one user request have a shared input. For each task segmentation plan, based on the shared input in each user request group, a scheme input tensor with redundant attributes determined to be redundant is determined, and for other initial data, a scheme input tensor with redundant attributes determined to be non-redundant is determined. According to the attribute signature of the specialized computation graph corresponding to each subgraph and the redundant attributes of the scheme input tensors in the execution order, the target specialized computation graph corresponding to each subgraph is sequentially selected to generate the corresponding candidate execution plan. Then, based on the multiple candidate execution plans corresponding to each task segmentation plan, the multiple candidate execution plans corresponding to the optimal task segmentation plan are selected as the task execution plan.

[0140] Exemplarily, when the electronic device receives three user requests, namely r1(a, x), r2(a, y), and r3(b, x), the electronic device can determine that the initial data a is the shared input of user request r1 and user request r2, and the initial data x is the shared input of user request r1 and user request r3. Based on the shared input a and the shared input x, the electronic device can determine that task segmentation plan 1 is {user request group 1: user request r1 and user request r2, user request group 2: user request r3}. Task segmentation plan 2 is {user request group 1: user request r1 and user request r3, user request group 2: user request r2}, and task segmentation plan 3 is {user request group 1: user request r1, user request r2, and user request r3}. For user request group 1 in task segmentation plan 1, the corresponding scheme input tensors can be (a, a) and (x, y), where the input data in the scheme input tensor (a, a) is the same, so the redundant attribute is redundant. The input data in the scheme input tensor (x, y) is different, so the redundant attribute is non-redundant.

[0141] Further, after the electronic device determines the solution input tensors corresponding to each user request group, it can select the target specialized computation graph corresponding to each subgraph in sequence according to the attribute signatures of the specialized computation graphs corresponding to each subgraph and the redundant attributes of the solution input tensors in the execution order, and generate the corresponding candidate execution solutions. For example, in the above example, for user request group 1 in task segmentation solution 1, if the input tensors are one redundant and one non-redundant, the specialized computation graph with attribute signatures (redundant, non-redundant) is selected as the target specialized computation graph. According to the execution order of the subgraphs, the above selection is performed for each subgraph in sequence, and each target specialized computation graph is obtained, constituting a candidate execution solution. Each user request group corresponds to a candidate execution solution. If there are multiple user request groups in a task segmentation solution, there are multiple corresponding candidate execution solutions. Among them, whether to perform batch processing optimization on the target specialized computation graph can also be selected according to the shapes of different solution input tensors. After obtaining the multiple candidate execution solutions corresponding to each task segmentation solution, the embodiments of the present disclosure can use dynamic programming to search for an optimal task execution solution from the multiple candidate execution solutions corresponding to a task segmentation solution based on user requests and corresponding data fingerprints. Each user request group can correspond to a task execution solution. The "optimal" criterion can be selected according to needs, and the present application does not limit this.

[0142] Step S50: Perform similarity optimization on the task processing computation graph according to each of the task execution solutions to obtain an optimized processing computation graph for executing the corresponding user request.

[0143] In a possible implementation manner, after the electronic device obtains multiple task execution solutions for processing at least one corresponding user request, for each task execution solution, it can replace the subgraphs in the task processing computation graph with the target specialized computation graphs corresponding to each subgraph therein to obtain an optimized processing computation graph for executing the corresponding user request. At the same time, according to the shapes of different solution input tensors, it can also sequentially determine whether the operators in each task processing computation graph need to be batch-processed and optimized. In the case where batch processing optimization is required, batch processing optimization is performed on the target specialized computation graph according to the operator type.

[0144] Figure 5 FIG. shows a schematic diagram of an optimized processing computation graph according to an embodiment of the present disclosure. As Figure 5 shown, the bold data is the shared data between different user requests. The electronic device can optimize the task processing computation graph based on the shared data in the manner of the embodiments of the present disclosure to obtain an optimized processing computation graph capable of batch-processing user requests.

[0145] Based on the above technical features, embodiments of the present disclosure automatically locate possible redundancies in parallel processing of user requests by fully combining the calculation process in the task processing computation graph with the information of the data in the user requests that need to be calculated, and generate a series of optimized processing computation graphs after eliminating redundancies to optimize the entire task processing process. At the same time, by means of a detailed attribute analysis of the tensors participating in the calculation, an adaptive and efficient execution scheme is formulated to schedule user requests with different similarities and generated shapes, so as to achieve better service efficiency.

[0146] Figure 6 FIG. shows a schematic diagram of a similarity calculation optimization device for a tensor program according to an embodiment of the present disclosure. As Figure 6 shown, the similarity calculation optimization device for the tensor program of the embodiment of the present disclosure may include:

[0147] A request acquisition module 60, configured to acquire a plurality of user requests each including at least one initial data with a corresponding position;

[0148] A shared input determination module 61, configured to determine the initial data located at the same position in at least two of the user requests as shared inputs;

[0149] A sub-graph specialization module 62, configured to determine a plurality of sub-graphs obtained by splitting the task processing computation graph, and a plurality of specialized computation graphs corresponding to each of the sub-graphs;

[0150] A scheme generation module 63, configured to determine at least one task execution scheme according to at least one of the shared inputs and the plurality of specialized computation graphs corresponding to each of the sub-graphs, where each task execution scheme is used to execute at least one user request, and includes a target specialized computation graph corresponding to each of the sub-graphs;

[0151] A computation graph optimization module 64, configured to perform similarity optimization on the task processing computation graph according to each of the task execution schemes to obtain an optimized processing computation graph for executing the corresponding user request.

[0152] In a possible implementation manner, the task processing computation graph includes a plurality of operators, and an input tensor and an output tensor of each of the operators. The input tensor includes at least one input data that needs to be processed by the operator, and the output tensor includes the output data obtained by calculating each of the input data in the input tensor through the operator. The sub-graph specialization module 62 is further configured to:

[0153] By performing consistency analysis on the task computation graph, at least one operator with the same optimization enable condition is clustered to obtain a corresponding sub-graph;

[0154] Perform sub-graph specialization based on the input tensors of each of the sub-graphs to obtain at least one corresponding candidate computation graph, and each of the candidate computation graphs has a corresponding attribute signature, and the attribute signature is used to characterize whether there is redundancy in each of the input tensors corresponding to the candidate computation graph;

[0155] Perform format conversion on the input tensors and output tensors of the operators in each of the candidate computation graphs according to the corresponding attribute signature to obtain the corresponding specialized computation graph.

[0156] In a possible implementation manner, the sub-graph specialization module 62 is further configured to:

[0157] Determine a plurality of attribute signatures according to the number of input tensors of each of the sub-graphs;

[0158] Perform sub-graph specialization on each of the sub-graphs to obtain a computation graph to be screened corresponding to each of the attribute signatures;

[0159] Screen the computation graphs to be screened corresponding to each of the sub-graphs according to a preset computation graph screening rule to obtain at least one corresponding candidate computation graph.

[0160] In a possible implementation manner, the computation graph screening rule includes at least one of the following:

[0161] Delete the computation graph to be screened that conflicts with the attribute signature;

[0162] In response to the existence of an inclusion relationship between two computation graphs to be screened, delete the included computation graph to be screened;

[0163] Delete the computation graph to be screened in which all input tensors have redundancy.

[0164] In a possible implementation manner, the sub-graph specialization module 62 is further configured to:

[0165] For each of the candidate computation graphs, perform redundancy propagation calculation on each of the operators included therein according to the corresponding attribute signature to obtain the redundancy attributes of the input tensors and output tensors of each operator;

[0166] Perform format conversion on the input tensors and output tensors of the operator according to the redundancy attributes of the input tensors and output tensors of each operator to obtain the corresponding specialized computation graph.

[0167] In a possible implementation manner, the sub-graph specialization module 62 is further configured to:

[0168] Determine the redundancy propagation function of each operator, and the redundancy propagation function is used to determine the redundancy attribute of the output tensor after calculation according to the redundancy attribute of the input tensor before calculation;

[0169] For each of the candidate computation graphs, determine the redundant attributes of the input tensors of the first operator therein according to the corresponding attribute signatures;

[0170] Starting from the first operator in the candidate computation graph, sequentially calculate the redundant attributes of the input tensors and output tensors of each operator therein according to the redundant attributes of the input tensors and the redundant propagation function.

[0171] In a possible implementation, the format conversion includes:

[0172] In response to the redundant attributes of the input tensors and output tensors of the operator being redundant, perform dimensionality reduction processing on the input tensors and the output tensors.

[0173] In a possible implementation, the task processing computation graph includes multiple operators and the input tensors of each operator. The sub-graph specialization module 62 is further configured to:

[0174] For each of the specialized computation graphs, when the number of input tensors of the operator is greater than 1 and there are at least two input tensors with different shapes, determine that the operator is an irregular operator;

[0175] Perform batch processing optimization on the irregular operator to obtain an optimized specialized computation graph.

[0176] In a possible implementation, the sub-graph specialization module 62 is further configured to:

[0177] Determine the type of each irregular operator;

[0178] Perform corresponding batch processing optimization according to the type of the irregular operator to obtain an optimized specialized computation graph.

[0179] In a possible implementation, the sub-graph specialization module 62 is further configured to:

[0180] In response to the type of the irregular operator being a data sharing operator, add corresponding shape conversion operators before and after the irregular operator according to the shapes of the input tensors, for converting the input tensors of the irregular operator into regular input tensors and performing reverse conversion on the output tensors of the irregular operator;

[0181] In response to the type of the irregular operator being a data-independent operator, split the input tensors to process each input data in the input tensors in parallel through the data-independent operator.

[0182] In a possible implementation, the scheme generation module 63 is further configured to:

[0183] Group the multiple user requests according to at least one of the shared inputs to obtain at least one task splitting scheme. Each task splitting scheme includes at least one user request group, and the user requests in the user request groups with more than one user request have shared inputs.

[0184] For each task splitting scheme, determine the scheme input tensors with redundant attributes as redundant according to the shared inputs in each user request group, and determine the scheme input tensors with non-redundant attributes as non-redundant according to other initial data.

[0185] According to the attribute signatures of the specialized computation graphs corresponding to each subgraph, and the redundant attributes of the scheme input tensors in the execution order, sequentially select the target specialized computation graphs corresponding to each subgraph to generate the corresponding candidate execution schemes.

[0186] According to the multiple candidate execution schemes corresponding to each task splitting scheme, select the multiple candidate execution schemes corresponding to an optimal task splitting scheme as the task execution scheme.

[0187] In a possible implementation manner, the computation graph optimization module 64 is further configured to:

[0188] For each task execution scheme, replace the subgraphs in the task processing computation graph with the target specialized computation graphs corresponding to each subgraph therein to obtain an optimized processing computation graph for executing the corresponding user requests.

[0189] In some embodiments, the functions or modules included in the device provided in the embodiments of the present disclosure can be used to execute the methods described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.

[0190] The embodiments of the present disclosure also propose a computer-readable storage medium, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the above methods are implemented. The computer-readable storage medium can be a volatile or non-volatile computer-readable storage medium.

[0191] The embodiments of the present disclosure also propose an electronic device, including: a processor; a memory for storing processor-executable instructions; wherein, the processor is configured to implement the above methods when executing the instructions stored in the memory.

[0192] The embodiments of the present disclosure also provide a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in the processor of an electronic device, the processor in the electronic device executes the above methods.

[0193] Figure 7 FIG. shows a schematic diagram of an electronic device 1900 according to an embodiment of the present disclosure. For example, the electronic device 1900 may be provided as a server or a terminal device. Referring to Figure 7 , the electronic device 1900 includes a processing component 1922, which further includes one or more processors, and memory resources represented by a memory 1932 for storing instructions executable by the processing component 1922, such as application programs. The application programs stored in the memory 1932 may include one or more modules each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute instructions to perform the above method.

[0194] The electronic device 1900 may further include a power supply component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and an input / output interface 1958 (I / O interface). The electronic device 1900 may operate based on an operating system stored in the memory 1932, such as Windows Server TM , Mac OS X TM , Unix TM , Linux TM , FreeBSD TM or the like.

[0195] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as the memory 1932 including computer program instructions, and the above computer program instructions can be executed by the processing component 1922 of the electronic device 1900 to complete the above method.

[0196] The present disclosure may be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.

[0197] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium may be, for example—but not limited to—an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device such as a punched card or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium as used herein is not construed as an instantaneous signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse through an optical fiber cable), or an electrical signal transmitted through a wire.

[0198] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to respective computing / processing devices, or can be downloaded to an external computer or an external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include a copper transmission cable, an optical fiber transmission, a wireless transmission, a router, a firewall, a switch, a gateway computer, and / or an edge server. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.

[0199] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine - related instructions, microcode, firmware instructions, state - setting data, or source code or object code written in any combination of one or more programming languages, including object - oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as the "C" language or similar programming languages. The computer - readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand - alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via an Internet service provider through the Internet). In some embodiments, by using the state information of the computer - readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field - programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer - readable program instructions to implement various aspects of the present disclosure.

[0200] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer - readable program instructions.

[0201] These computer - readable program instructions can be provided to a processor of a general - purpose computer, a special - purpose computer, or other programmable data - processing apparatus to produce a machine such that the instructions, when executed by the processor of the computer or other programmable data - processing apparatus, create a means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer - readable program instructions can also be stored in a computer - readable storage medium, which causes a computer, a programmable data - processing apparatus, and / or other devices to operate in a particular manner, so that the computer - readable medium storing the instructions comprises a manufacture, which includes instructions for implementing various aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0202] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices, causing a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other devices to generate a computer-implemented process, such that the instructions executed on the computer, other programmable data processing apparatus, or other devices implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram.

[0203] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur out of the order noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending upon the functionality involved. It should also be noted that each block of the block diagrams and / or flowchart, and combinations of blocks in the block diagrams and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified function or act, or by a combination of dedicated hardware and computer instructions.

[0204] The embodiments of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or improvements made to the technology in the market, or to enable other ordinary skilled persons in the art to understand the embodiments disclosed herein.

Claims

1. A similarity calculation optimization method for tensor programs, characterized in that The method includes: Obtaining a plurality of user requests, each user request including at least one initial data, and each initial data having a corresponding position in the user request; Determining the initial data located at the same position in at least two of the user requests as the shared input; Determining a plurality of subgraphs obtained by splitting a task processing computational graph, and a plurality of specialized computational graphs corresponding to each of the subgraphs; Determining at least one task execution plan according to at least one of the shared inputs and the plurality of specialized computational graphs corresponding to each of the subgraphs, each task execution plan being used to execute at least one user request, including the target specialized computational graph corresponding to each of the subgraphs; Performing similarity optimization on the task processing computational graph according to each of the task execution plans to obtain an optimized processing computational graph for executing the corresponding user request.

2. The method according to claim 1, characterized in that, The task processing computational graph includes a plurality of operators, and an input tensor and an output tensor of each operator. The input tensor includes at least one input data to be processed by the operator, and the output tensor includes the output data obtained by calculating each of the input data in the input tensor through the operator. The determining a plurality of subgraphs obtained by splitting the task processing computational graph, and a plurality of specialized computational graphs corresponding to each of the subgraphs, includes: By performing consistency analysis on the task processing computational graph, clustering at least one operator with the same optimization enabling condition to obtain a corresponding subgraph; Performing subgraph specialization according to the input tensor of each subgraph to obtain at least one corresponding candidate computational graph, and each candidate computational graph having a corresponding attribute signature, where the attribute signature is used to characterize whether there is redundancy in each input tensor corresponding to the candidate computational graph; Performing format conversion on the input tensor and output tensor of the operator in each candidate computational graph according to the corresponding attribute signature to obtain a corresponding specialized computational graph.

3. The method according to claim 2, characterized in that, The performing subgraph specialization according to the input tensor of each subgraph to obtain at least one corresponding candidate computational graph includes: Determining a plurality of attribute signatures according to the number of input tensors of each subgraph; Performing subgraph specialization on each subgraph to obtain a to-be-screened computational graph corresponding to each attribute signature; Screening the to-be-screened computational graphs corresponding to each subgraph according to a preset computational graph screening rule to obtain at least one corresponding candidate computational graph.

4. The method according to claim 3, characterized in that, The computational graph screening rule includes at least one of the following: Deleting the to-be-screened computational graph conflicting with the attribute signature; In response to the existence of an inclusion relationship between two to-be-screened computational graphs, deleting the included to-be-screened computational graph; Deleting the to-be-screened computational graph in which all input tensors have redundancy.

5. The method according to any one of claims 2-4, characterized in that, The performing format conversion on the input tensor and output tensor of the operator in each candidate computational graph according to the corresponding attribute signature to obtain a corresponding specialized computational graph includes: For each candidate computational graph, performing redundancy propagation calculation on each operator included therein according to the corresponding attribute signature to obtain the redundancy attributes of the input tensor and output tensor of each operator; According to the redundancy attributes of the input tensor and output tensor of each operator, perform format conversion on the input tensor and output tensor of the operator to obtain a corresponding specialized computation graph.

6. The method according to claim 5, characterized in that For each of the candidate computation graphs, perform redundancy propagation calculation on each operator included therein according to the corresponding attribute signature to obtain the redundancy attributes of the input tensor and output tensor of each operator, including: Determine the redundancy propagation function of each operator, where the redundancy propagation function is used to determine the redundancy attribute of the output tensor after calculation according to the redundancy attribute of the input tensor before calculation; For each of the candidate computation graphs, determine the redundancy attribute of the input tensor of the first operator therein according to the corresponding attribute signature; Starting from the first operator in the candidate computation graph, calculate the redundancy attributes of the input tensor and output tensor of each operator therein in sequence according to the redundancy attribute of the input tensor and the redundancy propagation function.

7. The method according to claim 5, characterized in that The format conversion includes: In response to the redundancy attributes of the input tensor and output tensor of the operator being redundant, perform dimensionality reduction processing on the input tensor and the output tensor.

8. The method according to any one of claims 2 to 4, characterized in that The task processing computation graph includes multiple operators and the input tensor of each operator. The determination of the multiple subgraphs obtained by splitting the task processing computation graph and the multiple specialized computation graphs corresponding to each subgraph further includes: For each of the specialized computation graphs, when the number of input tensors of the operator is greater than 1 and there are at least two input tensors with different shapes, determine that the operator is an irregular operator; Perform batch processing optimization on the irregular operator to obtain an optimized specialized computation graph.

9. The method according to claim 8, wherein The performing of batch processing optimization on the irregular operator to obtain an optimized specialized computation graph includes: Determine the type of each of the irregular operators; Perform corresponding batch processing optimization according to the type of the irregular operator to obtain an optimized specialized computation graph.

10. The method according to claim 9, characterized in that, The performing of corresponding batch processing optimization according to the type of the irregular operator includes: In response to the type of the irregular operator being a data sharing operator, add corresponding shape conversion operators before and after the irregular operator according to the shape of the input tensor, for converting the input tensor of the irregular operator into a regular input tensor and performing reverse conversion on the output tensor of the irregular operator; In response to the type of the irregular operator being a data independent operator, split the input tensor to parallel process each input data in the input tensor through the data independent operator.

11. The method according to any one of claims 1 to 4, characterized in that The determination of at least one task execution plan according to at least one of the shared inputs and the multiple specialized computation graphs corresponding to each subgraph includes: Group the multiple user requests according to at least one of the shared inputs to obtain at least one task splitting plan, where each task splitting plan includes at least one user request group, and the user requests in the user request group with the number of user requests greater than 1 have shared inputs; For each of the task splitting plans, determine the plan input tensors with redundant attributes according to the shared inputs in each user request group therein, and determine the plan input tensors with non-redundant attributes for the other initial data. Specialize the attribute signature of the computation graph according to each of the subgraphs, and sequentially select the target specialized computation graph corresponding to each of the subgraphs according to the redundant attributes of the program input tensors in the execution order to generate corresponding candidate execution plans; Select the candidate execution plans corresponding to an optimal task splitting plan from the candidate execution plans corresponding to each of the task splitting plans as the task execution plan.

12. The method according to any one of claims 1 to 4, characterized in that, The similarity optimization of the task processing computation graph according to each of the task execution plans to obtain an optimized processing computation graph for executing the corresponding user request includes: For each of the task execution plans, replace the subgraphs in the task processing computation graph with the target specialized computation graphs corresponding to each of the subgraphs therein to obtain an optimized processing computation graph for executing the corresponding user request.

13. An optimization device for similarity calculation of a tensor program, characterized in that, The apparatus includes: A request acquisition module, configured to acquire a plurality of user requests each including at least one piece of initial data with a corresponding position; A shared input determination module, configured to determine the initial data located at the same position in at least two of the user requests as the shared input; A subgraph specialization module, configured to determine a plurality of subgraphs obtained by splitting a task processing computation graph, and a plurality of specialized computation graphs corresponding to each of the subgraphs; A plan generation module, configured to determine at least one task execution plan according to at least one of the shared inputs and the plurality of specialized computation graphs corresponding to each of the subgraphs, each of the task execution plans being used to execute at least one user request, and including the target specialized computation graph corresponding to each of the subgraphs; A computation graph optimization module, configured to perform similarity optimization on the task processing computation graph according to each of the task execution plans to obtain an optimized processing computation graph for executing the corresponding user request.

14. An electronic device, characterized in that, Comprising: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to implement the method according to any one of claims 1 to 12 when executing the instructions stored in the memory.

15. A non-volatile computer-readable storage medium having computer program instructions stored thereon, characterized in that, The computer program instructions implement the method according to any one of claims 1 to 12 when executed by the processor.

Citation Information

Patent Citations

  • GPU graph neural network optimization method and device

    CN112767230A

  • Deep learning framework module design and construction method for large-format remote sensing image

    CN113869172A