Construction method of routing reverse operator, computer equipment and readable storage medium

By building and optimizing the routing reverse operator, the problem of low hardware computing resource utilization in the backpropagation of the routing operator is solved, and more efficient computing and faster model training are achieved.

CN119940405AActive Publication Date: 2025-05-06SHANGHAI BIREN TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510446318.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-05-06
Estimated Expiration
2045-04-10

AI Technical Summary

Technical Problem

The prior art has the problem of low utilization of hardware computing resources in the backpropagation of routing operators, mainly due to frequent global memory data access and redundant read and write operations.

Method used

By building a new routing inverse operator, determining the inverse submodules it includes, and calling these submodules from the operator library, connecting and mapping into the sub-core of the artificial intelligence chip by dependencies, optimizing the position of the transpose operator to improve the hardware parallelism rate.

Benefits of technology

It reduces access to global memory, avoids unnecessary intermediate result storage and reading, improves the utilization rate of hardware computing resources, improves the computing efficiency of backpropagation, and thus improves the training efficiency of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940405A_ABST
    Figure CN119940405A_ABST
Patent Text Reader

Abstract

The invention relates to a routing reverse operator construction method, computer equipment, a computer readable storage medium and a computer program product. The method comprises the following steps: determining a reverse sub-module included in a routing reverse operator, wherein the reverse sub-module is used for performing gradient calculation; calling each reverse sub-module from an operator library, connecting each reverse sub-module according to a dependency relationship, respectively mapping each reverse sub-module into a corresponding sub-kernel in an artificial intelligence chip, and constructing to obtain a routing reverse operator; wherein the reverse submodule is constructed on the basis of the deduced algorithm formula after deducing the algorithm formula corresponding to the forward submodule, and the forward submodule is a submodule obtained after dividing a routing forward operator. By adopting the method, the utilization rate of hardware resources can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a method for constructing a routing reverse operator, a computer device, a computer-readable storage medium, and a computer program product. Background Art

[0002] In large models, the Mixture of Experts (MoE) mode splits a complex model into multiple expert sub-models. These expert sub-models are relatively independent modules, each with its own professional capabilities, and can handle different types of tasks or data. MoE is an AI model based on the Transformer architecture, which replaces each feedforward network layer in the traditional Transformer model with an MoE layer, where the gating network (Router) is a key component in the Mixture of Experts mode. Its main function is to decide which expert sub-model to send the input tokens to for processing, based on the tokens.

[0003] When implemented on an artificial intelligence chip (e.g., GPGPU (General-purpose computing on graphics processing units)), the gated network can be called a Router operator (also known as a routing operator). Its typical structure includes single operators such as softmax and topk, which have fewer internal calculations but more memory accesses, and are memory-intensive in general. In different deep learning frameworks, the Router function is usually implemented by connecting the required single operators.

[0004] The reverse operator of the Router (also called the routing reverse operator) is usually implemented through the automatic reverse derivation of the deep learning framework, that is, for each single operator in the Router, the reverse gradient is transferred according to the chain rule. This method requires frequent data access to the global memory, there are a large number of redundant read and write operations, and the utilization rate of hardware computing resources is low. Summary of the invention

[0005] Based on this, it is necessary to provide a method for constructing a routing reverse operator, a computer device, a computer-readable storage medium and a computer program product that can improve the utilization of hardware computing resources in the reverse propagation of the routing operator in response to the above technical problems.

[0006] In a first aspect, the present application provides a method for constructing a routing reverse operator, the method comprising:

[0007] Determine a reverse submodule included in a route reverse operator, wherein the reverse submodule is used to perform gradient calculation;

[0008] Calling each of the reverse submodules from the operator library, connecting each of the reverse submodules according to the dependency relationship, and mapping each of the reverse submodules to the corresponding sub-kernel in the artificial intelligence chip to construct a routing reverse operator;

[0009] The reverse submodule is constructed based on the derived algorithm formula after the algorithm formula corresponding to the forward submodule is derived, and the forward submodule is a submodule obtained by dividing the routing forward operator.

[0010] In one embodiment, the method further comprises:

[0011] Divide the routing forward operator into modules to obtain multiple forward sub-modules;

[0012] For any of the forward submodules, perform the following operations:

[0013] According to the algorithm formula of the forward submodule, determine the algorithm formula of the reverse submodule corresponding to the forward submodule;

[0014] Based on the algorithm formula of the reverse submodule, the operators included in the reverse submodule and the dependencies between the operators are determined, and according to the operators and the dependencies between the operators, the reverse submodule is constructed.

[0015] In one embodiment, the input tensors, intermediate tensors and output tensors of the routing reverse operator are all stored in the global memory, and the internal tensors of the reverse submodule are stored in the designated memory, which is the memory corresponding to the sub-kernel that executes the reverse submodule.

[0016] In one embodiment, the reverse submodule included in the determining route reverse operator includes:

[0017] Acquire configuration parameters of the route reverse operator, and determine candidate reverse submodules whose configuration attributes in the configuration parameters are first attribute values;

[0018] The main path loss sub-template and each of the candidate reverse sub-modules are determined as the reverse sub-modules of the routing reverse operator.

[0019] In one embodiment, before mapping each of the reverse submodules to a corresponding sub-core in the artificial intelligence chip to construct a route reverse operator, the method further includes:

[0020] Based on the computing resources of the artificial intelligence chip, determining a position adjustment strategy of the transposition operator in the routing reverse operator;

[0021] The position of the transposition operator in the routing reverse operator is adjusted based on the position adjustment strategy, and the designated dimension of the first target operator in the reverse submodule is re-designated according to the position adjustment strategy.

[0022] In one embodiment, determining the position adjustment strategy of the transposition operator in the routing reverse operator based on the computing resources of the artificial intelligence chip includes:

[0023] Determining a first hardware parallelism rate of the artificial intelligence chip based on the non-specified dimension before adjustment and the computing resources of the artificial intelligence chip;

[0024] After swapping the designated dimension and the non-designated dimension, determining a second hardware parallelism rate of the artificial intelligence chip based on the swapped non-designated dimension and the computing resources of the artificial intelligence chip;

[0025] When the first hardware parallelism rate is less than the second hardware parallelism rate, a target position of a transposition operator in the route reversal operator is determined, and a position adjustment strategy is generated based on the target position of the transposition operator in the route reversal operator.

[0026] In one embodiment, determining a target position of a transposition operator in the route reversal operator includes:

[0027] Determining a second target operator from the route reverse operator, wherein the second target operator is an operator whose operation is dimension-independent;

[0028] Based on the position of the second target operator, a target position of the transposition operator in the route reversal operator is determined.

[0029] In one embodiment, the routing forward operator is divided into modules to obtain multiple forward sub-modules, including:

[0030] Get the routing forward operator graph corresponding to the routing forward operator;

[0031] Dividing the routing forward operator graph into sub-modules to obtain multiple operator graphs;

[0032] According to the multiple operator graphs, multiple forward submodules are obtained.

[0033] In a second aspect, the present application further provides a device for constructing a routing reverse operator, the device comprising:

[0034] A first determining module, used to determine a reverse submodule included in a route reverse operator, wherein the reverse submodule is used to perform gradient calculation;

[0035] The first construction module is used to call each of the reverse submodules from the operator library, connect each of the reverse submodules according to the dependency relationship, and then map each of the reverse submodules to the corresponding sub-kernel in the artificial intelligence chip to construct a routing reverse operator;

[0036] The reverse submodule is constructed based on the derived algorithm formula after the algorithm formula corresponding to the forward submodule is derived, and the forward submodule is a submodule obtained by dividing the routing forward operator.

[0037] In one embodiment, the device further comprises:

[0038] A partitioning module is used to partition the routing forward operator into modules to obtain multiple forward sub-modules;

[0039] For any of the forward submodules, perform the following operations:

[0040] A second determination module, used to determine the algorithm formula of the reverse submodule corresponding to the forward submodule according to the algorithm formula of the forward submodule;

[0041] The second construction module is used to determine the operators included in the reverse submodule and the dependencies between the operators based on the algorithm formula of the reverse submodule, and construct the reverse submodule according to the operators and the dependencies between the operators.

[0042] In one embodiment, the input tensors, intermediate tensors and output tensors of the routing reverse operator are all stored in the global memory, and the internal tensors of the reverse submodule are stored in the designated memory, which is the memory corresponding to the sub-kernel that executes the reverse submodule.

[0043] In one embodiment, the first determination module is specifically configured to:

[0044] Acquire configuration parameters of the route reverse operator, and determine candidate reverse submodules whose configuration attributes in the configuration parameters are first attribute values;

[0045] The main path loss sub-template and each of the candidate reverse sub-modules are determined as the reverse sub-modules of the routing reverse operator.

[0046] In one embodiment, the device further comprises:

[0047] A fourth determination module, used to determine a position adjustment strategy of the transposition operator in the routing reverse operator based on the computing resources of the artificial intelligence chip;

[0048] The optimization module is used to adjust the position of the transposition operator in the routing reverse operator based on the position adjustment strategy, and re-specify the specified dimension of the first target operator in the reverse submodule according to the position adjustment strategy.

[0049] In one embodiment, the optimization module is specifically used for:

[0050] Determining a first hardware parallelism rate of the artificial intelligence chip based on the non-specified dimension before adjustment and the computing resources of the artificial intelligence chip;

[0051] After swapping the designated dimension and the non-designated dimension, determining a second hardware parallelism rate of the artificial intelligence chip based on the swapped non-designated dimension and the computing resources of the artificial intelligence chip;

[0052] When the first hardware parallelism rate is less than the second hardware parallelism rate, a target position of a transposition operator in the route reversal operator is determined, and a position adjustment strategy is generated based on the target position of the transposition operator in the route reversal operator.

[0053] In one embodiment, the optimization module is specifically used for:

[0054] Determining a second target operator from the route reverse operator, wherein the second target operator is an operator whose operation is dimension-independent;

[0055] Based on the position of the second target operator, a target position of the transposition operator in the route reversal operator is determined.

[0056] In one embodiment, the partitioning module is specifically used for:

[0057] Get the routing forward operator graph corresponding to the routing forward operator;

[0058] Dividing the routing forward operator graph into sub-modules to obtain multiple operator graphs;

[0059] According to the multiple operator graphs, multiple forward submodules are obtained.

[0060] In a third aspect, the present application further provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor implements any of the above methods for constructing a routing reverse operator when executing the computer program.

[0061] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements any of the above methods for constructing a routing reverse operator.

[0062] In a fifth aspect, the present application also provides a computer program product, including a computer program, which, when executed by a processor, implements any of the above methods for constructing a routing reverse operator.

[0063] The construction method, computer device, computer-readable storage medium and computer program product of the above-mentioned routing reverse operator are obtained by constructing the routing reverse operator based on the reverse submodule in the operator library, and the reverse submodule is constructed based on the algorithm formula derived from the algorithm formula corresponding to the forward submodule. This specially derived reverse operator formula enables the reverse submodule to reduce certain transitional intermediate variables in the forward calculation through derivation. In this way, in the calculation of reverse propagation, there is no need to re-read and store this part of the data, which reduces the access to these data in the global memory, that is, when performing reverse propagation, there is no need to perform gradient calculation on the entire operator graph of the routing operator, which reduces unnecessary storage and reading of intermediate results, can greatly improve the utilization rate of hardware computing resources, improve the calculation efficiency of reverse propagation, and thus improve the training efficiency of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the drawings required for use in the embodiments of the present application or related technical descriptions will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0065] Figure 1 is a forward propagation operator graph of a routing operator in a MoE model in one embodiment;

[0066] Figure 2 A schematic diagram of a process of constructing a reverse routing operator in one embodiment;

[0067] Figure 3 A schematic diagram of a process of constructing a reverse routing operator in one embodiment;

[0068] Figure 4 is a flow chart of step 302 in another embodiment;

[0069] Figure 5 An operator graph of a forward submodule obtained by dividing a routing operator in one embodiment;

[0070] Figure 6a is an operator graph of a routing reverse operator in one embodiment;

[0071] Figure 6b is an operator graph of a main path submodule in one embodiment;

[0072] Figure 6c is an operator graph of a z-loss loss submodule in one embodiment;

[0073] Figure 6d is an operator graph of a load balancing submodule in one embodiment;

[0074] Figure 7 is a flow chart of step 202 in one embodiment;

[0075] Figure 8 A schematic diagram of a flow chart of a method for constructing a reverse routing operator in another embodiment;

[0076] Fig. 9 is a flow chart of step 802 in one embodiment;

[0077] Fig.10 is a flow chart of step 906 in one embodiment;

[0078] Fig.11a An operator graph after optimization of a main path submodule in one embodiment;

[0079] Fig.11b It is an operator graph after optimization of the z-loss loss submodule in one embodiment;

[0080] Fig.11c An operator graph after optimization of a load balancing submodule in one embodiment;

[0081] Fig.12 A schematic diagram of the structure of a GPGPU in one embodiment;

[0082] Fig.13 is a structural block diagram of a device for constructing a routing reverse operator in one embodiment;

[0083] Fig.14 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0084] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0085] In the Mixture of Experts (MoE), the routing operator is a key component, which is usually implemented with the help of lightweight neural networks (such as linear layers) or other structures. Different MoE models have some subtle differences in the specific design of routing operators. Figure 1 As shown in Figure 1, it shows the forward propagation process of the routing operator in a typical MoE model. Figure 1In the forward operator of the routing shown, the input tensor (hereinafter referred to as Input) has a shape of [b, s / T, e], which is the input of the entire calculation process. Among them, b represents the batch size, that is, the number of data samples processed in one calculation; s is the sequence length seqlen, T is the total number of tokens, and e is the number of experts expert num.

[0086] It should be noted that the routing operators and the shape expressions of different tensors in the embodiments of the present application are only implemented as an example in the embodiments of the present application, and are not to be understood as a limitation on the shape expressions of routing operators and tensors. In fact, for routing operators of different structures, there is a corresponding relationship between the shape expressions of their tensors and the parameters of the routing operators, which is not specifically limited in the embodiments of the present application.

[0087] The weight tensor (hereinafter referred to as weight) enters the gating function together with the input. The weight tensor is used in deep learning to adjust the model's emphasis on different input features. Among them, gating operates on the input and weight, and outputs the logistic regression value logits tensor 1 (hereinafter referred to as logits1), with a shape of [b, s / T, e]. Logits is the original predicted value that has not been processed by the activation function, and will be used to calculate probability, etc.

[0088] The view function (hereinafter referred to as view) performs a view transformation operation on logits1, changes the shape of the data, and outputs the logistic regression value logits tensor 2 (hereinafter referred to as logits2), with a shape of [b×s, e]. The seq_gather function is a sequence parallel gather function, where gather is a basic function used to obtain data of a specified dimension (dim) and a specified index (index) from the original tensor. It processes logits2 to obtain the logistic regression value logits tensor 3 (hereinafter referred to as logits3). This operation includes extracting information about specific locations or features from sequence data.

[0089] The topk function is used to obtain the largest or smallest k elements and index positions in a tensor or array. For example, it performs a topk operation on logits3 to find the k elements with the highest probability and index positions in each sample, and outputs a score value of 1 (hereinafter referred to as scores1) with a shape of [b×s, topk]. After the topk operation finds the k elements with the highest probability in each sample, the argsort function sorts these elements and generates an index (hereinafter referred to as indices) with a shape of [b×s, topk]. Indices records the index position information of the relevant elements in the original data after the topk operation.

[0090] The normalized exponential function (hereinafter referred to as softmax) function applies the softmax activation function to the result of the topk operation, converting scores1 into probability values ​​(hereinafter referred to as probs) with a shape of [b×s, topk]. The scatter function is a basic function that transfers data from one tensor to another according to the index. It transfers probs to another tensor according to the index obtained by indices, and obtains the topk_masked_gates tensor (hereinafter referred to as topk_masked_gates), which is a mask tensor used to mark the position of the topk elements.

[0091] The topk function performs the topk operation on the input topk_masked_gates, where k=capacity, and the operation is performed along the dimension dim=0. This operation finds the capacity (capacity is the expert capacity) elements with the highest probability in each dimension and outputs two results: the capacity_indices tensor (hereinafter referred to as capacity_indices), which represents the capacity index: the shape is [capacity, e], which records the index of the capacity highest probability elements found in topk_masked_gates; the capacity_probs tensor (hereinafter referred to as capacity_probs), which represents the capacity probability: the shape is [capacity, e], which represents the probability value of the corresponding capacity highest probability elements.

[0092] Perform a transpose operation on capacity_probs to swap the order of its dimensions and obtain the final probability final_probs tensor (hereinafter referred to as final_probs), with a shape of [e, capacity], which is the probability value of the final output. Perform a transpose operation on capacity_indices to obtain the final probability final_indices tensor (hereinafter referred to as final_indices), with a shape of [e, capacity], which is used to indicate the index value related to the probability value of the final output.

[0093] In the z_loss loss function, the smooth maximum function (hereinafter referred to as logsumexp) performs a logarithmic sum exponential operation on logits3 along the specified dimension (dim=1), and the output shape is [b×s,1]; the square function (hereinafter referred to as square) squares the result of logsumexp, and the output shape is [b×s,1]; the average function (hereinafter referred to as mean) calculates the average of the result of square, and is transformed by the coefficient transformation function (e×moe_aux_loss_coeff) to obtain the z_loss loss with a shape of [1,1].

[0094] In aux_loss_load_balancing (load balancing loss), the softmax function applies the softmax activation function to logits3 again, with the data type specified as fp32, and outputs score 2 (hereinafter referred to as scores2), with a shape of [b×s, e]. The mean function calculates the average value of scores2 along the dimension dim=0 to obtain the probs_mean_per_expert tensor (hereinafter referred to as probs_mean_per_expert), with a shape of [1, e]. The scatter function scatters the data according to indices to generate a mask topk_mask tensor (hereinafter referred to as topk_mask), with a shape of [b×s, e]. The sum function (hereinafter referred to as sum) sums the topk_mask along the dimension dim=0 to obtain the tokens_per_expert tensor (hereinafter referred to as tokens_per_expert), with a shape of [1, e]. This part of the data is related to the number of tokens processed by each expert and is used for load balancing calculations. Finally, probs_mean_per_expert and tokens_per_expert are dot-multiplied and summed by the mean dot multiplication + sum function (hereinafter referred to as dot mul+sum), and transformed by the coefficient transformation function (e×moe_aux_loss_coeff), and finally the load balancing loss (hereinafter referred to as aux_loss, also called aux loss auxiliary loss function) is obtained, with a shape of [1, 1].

[0095] That is, the routing operator involves multiple links such as input processing, probability calculation, loss function calculation, and final result generation, and these operations serve the training and reasoning process of the model together. The input, output, and intermediate data of the above multiple links are all stored in the global memory. When the chain rule is used in the related technology to implement the back propagation process of the routing operator on GPGPU, the gradient calculation of the entire operator graph of the routing operator is usually performed according to fixed rules, which requires frequent read and write operations of a large amount of data to the global memory, and some of the same intermediate results may be calculated and stored multiple times. A large number of redundant read and write operations will lead to low utilization of hardware computing resources.

[0096] The embodiment of the present application provides a method for constructing a routing reverse operator for gradient calculation during back propagation of a routing operator. The routing reverse operator constructed based on the embodiment of the present application does not need to perform gradient calculation on the entire operator graph of the routing operator during back propagation, thereby reducing unnecessary storage and reading of intermediate results, greatly improving the utilization rate of hardware computing resources, improving the computing efficiency of back propagation, and thereby improving the training efficiency of the model.

[0097] In one embodiment, Figure 2 As shown, a method for constructing a reverse routing operator is provided. This embodiment uses the method applied to the host side as an example. It can be understood that the host side may include a CPU (Central Processing Unit). In this embodiment, the method includes the following steps 202 to 204, wherein:

[0098] Step 202: determine a reverse submodule included in the route reverse operator, where the reverse submodule is used to perform gradient calculation.

[0099] In the embodiment of the present application, the forward routing operator is used for the forward propagation of the routing operator (ie, Router), and the reverse routing operator is used for the reverse propagation of the routing operator. The reverse routing operator includes at least one reverse submodule, which is an operator module used for gradient calculation in the reverse propagation. Figure 1 The embodiment of the present application is described by taking the shown routing forward operator as an example.

[0100] In the embodiment of the present application, the routing forward operator can be divided into modules in advance to obtain at least one forward sub-module, and reverse deduction can be performed based on the algorithm formula of each forward sub-module to obtain the algorithm formula of the reverse sub-module corresponding to each forward sub-module, and each reverse sub-module can be constructed based on the algorithm formula of each reverse sub-module and stored in the operator library.

[0101] For example, the routing forward operator can be analyzed to determine how the routing forward operator assigns weights to different expert networks based on input data, and the decision-making role it plays in the forward propagation process of the entire model. For example, analyze what features (such as the dimension of the input vector, numerical distribution, etc.) the routing forward operator is based on to select experts and assign weights. Based on the analysis results, identify the parts of the routing forward operator that have relatively independent functions, which usually correspond to different calculation steps or logic units. For example: refer to the aforementioned Figure 1 The example routing forward operator includes an input feature extraction subfunction, which is used to extract key features for routing decisions from the original input data, which may involve operations such as linear transformation and nonlinear activation; a weight calculation subfunction, which is used to calculate the assigned weights of each expert network based on the extracted features, which may use the softmax function, attention mechanism, etc.; an expert selection subfunction, which is used to select the expert network participating in the subsequent calculation based on the calculated weights, which may use the topk function, etc.

[0102] According to the identified sub-functions, the routing forward operator can be divided into at least one forward sub-module. It should be noted that when dividing, it is necessary to ensure that each forward sub-module has clear inputs and outputs and relatively independent functions. For example, Figure 1 The routing forward operator shown in the figure has the following inputs: the input is the original input data, and the output is the extracted feature vector. The weight calculation module has the following inputs: the input is the extracted feature vector, and the output is the assigned weight for each expert network. The expert selection module has the following inputs: the input is the expert assigned weight, and the output is the selected expert network index.

[0103] Furthermore, for each forward submodule, reverse deduction can be performed based on its algorithm formula to obtain the algorithm formula of each reverse submodule in the back propagation. Exemplarily, the algorithm formula of each forward submodule can be determined, and the variables, constants and operational relationships involved in the formula can be clarified. For example: for the input feature extraction module, the linear transformation y=Wx+b is often used (where x is the input vector, W is the weight matrix, b is the bias vector, and y is the output feature vector), and the meaning and dimension of each parameter must be clarified.

[0104] After determining the algorithm formula of each forward submodule, the chain rule can be used for reverse deduction to obtain the reverse algorithm formula. For example: for the linear transformation y=Wx+b, assuming the loss function is L, according to the chain rule, the reverse algorithm formula can be deduced as .

[0105] After the algorithm formula is obtained by reverse deduction for each forward submodule, the corresponding reverse submodule can be constructed based on the algorithm formula and stored in the operator library. In this way, when the routing reverse operator is implemented later, the reverse submodule included in the routing reverse operator can be determined based on the forward operator module included in the routing forward operator, and each reverse submodule can be called from the operator library to implement the routing reverse operator.

[0106] The process of deriving the reverse submodule will be introduced in detail below.

[0107] In an exemplary embodiment, referring to Figure 3 As shown, the process of deriving the reverse submodule may include the following steps 302 to 306, wherein:

[0108] Step 302, dividing the routing forward operator into modules to obtain multiple forward sub-modules;

[0109] For any forward submodule, do the following:

[0110] Step 304, determining the algorithm formula of the reverse submodule corresponding to the forward submodule according to the algorithm formula of the forward submodule;

[0111] Step 306: Based on the algorithm formula of the reverse submodule, determine the operators included in the reverse submodule and the dependencies between the operators, and construct the reverse submodule according to the dependencies between the operators.

[0112] The typical routing forward operator mainly includes the following core structures: softmax module, topk module, load balancing loss module and z-loss loss module, among which:

[0113] Softmax module: The router module calculates the scores of all experts for each input data. In order to reasonably assign the input to the appropriate experts, these scores need to be normalized, and this task is completed by softmax. After softmax normalization, the input data will be assigned to the expert with the highest score for subsequent calculations. This process helps to assign reasonable weights to each expert based on the input, so that the model can select the most matching expert for processing based on the input features.

[0114] Topk module: In a large-scale MoE model, if all experts are involved in the calculation, it will bring huge computing overhead. In order to effectively reduce the computing cost, the routing module will use topk. Topk will select the k experts with the highest scores from all experts for activation, and the unselected experts will not participate in the subsequent calculation process. In this way, while ensuring the performance of the model, the unnecessary amount of calculation is significantly reduced, and the calculation efficiency of the model is improved.

[0115] Load balancing loss module: In order to ensure load balance among experts and avoid some experts being overused while others are idle, the MoE model usually introduces a load balancing term in the loss function. The role of the load balancing loss is to calculate this load balancing loss and impose constraints on the routing module to encourage it to distribute the input data as evenly as possible to different experts, thereby giving full play to the capabilities of each expert and improving the overall generalization performance of the model.

[0116] z-loss loss module: In order to further assist the training process of the routing module and ensure the stability and diversity of expert selection, some MoE models will introduce other loss functions, one of which is z-loss. z-loss is a commonly used regularization loss used to assist in training the Router to ensure that the selection of experts is more stable and diverse. The main purpose is to alleviate the problem of unreasonable distribution of Router outputs, avoid over-reliance on a few experts or extreme allocation probabilities. The core idea is to avoid excessively extreme distribution of routing outputs by penalizing excessively large or small probability values. z-loss can constrain the routing module from different angles, making the model more comprehensive and reasonable when selecting experts, avoiding over-reliance on certain specific experts, thereby enhancing the robustness and generalization ability of the model.

[0117] These modules work together to form a complete routing forward operator. Based on this, the routing forward operator can be divided into modules to obtain corresponding multiple forward sub-modules.

[0118] In an exemplary embodiment, referring to Figure 4 As shown, in step 302, the routing forward operator is divided into modules to obtain multiple forward sub-modules, which may include the following steps 402 to 406, wherein:

[0119] Step 402, obtaining a routing forward operator graph corresponding to the routing forward operator;

[0120] Step 404, dividing the routing forward operator graph into sub-modules to obtain multiple operator graphs;

[0121] Step 406, obtaining multiple forward submodules according to the multiple operator graphs.

[0122] In the embodiment of the present application, a routing forward operator graph corresponding to the routing forward operator can be obtained, and the routing forward operator graph can be further divided into sub-modules based on the structure of the routing forward operator to obtain multiple operator graphs. Figure 1 After dividing the routing forward operator into submodules, the division result can be shown as follows: Figure 5 As shown, Figure 5 Each dotted box in the figure corresponds to an operator graph of a submodule.

[0123] The routing forward operator graph is divided into multiple operator graphs, each of which corresponds to a forward submodule. Taking the typical Router operator structure used in the switch transformer model structure as an example, the forward submodules included in the routing forward operator graph may include a main path loss submodule, a load balancing submodule, and a z-loss loss submodule, which can be expressed as the following formula (I):

[0124] Formula (I)

[0125] in, Characterizes the total loss of the routing forward operator, Characterize the main path loss submodule, Characterize the load balancing loss submodule, Characterize the z-loss loss submodule, Characterizes the weight of the load balancing loss submodule, Characterize the weight of the z-loss submodule. The main path loss submodule is the main part of the routing forward operator, which is mainly composed of single operators such as topk and softmax. Therefore, the main path loss submodule can also be further divided into more submodules, such as softmax submodule, topk submodule, etc., which is not specifically limited in the embodiments of the present application.

[0126] The load balancing loss submodule is usually an implementation of the auxiliary loss auxiliary loss function, which is recorded as the aux loss function in the embodiment of the present application. The load balancing loss submodule is a key regularization term used to optimize the Router in the MoE model. The typical aux loss function forward algorithm formula refers to the following formula (II):

[0127] Formula (II)

[0128] Among them, α is the scaling factor used to control the weight of load balancing loss in the total loss, N is the number of experts, Represents the probability that the Router assigns the current token to expert i, , The result of Router input x after softmax normalization calculation. , T is the total number of Tokens, x is the input of the Router, .

[0129] The typical forward calculation formula of the z-loss function is shown in formula (3).

[0130] Formula (III)

[0131] After obtaining the algorithm formula of each forward submodule, for any forward submodule, the algorithm formula of the corresponding reverse submodule can be inferred based on the algorithm formula of the forward submodule. Figure 5 The reverse process of each forward submodule of the routing forward operator shown is as follows:

[0132] The algorithm formula of the main path loss submodule is reversed as follows: Since the implementation of the main path is different in different routers, the reverse operator of the main path loss operator, with softmax and topk as the key parts, can be fused into multiple sub-paths, that is, fused into multiple reverse sub-modules. In the embodiment of the present application, the reverse process of key operators such as softmax and topk is not elaborated in detail, and the reverse process can be performed based on the chain rule.

[0133] The reverse algorithm formula of the load balancing loss submodule is , for formula (II), according to the chain rule, we can get ,because The computational characteristics of It is not derivable, so we can get .in, , .

[0134] Therefore, the back propagation of the aux loss function is implemented as follows:

[0135] Formula (IV).

[0136] Among them, (meanreduce_bwd) represents the reverse implementation of meanreduce. Meanreduce refers to the function that performs aggregation operations on input tensors. Here, it refers to the above formula for The operation of finding the mean. The reverse implementation of meanreduce requires evenly distributing the incoming gradient back to each element of the input. (softmax_bwd) refers to the reverse implementation of softmax, which can be referred to the general softmax reverse implementation method. The specific derivation formula is not introduced in the embodiment of this application.

[0137] The algorithm formula of the z-loss loss function is reversed as follows: For ease of description, the calculation part of logsumexp in z-loss is recorded as LSE. LSE calculates the exponential sum of the input x and takes the logarithm, and records its square as SLSE. The formula is as follows: , .

[0138] Therefore, the forward formula of z-loss can be expressed as shown in formula (V):

[0139] Formula (V)

[0140] The derivation process of the reverse calculation formula corresponding to z-loss is as follows: ,in, , , .

[0141] Therefore, the back propagation of the z-loss function is implemented as shown in the following formula (VI):

[0142] Formula (VI)

[0143] In this way, by reversely deducing the forward formula of each forward submodule in the Router, the reverse calculation formula of each component in the Router is derived. By analyzing each item in the reverse calculation formula, the operators contained in the reverse submodule and the input and output of the operator are obtained, that is, the dependency relationship between the operators included in each reverse submodule can be obtained, and then the reverse submodule can be constructed based on the dependency relationship between the operators.

[0144] For example, the obtained reverse calculation formula can be disassembled in detail and decomposed into basic operation units, which are the operators contained in the reverse submodule. For example, operations such as multiplication, addition, and transposition in the formula can be regarded as different operators.

[0145] For each operator, its input and output can be determined according to the source and destination of the item in the reverse calculation formula. By analyzing the entire reverse calculation formula, all operators contained in each reverse submodule and their respective inputs and outputs can be determined.

[0146] Based on the input and output of the operator, we can further analyze which operator outputs are the inputs of other operators, thereby determining the dependency between operators. For example, if the input of operator B depends on the output of operator A, then operator B depends on operator A. When building the reverse submodule, operator A must be executed before operator B can be executed.

[0147] A graph structure (such as a directed acyclic graph) can be used to intuitively represent the dependency between operators. The nodes in the graph represent operators, and the directed edges represent the dependency direction between operators, from the output operator to the input operator. In this way, the logical relationship between the operators in the reverse submodule can be clearly displayed.

[0148] Furthermore, by topologically sorting the dependency graph, a linear operator execution order can be obtained, that is, the operator graph of the reverse submodule. The result of topological sorting ensures that when each operator is executed, all its dependent operators have been executed, thereby satisfying the dependency relationship between operators. According to the operator execution order obtained by topological sorting, the reverse submodule can be implemented by writing code. In the code, each operator is called in sequence and the corresponding input data is passed to finally realize the function of the reverse submodule. The operator library provided by the deep learning framework (such as TensorFlow, PyTorch, etc.) can be used to implement specific operator operations. These frameworks usually have optimized and encapsulated various common operators and can be called directly.

[0149] For example, refer to Figure 5The forward routing operator, based on the operator graph of the reverse routing operator constructed by the derived reverse submodules, refers to Figures 6a to 6d As shown. Figure 6a The schematic diagram of the routing reverse operator shown in the figure shows that the routing reverse operator includes a main path submodule, a load balancing loss submodule and a z-loss loss submodule, wherein d-logits-z_loss is the aggregated z_loss logits gradient tensor output by the z-loss loss submodule, with a shape of [b×s,e], and d-logtis-aux_loss is the aggregated aux-loss logits gradient tensor output by the load balancing loss submodule, with a shape of [b×s,e]. For example, the operator graph of the main path submodule is referred to Figure 6b As shown in the figure, it mainly includes the reverse operator of the normalized exponential function (hereinafter referred to as softmax bwd, also known as the normalized exponential reverse function), the reverse operator of the topk function (hereinafter referred to as scatter), and other operators such as view function view, device transpose, and addition add. The operator diagram of the z-loss loss submodule can be found in Figure 6c As shown in Figure 1, this is the operator graph constructed based on the z-loss reverse function derived above. The operator graph of the load balancing loss submodule is shown in Figure 1. Figure 6d As shown, this is the operator graph constructed based on the aux loss reverse function derived above.

[0150] in, Figure 6b The corresponding main path submodule is composed of Figure 5 The main path loss submodule in the reverse is obtained. Figure 6c The corresponding z-loss loss submodule is composed of Figure 5 The derivation result is derived from the z-loss loss submodule in , and the derivation result refers to formula (six), Figure 6d The corresponding load balancing loss submodule is composed of Figure 5 The load balancing loss submodule in is deduced to obtain the derivation result, and the derivation result is described in formula (IV). Figures 6a to 6d , where d-tensor (including d-Input, d-logits2, d-scores1, etc.) refers to the corresponding gradient tensor, for example, d-Input is the input gradient tensor.

[0151] exist Figure 6bIn the corresponding operator graph, since the topk operator is used in the forward propagation, the gradient information needs to be passed back from the output of the topk operation to the input during the backward propagation. Since the topk operation only selects some elements, the output gradient needs to be correctly assigned back to the corresponding position in the input tensor during the backward propagation, and the scatter operator can complete this task. The scatter operator can disperse the gradient information to the corresponding position of the input tensor according to the index information recorded during the forward propagation. Therefore, the scatter operator is used for the topk position in the backward propagation process, that is, the scatter operator is the reverse implementation of the topk operator. The softmax bwd operator is the reverse function of softmax. When the scatter operator is used in forward propagation, the gradient information needs to be passed back from the output of the scatter operation to the input during backward propagation. Since the scatter operation disperses the elements to different positions, the output gradient needs to be collected back to the corresponding position of the input tensor according to the index rule during forward propagation during backward propagation. The gather operator can collect elements at specific positions from the input tensor according to the given index. Therefore, the gather operator is used at the position corresponding to the scatter in the backward propagation process, that is, the gather operator is the reverse implementation of the scatter operator.

[0152] exist Figure 6c In the corresponding operator graph, based on formula (6), we can see that the reverse submodule corresponding to the operator graph includes three operators, one is the reverse implementation of softmax, which is represented as softmax in the operator graph, and the other is 2×LSE i The corresponding representation is logsumexp in the operator graph, and the other is meanreduce_bwd (can also be expressed as reduce_bwd), which corresponds to mean_bwd in the operator graph, where mul is the fusion of the three operators.

[0153] exist Figure 6d In the corresponding operator graph, based on formula (IV), we can see that the reverse submodule corresponding to the operator graph includes three operators, one is the reverse implementation of softmax, which corresponds to softmax_bwd in the operator graph, one is meanreduce_bwd, which corresponds to mean-bwd in the operator graph, and one is the implementation The operator corresponds to the dot mul sum in the operator graph.

[0154] In this way, based on the dependency relationship between tensors of operators, each reverse sub-module is constructed with the goal of reducing redundant memory transfer. Each reverse sub-module can be stored in the operator library for subsequent calls.

[0155] Step 204, calling each reverse sub-module from the operator library, connecting each reverse sub-module according to the dependency relationship, and then mapping each reverse sub-module to the corresponding sub-kernel in the artificial intelligence chip to construct a routing reverse operator; wherein, the reverse sub-module is constructed based on the derived algorithm formula after deriving the algorithm formula corresponding to the forward sub-module, and the forward sub-module is a sub-module obtained by dividing the routing forward operator.

[0156] In the embodiment of the present application, after determining each reverse submodule of the routing operator, each reverse submodule can be called from the operator library, and each reverse submodule can be connected in the order of the dependency relationship between the reverse submodules to construct a routing reverse operator, and further map each reverse submodule in each routing reverse operator to the corresponding sub-kernel in the artificial intelligence chip, and then perform reverse propagation of the routing operator based on the routing reverse operator. For example, if the output of a reverse submodule is the input of another reverse submodule, then connect them according to this logic, and finally obtain a complete routing reverse operator, and map each reverse submodule to the sub-kernel of the artificial intelligence chip, for example, Figure 6b The multiple sub-modules divided by the main path loss sub-module are mapped to multiple sub-kernels (including sub-kernel function 1 (which can be expressed as subKernel-Main1), sub-kernel function 2 (which can be expressed as subKernel-Main2), sub-kernel function 3 (which can be expressed as subKernel-Main3), and sub-kernel function 4 (which can be expressed as subKernel-Main4)), the load balancing loss sub-module is mapped to the sub-kernel (including sub-kernel function-auxloss, which can also be expressed as subKernel-auxloss), and the z-loss loss sub-module is mapped to the sub-kernel (including sub-kernel function-auxloss, which can also be expressed as subKernel-zloss). In this way, the back propagation of the routing operator can be efficiently implemented on the artificial intelligence chip.

[0157] The method for constructing the routing reverse operator provided by the embodiment of the present application is adopted. Since the routing reverse operator is constructed based on the reverse submodule in the operator library, and the reverse submodule is constructed based on the derived algorithm formula after the algorithm formula corresponding to the forward submodule is deduced. This specially derived reverse operator formula enables the reverse submodule to reduce certain transitional intermediate variables in the forward calculation through deduction. In this way, in the calculation of reverse propagation, there is no need to re-read and store this part of the data, which reduces the access to these data in the global memory, that is, when performing reverse propagation, there is no need to perform gradient calculation on the entire operator graph of the routing operator, which reduces unnecessary storage and reading of intermediate results, can greatly improve the utilization rate of hardware computing resources, improve the calculation efficiency of reverse propagation, and thus improve the training efficiency of the model.

[0158] In an exemplary embodiment, the input tensors, intermediate tensors, and output tensors of the routing reverse operator are all stored in the global memory, and the internal tensors of the reverse submodule are stored in the designated memory, wherein the designated memory is the memory corresponding to the sub-kernel that executes the reverse submodule.

[0159] In the present application, refer to Figure 6b to Figure 6d As shown in the figure, the tensor marked as 4 is the input tensor, which is located in the global memory, the tensor marked as 5 is the output tensor, which is located in the global memory, the tensor marked as 6 is the intermediate tensor, which is the data exchanged between different subKernels and is located in the global memory, and the tensor marked as 7 is the internal tensor, which refers to the tensor calculated in the sub-kernel and is located in the specified memory corresponding to the sub-kernel. The specified memory is a memory structure closer to the core that is pre-divided for the sub-kernel, such as a shared cache (Shared Mem).

[0160] In this way, after building a routing reverse operator for the reverse propagation of the routing operator, for any module in the routing reverse operator, the computation tensors in the sub-kernel that executes the module can be stored in a memory structure closer to the core (such as a shared cache), thereby making full use of its high-speed memory access performance. For example, in subKernel-Main, when performing operations such as softmax reverse and topk reverse, storing the intermediate results in the shared cache can reduce frequent access to the global memory, thereby improving computing efficiency.

[0161] In the embodiment of the present application, after the reverse submodules are pre-derived and constructed, in the process of constructing the route reverse operator, the corresponding route reverse operator can be constructed by freely and flexibly combining the reverse submodules. Figure 7 As shown, in step 202, determining the reverse submodule included in the route reverse operator may include the following steps 702 to 704, wherein:

[0162] Step 702, obtaining configuration parameters of the route reverse operator, and determining each candidate reverse submodule whose configuration attribute in the configuration parameters is a first attribute value;

[0163] Step 704: determine the main path loss sub-template and each candidate reverse sub-module as the reverse sub-module of the route reverse operator.

[0164] In the embodiment of the present application, in actual application, the user can control whether to use the corresponding reverse submodule through the configuration parameters of the route reverse operator. For example, the configuration attributes of each route reverse operator can be configured in the configuration parameters of the route reverse operator. When the configuration attribute is the first attribute (for example, True), it can be determined that the reverse submodule is used and the reverse submodule is determined as a candidate reverse submodule. When the configuration attribute is the second attribute (for example, False), it can be determined that the reverse submodule is not used.

[0165] After determining each candidate reverse submodule, the main path loss submodule and each candidate reverse submodule may be connected in a corresponding order according to the dependency relationship, so that a routing reverse operator may be constructed.

[0166] Exemplarily, referring to Table 1 below, different operator combinations corresponding to different configuration attributes of the reverse submodule are shown when the reverse submodule includes a load balancing loss submodule and a z-loss loss submodule. Among them, when you want to ensure that the load between multiple expert modules is relatively balanced, you can use the load balancing loss module. For example, in large-scale natural language processing tasks, different experts may be good at processing texts of different semantic types. If some experts always process texts with common semantics, while other experts are rarely used, you can introduce a load balancing loss term to optimize this situation, so that each expert can play a full role and improve the overall generalization ability and stability of the model. When it is necessary to adjust the probability distribution of the model output to make it more in line with the actual situation, you can use the z-loss loss module. For example, in applications such as medical diagnosis assistance and financial risk assessment, if the probability of the model output cannot accurately reflect the real uncertainty, it may lead to serious consequences. Z-loss can help adjust the model output so that the predicted probability distribution is more reasonable and improve the credibility and accuracy of the model prediction.

[0167] Table 1

[0168] serial number combination Load Balancing Loss Module z-loss loss module 1 Main path loss submodule False False 2 Main path loss submodule + load balancing loss module True False 3 Main path loss submodule + z-loss loss module False True 4 Main path loss submodule + load balancing loss module + z-loss loss module True True

[0169] When deploying forward and reverse routing operators on AI chips, the problem of low computing resource utilization often arises. This is because the softmax and topk operators in the routing operators need to calculate data of specified dimensions (the specified dimensions are often the expert dimension (e) or the topk dimension (k)), and the hardware resources of the AI ​​chip process data of other non-specified dimensions (such as batch size batch_size or sequence length sequence_length) in parallel.

[0170] In the routing operator, the shape of the input data is usually [batch_size, sequence_length, e], and in these dimensions, batch_size and sequence_length are usually much larger than e or k. For example, batch_size may be 128, sequence_length may be 256, while e may be 16 and k may be 2.

[0171] The softmax and topk operations are performed in the e dimension, which means that each input data (indexed by batch_size and sequence_length) needs to independently calculate the weights of e experts. The topk operation is performed in the k dimension, which means that each input data needs to select k experts from e. In this way, if e=16, only 16 data points can be processed in parallel in the e dimension. If k=2, only 2 data points can be processed in parallel in the k dimension. For artificial intelligence chips, the hardware capacity (hardware_capacity) is usually at the level of tens to thousands. This low-dimensional parallelization cannot fully utilize the computing resources of the hardware.

[0172] For example, assuming that the shape of the input data is [batch_size=128, sequence_length=256, e=16], and softmax calculation is performed in the e dimension, each input data needs to calculate the weights of 16 experts, then the parallelization degree is only 16, while the hardware capacity of the AI ​​chip may be 1024, which means that the utilization rate of hardware resources is only 16 / 1024. For topk calculation in the k dimension, each input data needs to select 2 from 16 experts, the parallelization degree is only 2, and the utilization rate of hardware resources is only 2 / 1024.

[0173] In the dimensions of batch_size and sequence_length, the amount of data is large (128×256=32768), but these dimensions are not used for the calculation of softmax and topk, resulting in a waste of hardware resources.

[0174] Therefore, in order to further improve the utilization rate of hardware resources, in the embodiment of the present application, after the initial routing reverse operator is obtained by connection based on the determined reverse sub-module, the structure of each initial routing reverse operator can be optimized, the position of the transpose function in the initial routing reverse operator can be changed, and the e or topk dimension can be swapped with other dimensions to maximize the utilization rate of hardware resources.

[0175] The structural optimization process of the initial routing reverse operator is introduced below.

[0176] In an exemplary embodiment, referring to Figure 8 As shown, the method may further include the following steps 802 to 804, wherein:

[0177] Step 802, based on the computing resources of the artificial intelligence chip, determine the position adjustment strategy of the transposition operator in the routing reverse operator.

[0178] In an embodiment of the present application, the designated dimensions and non-designated dimensions of tensor data can be determined, and the hardware parallelism of the artificial intelligence chip can be determined based on the designated dimensions and non-designated dimensions and the computing resources of the artificial intelligence chip. The hardware parallelism reflects the degree to which the hardware can achieve parallel computing when the artificial intelligence chip processes related computing tasks. The higher the parallelism, the higher the computing efficiency may be in theory, thereby determining the position adjustment strategy for the transpose operator based on the hardware parallelism.

[0179] In an exemplary embodiment, referring to Fig. 9 As shown, in step 802, based on the computing resources of the artificial intelligence chip, determining the position adjustment strategy of the transposition operator in the routing reverse operator may include the following steps 902 to 906, wherein:

[0180] Step 902, determining a first hardware parallelism rate of the artificial intelligence chip based on the non-specified dimension before adjustment and the computing resources of the artificial intelligence chip;

[0181] Step 904: after swapping the designated dimension and the non-designated dimension, determine a second hardware parallelism rate of the artificial intelligence chip based on the swapped non-designated dimension and the computing resources of the artificial intelligence chip;

[0182] Step 906, when the first hardware parallelism rate is less than the second hardware parallelism rate, determine the target position of the transposition operator in the routing reverse operator, and generate a position adjustment strategy based on the target position of the transposition operator in the routing reverse operator.

[0183] In the embodiment of the present application, the first hardware parallelism rate can be calculated for the current non-specified dimension and the computing resources of the artificial intelligence chip, and the first hardware parallelism rate represents the hardware parallelism rate of the artificial intelligence chip before the structure is optimized. Afterwards, the non-specified dimension and the specified dimension can be swapped, that is, the specified dimension can be used as a new non-specified dimension, and based on the new non-specified dimension and the computing resources of the artificial intelligence chip, the second hardware parallelism rate can be calculated, and the second hardware parallelism rate represents the hardware parallelism rate of the artificial intelligence chip before and after the structure is optimized.

[0184] If the first hardware parallelism rate is less than the second hardware parallelism rate, it means that the hardware parallelism rate of the artificial intelligence chip is higher after the optimized structure is characterized. In this case, the structure of the routing reverse operator can be optimized, and the position of the transpose operator in the routing reverse operator can be adjusted to achieve the exchange of specified dimensions and non-specified dimensions, thereby improving the hardware parallelism rate of the artificial intelligence chip.

[0185] In an exemplary embodiment, referring to Fig.10 As shown, in step 906, determining the target position of the transposition operator in the route reverse operator may include the following steps 1002 to 1004, wherein:

[0186] Step 1002, determining a second target operator from the route reverse operator, where the second target operator is an operator whose operation is independent of dimension;

[0187] Step 1004: Determine the target position of the transposition operator in the route reverse operator based on the position of the second target operator.

[0188] In an embodiment of the present application, a second target operator can be determined from the route reverse operator. The second target operator is an operator whose operation is independent of dimension. The calculation characteristics of this type of operator do not depend on the specific dimension setting and are relatively stable when the dimension is adjusted. After determining the second target operator, a transposition operator can be set before the second target operator to re-swap the swapped specified dimension and non-specified dimension and perform subsequent calculation operations. That is, it can be determined that the target position of the transposition operator is before the second target operator, thereby generating a position adjustment strategy based on the target position.

[0189] Step 804: adjust the position of the transposition operator in the routing reverse operator based on the position adjustment strategy, and re-specify the specified dimension of the first target operator in the reverse submodule according to the position adjustment strategy.

[0190] In the embodiment of the present application, the position of the transposition operator in the route reverse operator is adjusted to the target position based on the position adjustment strategy, and the specified dimension of the first target operator in the reverse submodule can be re-specified to achieve the swapping of the original specified dimension and the non-specified dimension in the first target operator. Figures 6a to 6d The route reverse operator shown is Figure 6b As shown in the figure, before the specified dimension and the non-specified dimension are swapped, the transpose operator is located before scatter, then scatter can be determined as the first target operator, and the sequence parallel gather reverse function (also expressed as seq_gatherbwd) is the second target operator. At this time, the input shape of scatter is [capacity, e], the specified dimension is e, and the non-specified dimension is capacity. After the specified dimension and the non-specified dimension are swapped, the obtained routing reverse operator is referenced Figures 11a to 11c As shown, Fig.11a is the adjusted main path submodule, Fig.11b is the adjusted z-loss loss submodule, Fig.11c It is the adjusted load balancing loss submodule.

[0191] In the back propagation path, the transpose operator will be located before seq_gather bwd. At this time, the input shape of scatter is [e, capacity]. Fig.11a In the scatter operator, you can re-set the specified dimension by using dim=1 to re-specify capacity as the specified dimension. Fig.11b In the output of d-logtis-aux_loss, the shape is adjusted to [e,b×s], and in Fig.11c In the mean-bwd operator of the load balancing loss submodule in , reassign e to the specified dimension by dim=1.

[0192] The construction method of the routing reverse operator provided in the embodiment of the present application is aimed at the Router reverse operator of the mixed expert mode in the large model training, and an efficient implementation method based on an artificial intelligence chip is designed to help improve the computational efficiency of the Router reverse operator. According to the basic structure of the Router forward operator, the embodiment of the present application derives the calculation formula of its reverse operator through the forward calculation formula, and splits the reverse operator into multiple freely combinable sub-modules. After the multiple sub-modules are connected in the corresponding order, the function of the Router reverse operator can be efficiently realized when running on the artificial intelligence chip, and the frequent access to the global memory inside the Router operator is reduced as much as possible, and the calculation and memory access efficiency inside the Router operator is optimized. And the embodiment of the present application optimizes the structure of the existing Router reverse operator to further improve the parallel computing efficiency of hardware resources. Compared with the existing scheme, the scheme provided by the embodiment of the present application can greatly improve the computational efficiency of back propagation and improve the training effect of the model when running on the artificial intelligence chip.

[0193] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.

[0194] In addition, it should be noted that the method for constructing a routing reverse operator provided in the embodiment of the present application can be applied at least to the fields of speech processing, image processing, text processing, video processing, etc.

[0195] Exemplarily, in the field of speech processing, the MOE model can be trained to obtain a speech processing model, in the field of image processing, the MOE model can be trained to obtain an image processing model, in the field of text processing, the MOE model can be trained to obtain a text processing model, and in the field of video processing, the MOE model can be trained to obtain a video processing model. The construction method of the route reverse operator provided in the embodiment of the present application can be applied to the training process of the speech processing model, the text processing model, the image processing model, and the video processing model, while accelerating the training efficiency of the model and improving the model accuracy, which can effectively improve the text processing accuracy, image processing accuracy, video processing accuracy, and speech processing accuracy.

[0196] In order to enable those skilled in the art to better understand the embodiments of the present application, the following will take the application of the MOE model to a large-scale text classification task as an example to illustrate the embodiments of the present application. The application process in the fields of speech processing, image processing, and video processing can refer to this example and will not be repeated in the embodiments of the present application.

[0197] In this example, in the data preprocessing stage, data collection, cleaning, vectorization and other processing can be performed, among which:

[0198] Collect data: Collect a large amount of text data covering various categories, such as news article classification, which may include different categories of news such as economy and culture. Text cleaning: Remove noise from the text, such as special characters, stop words, etc., to improve data quality. Text vectorization: Convert the cleaned text into a vector form that can be processed by a computer. Common methods include bag-of-words model, word frequency-inverse document frequency, or use a pre-trained word vector model to map each word to a low-dimensional vector. For long texts, deep learning models can also be used to obtain contextual representations of the text.

[0199] The MOE model process includes an input layer, a gating network (a routing operator in the embodiment of the present application, for the sake of clarity, the routing operator will be used in place of the gating network in the following description), an expert network and an output layer, wherein:

[0200] Input layer: The preprocessed text vector is used as the input of the MOE model. The shape may be [batch_size, sequence_length, embedding_dim], where batch_size is the number of samples processed at a time, sequence_length is the length of the text sequence, and embedding_dim is the vector dimension.

[0201] Routing operator: responsible for assigning input text to different experts. It receives the input vector and calculates the "fitness" (probability) of each expert to the input through linear transformation and activation function (such as Softmax). For example, for a certain input text, the gating network may calculate that the fitness probability of expert 1 is 0.2, the fitness probability of expert 2 is 0.3, etc. At the same time, in order to balance the expert load and avoid overuse of some experts, load balancing loss terms and z-loss loss terms are added to the routing operator.

[0202] Expert Network: Multiple expert networks work in parallel, with each expert focusing on the classification of a specific type of text. Each expert receives part of the input data assigned by the gating network (weighted by the probability of adaptation), extracts features and classifies them. For example, expert 1 may be good at processing economic news, and expert 2 may be good at cultural news. The expert network can be a simple multilayer perceptron (MLP) or a complex convolutional neural network or recurrent neural network.

[0203] Output layer: The outputs (classification results) of each expert network are weighted and summed according to the probability of the gated network to obtain the final classification result. For example, if the probability of expert 1 outputting the classification as "economy" is 0.8, and the probability of expert 2 outputting the classification as "culture" is 0.6, the probability assigned by the gated network to expert 1 is 0.3, and the probability assigned to expert 2 is 0.7, then the final classification result is the weighted sum of the two expert results, and after the Softmax function, the final probability distribution of each category is obtained.

[0204] During the training process, the cross entropy loss function can be used to measure the difference between the model prediction results and the true labels. Stochastic gradient descent is used to update the parameters of the routing operator and the expert network to minimize the loss function. During the training process, the gradient is calculated through the back propagation algorithm and the model parameters are adjusted to make the model more accurate in classifying various types of text.

[0205] In the reverse propagation process of the routing operator, in the embodiment of the present application, the corresponding reverse submodule is called from the operator library based on the functional module adopted by the routing operator to construct the routing reverse operator. The reverse submodule is an operator module for gradient calculation that is constructed based on the algorithm formula derived after reverse deduction based on the algorithm formula of each functional module in advance. The specific derivation process and the process of constructing the routing reverse operator based on the reverse submodule can refer to the relevant description of the aforementioned embodiment, which will not be repeated here in the embodiment of the present disclosure. The reverse submodules of the routing reverse operator are mapped to multiple sub-kernels of the GPGPU for gradient calculation, so as to update the parameters of the routing operator by gradient descent to minimize the loss function.

[0206] It should be noted that artificial intelligence chips have been widely used in deep learning model training. Artificial intelligence chips can include GPU (Graphics Processing Unit), GPGPU (General-Purpose Computing on Graphics Processing Units), TPU (Tensor Processing Unit), NPU (Neural Network Processing Unit) and other chips. Taking GPGPU as an example, Fig.12 A schematic structural diagram of a general-purpose graphics processing unit (GPGPU) is shown.

[0207] like Fig.12 As shown, a general-purpose graphics processor is actually an array of streaming processor clusters (SPC), including Fig.12 The stream processor clusters 1, ..., and stream processor cluster M are shown, where M is a positive integer greater than 1. In a graphics processor, one stream processor cluster processes one computing task, or multiple stream processor clusters process one computing task. Multiple stream processor clusters share data through a global cache or a global memory.

[0208] like Fig.12 As shown, taking stream processor cluster 1 as an example, one stream processor cluster includes multiple computing units (also called), such as Fig.12In the CU, there are CU 1, CU 2, ..., CU N, where N is a positive integer. Each CU is used to perform arithmetic and logic operations other than matrix calculations such as matrix multiplication and convolution operations, such as accumulation, reduction, conventional addition, subtraction, multiplication, and division. A CU includes multiple cores (also called computing cores or computing cores), each of which includes an arithmetic logic unit (ALU), a floating-point computing unit, etc. The computing core is used to perform specific computing tasks. In addition, the CU also includes registers (e.g. Fig.12 The register file in the computing unit and the shared cache are used to hierarchically store source data and destination data related to computing tasks. The shared cache in a computing unit is used to share data between the cores of the computing unit.

[0209] In parallel computing, computing tasks are generally performed by multiple threads. These threads are divided into multiple thread blocks before being executed in a general-purpose graphics processor (or parallel computing processor), and then distributed through the thread block distribution module ( Fig.12 The CPU (not shown) distributes multiple thread blocks to each computing unit. All threads in a thread block must be assigned to the same computing unit for execution. At the same time, the thread block is split into minimum execution thread bundles (or simply thread bundles, warps), each of which contains a fixed number of threads (or less than this fixed number), for example, 32 threads. Multiple thread blocks can be executed in the same computing unit or in different computing units.

[0210] In each compute unit, the warp scheduling / dispatching module ( Fig.12 The thread warps are scheduled and allocated (not shown) so that multiple computing cores of the computing unit can run the thread warps. Depending on the number of computing cores in the computing unit, multiple thread warps in a thread block can be executed simultaneously or in time-sharing. Multiple threads in each thread warp will execute the same instruction. The memory execution instruction will be emitted to the shared cache in the computing unit or further emitted to the intermediate cache or global cache or global memory for read and write operations, etc.

[0211] In the embodiment of the present application, the parallel computing capability of GPGPU can be used to complete the training of the MOE model on GPGPU. Including: storing the input data, intermediate results and output data related to the Router operator in the forward propagation process in the global memory of GPGPU. These data need to be correctly loaded into the appropriate memory location before the back propagation begins for subsequent calculations. For example, the input feature tensor, the gating signal generated by the Router operator, and the output tensor after the Router processing, etc.

[0212] In order to better utilize the parallelism of GPGPU, data is usually divided into multiple small blocks. Each small block can be processed by one or more thread blocks. In this way, gradient calculations can be performed simultaneously on different thread blocks, improving computational efficiency.

[0213] The reverse submodules included in the routing reverse operator can be determined based on the configuration of the Router operator, and each reverse submodule can be called from the operator library, and the routing reverse operator corresponding to the Router operator can be constructed according to the dependency connection, and each reverse submodule in the routing reverse operator can be mapped to the corresponding subkernel, and the subkernel can be executed in each computing unit of the GPGPU to implement the reverse propagation calculation, so as to use each computing unit on the GPGPU to calculate the gradient of a part of the data. Exemplarily, for operations such as matrix multiplication and activation functions involved in the Router operator, the corresponding reverse calculation kernel functions are implemented respectively. These kernel functions will calculate the gradient of the input of the current layer (Router operator layer) relative to the loss function based on the input gradient information (the gradient transmitted in reverse from the subsequent layer).

[0214] Since computations on GPGPU are performed in parallel, different thread blocks may complete computations at different speeds. In order to ensure that all thread blocks have completed their current computations before performing the next computation (such as gradient aggregation or data transfer), synchronization operations are required.

[0215] After calculating the local gradients, each computing unit needs to summarize these local gradients to get the final gradient value. The summarized gradient value will be used as the back propagation result of the router operator and passed to the previous layer for further calculation.

[0216] Through the above steps, the back propagation of the Router operator is implemented on GPGPU, and the gradient is efficiently calculated by using the parallel computing capability of GPGPU, thereby improving the training efficiency of the MOE model.

[0217] Based on the same inventive concept, the embodiment of the present application also provides a device for constructing a routing reverse operator for implementing the method for constructing a routing reverse operator involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in the embodiments of the device for constructing one or more routing reverse operators provided below can refer to the limitations of the method for constructing a routing reverse operator above, and will not be repeated here.

[0218] In an exemplary embodiment, Fig.13As shown, a device 1300 for constructing a route reverse operator is provided, comprising: a first determining module 1302 and a first constructing module 1304, wherein:

[0219] A first determining module 1302 is used to determine a reverse submodule included in a route reverse operator, wherein the reverse submodule is used to perform gradient calculation;

[0220] The first construction module 1304 is used to call each of the reverse submodules from the operator library, connect each of the reverse submodules according to the dependency relationship, and then map each of the reverse submodules to a corresponding sub-kernel in the artificial intelligence chip to construct a routing reverse operator;

[0221] The reverse submodule is constructed based on the derived algorithm formula after the algorithm formula corresponding to the forward submodule is derived, and the forward submodule is a submodule obtained by dividing the routing forward operator.

[0222] The construction device of the above-mentioned routing reverse operator is obtained by constructing the routing reverse operator based on the reverse submodule in the operator library, and the reverse submodule is constructed based on the algorithm formula derived from the algorithm formula corresponding to the forward submodule. This specially derived reverse operator formula enables the reverse submodule to reduce certain transitional intermediate variables in the forward calculation through derivation. In this way, in the calculation of reverse propagation, there is no need to re-read and store this part of the data, which reduces the access to this data in the global memory, that is, when performing reverse propagation, there is no need to perform gradient calculation on the entire operator graph of the routing operator, which reduces unnecessary storage and reading of intermediate results, can greatly improve the utilization rate of hardware computing resources, improve the calculation efficiency of reverse propagation, and thus improve the training efficiency of the model.

[0223] In one embodiment, the device further comprises:

[0224] A partitioning module is used to partition the routing forward operator into modules to obtain multiple forward sub-modules;

[0225] For any of the forward submodules, perform the following operations:

[0226] A second determination module, used to determine the algorithm formula of the reverse submodule corresponding to the forward submodule according to the algorithm formula of the forward submodule;

[0227] The second construction module is used to determine the operators included in the reverse submodule and the dependencies between the operators based on the algorithm formula of the reverse submodule, and construct the reverse submodule according to the operators and the dependencies between the operators.

[0228] In one embodiment, the input tensors, intermediate tensors and output tensors of the routing reverse operator are all stored in the global memory, and the internal tensors of the reverse submodule are stored in the designated memory, which is the memory corresponding to the sub-kernel that executes the reverse submodule.

[0229] In one embodiment, the first determination module is specifically configured to:

[0230] Acquire configuration parameters of the route reverse operator, and determine candidate reverse submodules whose configuration attributes in the configuration parameters are first attribute values;

[0231] The main path loss sub-template and each of the candidate reverse sub-modules are determined as the reverse sub-modules of the routing reverse operator.

[0232] In one embodiment, the device further comprises:

[0233] A fourth determination module, used to determine a position adjustment strategy of the transposition operator in the routing reverse operator based on the computing resources of the artificial intelligence chip;

[0234] The optimization module is used to adjust the position of the transposition operator in the routing reverse operator based on the position adjustment strategy, and re-specify the specified dimension of the first target operator in the reverse submodule according to the position adjustment strategy.

[0235] In one embodiment, the optimization module is specifically used for:

[0236] Determining a first hardware parallelism rate of the artificial intelligence chip based on the non-specified dimension before adjustment and the computing resources of the artificial intelligence chip;

[0237] After swapping the designated dimension and the non-designated dimension, determining a second hardware parallelism rate of the artificial intelligence chip based on the swapped non-designated dimension and the computing resources of the artificial intelligence chip;

[0238] When the first hardware parallelism rate is less than the second hardware parallelism rate, a target position of a transposition operator in the route reversal operator is determined, and a position adjustment strategy is generated based on the target position of the transposition operator in the route reversal operator.

[0239] In one embodiment, the optimization module is specifically used for:

[0240] Determining a second target operator from the route reverse operator, wherein the second target operator is an operator whose operation is dimension-independent;

[0241] Based on the position of the second target operator, a target position of the transposition operator in the route reversal operator is determined.

[0242] In one embodiment, the partitioning module is specifically used for:

[0243] Get the routing forward operator graph corresponding to the routing forward operator;

[0244] Dividing the routing forward operator graph into sub-modules to obtain multiple operator graphs;

[0245] According to the multiple operator graphs, multiple forward submodules are obtained.

[0246] Each module in the construction device of the above-mentioned route reverse operator can be implemented in whole or in part by software, hardware and their combination. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0247] In an exemplary embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as shown in FIG. Fig.14 As shown. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit and an input device. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface, the display unit and the input device are connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and the external device. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, near field communication (NFC) or other technologies. When the computer program is executed by the processor, a method for constructing a routing reverse operator is implemented. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device shell, or an external keyboard, touchpad or mouse.

[0248] Those skilled in the art will understand that Fig.14The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0249] In one embodiment, a computer device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above method embodiments when executing the computer program.

[0250] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0251] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.

[0252] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0253] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment method can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (AI) processor, etc., but are not limited to this.

[0254] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0255] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.

Claims

1. A method for constructing a routing reverse operator, characterized in that: The method comprises: Determine a reverse submodule included in a route reverse operator, wherein the reverse submodule is used to perform gradient calculation; Calling each of the reverse submodules from the operator library, connecting each of the reverse submodules according to the dependency relationship, and mapping each of the reverse submodules to the corresponding sub-core in the artificial intelligence chip to construct a routing reverse operator; The reverse submodule is constructed based on the derived algorithm formula after the algorithm formula corresponding to the forward submodule is derived, and the forward submodule is a submodule obtained by dividing the routing forward operator.

2. The method according to claim 1, characterized in that The method further comprises: Divide the routing forward operator into modules to obtain multiple forward sub-modules; For any of the forward submodules, perform the following operations: According to the algorithm formula of the forward submodule, determine the algorithm formula of the reverse submodule corresponding to the forward submodule; Based on the algorithm formula of the reverse submodule, the operators included in the reverse submodule and the dependencies between the operators are determined, and according to the operators and the dependencies between the operators, the reverse submodule is constructed.

3. The method according to claim 1 or 2, characterized in that: The input tensors, intermediate tensors and output tensors of the routing reverse operator are all stored in the global memory, and the internal tensors of the reverse submodule are stored in the specified memory, which is the memory corresponding to the sub-kernel that executes the reverse submodule.

4. The method according to claim 1 or 2, characterized in that: The reverse submodule included in the determining route reverse operator includes: Acquire configuration parameters of the route reverse operator, and determine candidate reverse submodules whose configuration attributes in the configuration parameters are first attribute values; The main path loss sub-template and each of the candidate reverse sub-modules are determined as the reverse sub-modules of the routing reverse operator.

5. The method according to claim 1 or 2, characterized in that: Before mapping each of the reverse submodules to a corresponding sub-core in the artificial intelligence chip to construct a route reverse operator, the method further includes: Based on the computing resources of the artificial intelligence chip, determining a position adjustment strategy of the transposition operator in the routing reverse operator; The position of the transposition operator in the routing reverse operator is adjusted based on the position adjustment strategy, and the designated dimension of the first target operator in the reverse submodule is re-designated according to the position adjustment strategy.

6. The method according to claim 5, characterized in that The determining of the position adjustment strategy of the transposition operator in the routing reverse operator based on the computing resources of the artificial intelligence chip includes: Determining a first hardware parallelism rate of the artificial intelligence chip based on the non-specified dimension before adjustment and the computing resources of the artificial intelligence chip; After swapping the designated dimension and the non-designated dimension, determining a second hardware parallelism rate of the artificial intelligence chip based on the swapped non-designated dimension and the computing resources of the artificial intelligence chip; When the first hardware parallelism rate is less than the second hardware parallelism rate, a target position of a transposition operator in the route reversal operator is determined, and a position adjustment strategy is generated based on the target position of the transposition operator in the route reversal operator.

7. The method according to claim 6, characterized in that The determining of the target position of the transposition operator in the route reversal operator comprises: Determining a second target operator from the route reverse operator, wherein the second target operator is an operator whose operation is dimension-independent; Based on the position of the second target operator, a target position of the transposition operator in the route reversal operator is determined.

8. The method according to claim 4, characterized in that The routing forward operator is divided into modules to obtain multiple forward submodules, including: Get the routing forward operator graph corresponding to the routing forward operator; Dividing the routing forward operator graph into sub-modules to obtain multiple operator graphs; According to the multiple operator graphs, multiple forward submodules are obtained.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.

11. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Static generation method and device of reverse calculation graph, equipment and medium

    CN117273115A

  • Model exception handling method and device, electronic equipment and storage medium

    CN117827523A

  • Allocating and accessing website resources via domain name routing rules

    US20150304235A1