Neural network tensor shape tracking method and computing platform
By constructing shape flow graphs and constraint condition judgments, the problem of resource waste caused by tensor shape mismatch in neural network calculations is solved, shape errors are discovered in advance, and calculation efficiency is improved.
Patent Information
- Application Number
- CN202111353668.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-16
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2041-11-16
AI Technical Summary
In the prior art, calculations during neural network calculations are terminated due to tensor shape mismatch, resulting in a waste of transmission, storage, and computing resources, and shape errors cannot be discovered in pre-checks.
Build a shape flow graph based on the computational graph, extract the constraints of tensor shape information, determine shape errors or limitations, and implement pre-inspection to discover and report shape problems.
It avoids the waste of resources caused by shape errors and improves the efficiency of neural network calculation and resource utilization.
Smart Images

Figure CN114492772B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of deep learning, and in particular to a neural network tensor shape tracking method and computing platform. Background Art
[0002] Artificial intelligence has developed rapidly in recent years, achieving promising results in areas such as image classification, detection, video, and speech processing, and continues to hold great promise for future development. Neural networks are at the core of AI applications, with deep learning neural network algorithms being the most common neural network model. Neural network workloads are characterized by being computationally and data intensive. The multiplication and addition operations required for neural network computations are typically on the order of gigabytes. For example, the computational complexity of an SSD neural network for object detection can reach 120 gigabytes. The parameters required for training a neural network and subsequent inference are typically on the order of megabytes to hundreds of megabytes. For example, the parameter count for the VGG classification neural network is as high as 480 megabytes. Because standard computers generally lack the computing power and memory to achieve this level of performance, neural network training and inference are increasingly performed on specialized neural network platforms.
[0003] To perform calculations on a neural network platform, users need to transfer data, such as training data, that can generate computational graphs to the platform. In actual neural network calculations, eigenvalues are passed layer by layer between multiple layers of the neural network. If the shape of the eigenvalue being transferred does not match the shape of the eigenvalue to be received in the next layer, the calculation will terminate. In existing technologies, such shape mismatches cannot be detected during pre-check operations, resulting in significant waste of transmission, storage, and computing resources.
[0004] To this end, the present invention requires an improved neural network calculation solution. Summary of the Invention
[0005] A technical problem to be solved by the present disclosure is to provide an improved neural network tensor shape tracking scheme, which constructs a shape flow graph based on the existing computational graph, obtains the constraints of the eigenvalue shape flow and discovers shape errors or limitations accordingly, so that shape errors that could only be discovered in actual computing operations due to termination of operations can be found through pre-check, thereby avoiding unnecessary transmission, storage and computing resource consumption.
[0006] According to a first aspect of the present disclosure, a method for tracking the shape of a neural network tensor is provided, comprising: constructing a shape flow graph based on an intermediate representation (IR) of a computational graph of a neural network, the shape flow graph comprising nodes for performing neural network computational operations and edges representing dependencies between the nodes, the edges being marked with shape information of tensors to flow along the edges; extracting constraints on the shape information of the tensor from the node operations and newly introduced operands of the shape flow graph; determining whether the shape information of the tensor has errors or restrictions based on the constraints; and reporting when it is determined that the shape information of the tensor has errors or restrictions.
[0007] Optionally, there is unknown information in the shape information of the added corresponding tensor, and the constraints for the tensor shape information are extracted from the node operations and newly introduced operands of the shape flow graph, including: listing multiple constraints for the unknown information based on the relationship between each node operation and operand of the shape flow graph and the unknown information; and solving the unknown information based on the multiple constraints: if the unknown information has no solution, it is determined that there is a shape error; if the unknown information has a specific solution, it is determined that there is a shape restriction.
[0008] Optionally, constructing the shape flow graph includes: slicing backward from the input tensor of the intermediate representation of the computational graph to construct the shape flow graph, and collecting the operands required for node operations during the backward slicing process, wherein the operands include constants and scalar variables.
[0009] Optionally, constructing the shape flow graph includes: adding a return value of a call node operation as a new node to the shape flow graph in backward slicing, and continuing backward slicing after the newly added node.
[0010] Optionally, constructing the shape flow graph includes: copying all previous slices of a confluence node with n incoming values into n slices; and each copied previous slice selects a different incoming value from the n incoming values to continue slicing, and constructing a shape flow graph for each path in the computation graph.
[0011] Optionally, the reporting when determining that there are errors or limitations in the shape information of the tensor includes: deleting node transformation operations and constraints introduced by the node transformation operations one by one; and locating the node operations that cause the errors or limitations when the remaining constraints are met.
[0012] The tensor shape tracking method makes the determination without actually having eigenvalues flowing.
[0013] According to a second aspect of the present disclosure, a neural network computing platform is provided, comprising: a tensor shape tracking module, comprising: a construction submodule for constructing a shape flow graph based on an intermediate representation (IR) of a computational graph of a neural network, wherein the shape flow graph includes nodes for performing neural network computing operations and edges representing dependencies between nodes, wherein the edges are marked with shape information of tensors to flow along the edges; a constraint determination submodule for extracting constraints on tensor shape information from node operations and newly introduced operands of the shape flow graph, and determining whether there are errors or restrictions in the shape information of the tensor based on the constraints; and a reporting submodule for reporting when it is determined that there are errors or restrictions in the shape information of the tensor; and a neural network training module for performing neural network training calculations based on the intermediate representation of the computational graph in which the errors or restrictions have been eliminated according to the operations of the tensor shape tracking module.
[0014] Optionally, the computing platform further includes: a training data acquisition module for acquiring training data and training labels, wherein the training data serves as the initial tensor of the flow tensor in the computing graph, and the training labels are used to reversely adjust the classification results output by the computing graph, wherein, in response to the tensor shape tracking module determining that there are no errors or limitations, the neural network training module loads the training data and the training labels to perform neural network training calculations.
[0015] According to a third aspect of the present disclosure, a non-transitory machine-readable storage medium is provided, on which executable code is stored. When the executable code is executed by a processor of an electronic device, the processor is caused to execute the method as described in the first aspect.
[0016] This paper proposes a shape tracking method specifically suited for tensor shapes with unknown shapes in computational graphs. This method symbolically represents tensor shapes by introducing specific symbols for tensor shapes of unknown rank or dimension size. Constraints can be introduced from tensor operators, scalar variables, and conditional branches. Finally, the satisfiability of these constraints is checked and reported. This allows shape problems to be identified through pre-checks before actual neural network calculations are performed, avoiding unnecessary computations that waste transmission, access, and computing resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The above and other objects, features and advantages of the present disclosure will become more apparent through a more detailed description of exemplary embodiments of the present disclosure with reference to the accompanying drawings, wherein like reference numerals generally represent like components in the exemplary embodiments of the present disclosure.
[0018] Figure 1 An example of the composition of a typical CNN is shown.
[0019] Figure 2 An example of a computational graph for training a CNN network built on a deep computing framework is shown.
[0020] Figure 3 An example of an input feature map as a tensor is shown.
[0021] Figure 4 A schematic diagram of a neural network computing system to which the present invention is applied is shown.
[0022] Figure 5 A schematic flowchart of a tensor shape tracking method for a neural network according to an embodiment of the present invention is shown.
[0023] Figure 6 An implementation example of a neural network tensor shape tracker according to an embodiment of the present invention is shown.
[0024] Figure 7 An example of a shape flow graph constructed according to the present invention is shown.
[0025] Figure 8 Another example of a shape flow graph constructed according to the present invention is shown.
[0026] Figure 9 A schematic diagram of the composition of a neural network computing platform according to an embodiment of the present invention is shown.
[0027] Figure 10 A schematic structural diagram of a computing device that can be used to implement the above-mentioned shape tracking method according to an embodiment of the present invention is shown. DETAILED DESCRIPTION
[0028] The preferred embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although preferred embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments described herein. Rather, these embodiments are provided to make the present disclosure more thorough and complete, and to fully convey the scope of the present disclosure to those skilled in the art.
[0029] Artificial intelligence has developed rapidly in recent years, achieving promising results in areas such as image classification, detection, video, and speech processing, and continues to hold great promise for future development. Neural networks are at the core of AI applications, with deep learning neural network algorithms being the most common neural network model. Neural network workloads are characterized by being computationally and data intensive. The multiplication and addition operations required for neural network computations are typically on the order of gigabytes. For example, the computational complexity of an SSD neural network for object detection can reach 120 gigabytes. The parameters required for training a neural network and subsequent inference are typically on the order of megabytes to hundreds of megabytes. For example, the parameters of a VGG neural network for classification use 480 megabytes.
[0030] Common artificial neural networks (ANNs) include deep neural networks (DNNs), recurrent neural networks (RNNs), and convolutional neural networks (CNNs). CNNs are a type of artificial neural network that has become a research hotspot in the fields of speech analysis and image recognition. Their weight-sharing network structure makes them more similar to biological neural networks, reducing the complexity of the network model and the number of weights. This advantage is particularly evident when the network input is a multidimensional image, allowing the image to be directly used as the network input, avoiding the complex feature extraction and data reconstruction processes of traditional recognition algorithms. A convolutional network is a multilayer perceptron specifically designed for recognizing two-dimensional shapes. This network structure is highly invariant to translation, scaling, tilting, or other forms of deformation. The following will provide a background explanation of neural network computing and related concepts, particularly in conjunction with the accompanying figures.
[0031] Basic Concepts of CNN
[0032] Figure 1 Figure 2 shows an example of the composition of a typical CNN. Figure 1 As shown, a typical CNN consists of a series of layers that run in order.
[0033] The parameters of a CNN model are called "weights." The first layer of a CNN reads an input image and outputs a series of feature maps. The following layers read the feature maps generated by the previous layer and output new feature maps. The final classifier outputs the probability that the input image belongs to a certain class. CONV layers (convolutional layers) and FC layers (fully connected layers) are the two basic layer types in a CNN. After the CONV layer, there is usually a pooling layer.
[0034] In this application, for a CNN layer, represents the j-th input feature map, represents the i-th output feature map, b i represents the bias term of the i-th output graph.
[0035] For CONV layers, n in and n out Represent the number of input and output feature maps respectively.
[0036] For the FC layer, n in and n out Represent the length of the input and output feature vectors respectively.
[0037] Definition of CONV layer (Convolutional layers): The CONV layer takes a series of feature maps as input and convolves them with a convolution kernel to obtain an output feature map.
[0038] A nonlinear layer, i.e., a nonlinear activation function, is usually connected to the CONV layer and applied to each element in the output feature map. The activation function used is generally the ReLU function, and this layer is also commonly called a ReLU layer.
[0039] The CONV layer can be expressed by Expression 1:
[0040] (1)
[0041] where g i,j It is the convolution kernel applied to the j-th input feature map and the i-th output feature map. Definition of FC layer (Fully-Connected layers): FC layer applies a linear transformation upward to the input features:
[0042] (2)f out =Wf in +b
[0043] W is a n out ×n in The transformation matrix, b is the bias term. It is worth noting that for the FC layer, the input is not a combination of several two-dimensional feature maps, but a feature vector. Therefore, in Expression 2, the parameter n in and n out Actually corresponds to the length of the input and output feature vectors.
[0044] Pooling layer: Usually connected to the CONV layer, it is used to output the maximum or average value of each subarea in each feature map. The maximum value of Pooling can be expressed by Expression 3:
[0045] (3) where p is the size of the pooling kernel. This nonlinear "downsampling" not only reduces the size and computation of the feature maps for the next layer but also provides translation invariance. CNNs can be used for image classification during forward inference.
[0046] Before deploying a CNN for inference (e.g., image classification), it must first be trained. By importing a large amount of training data, the parameters of each layer of the neural network model, such as weights and biases, are determined.
[0047] CNN training
[0048] Training a model means learning (determining) the ideal values of all weights and biases through labeled samples. These determined weights and biases can then be used to perform high-accuracy inference on the input feature values during the neural network deployment phase, for example, correctly classifying the input image.
[0049] In supervised learning, a machine learning algorithm learns parameters by examining many examples and trying to find a model that minimizes loss, a process called empirical risk minimization.
[0050] Loss is a penalty for poor predictions. In other words, loss can be a numerical value that represents how accurate the model's predictions are for a single example. If the model's predictions are perfectly accurate, the loss is zero; otherwise, the loss is higher. The goal of training a model is to find a set of weights and biases that, across all examples, results in a "low" average loss.
[0051] During the training of a neural network, in order to quantify whether the current weights and biases can make the network input fit all network inputs, a loss function needs to be defined (as shown below): Figure 2 The purpose of training a network can thus be transformed into a process of minimizing the loss function of weights and biases. Typically, the gradient descent algorithm (or backpropagation algorithm in multi-layer neural network training) is used to implement the above minimization process.
[0052] The backpropagation algorithm involves repeated iterations of forward and backward propagation. Forward propagation involves connecting neurons between layers via a weight matrix, allowing stimuli (eigenvalues) to be continuously transferred from the previous layer to the next via the activation function of each layer. In reverse propagation, the error of a layer is inferred from the error of the next layer. This iterative process of forward and backward propagation continuously adjusts weights and biases, gradually minimizing the loss function and thus completing neural network training.
[0053] Deep Learning Frameworks and Computational Graphs
[0054] Deep learning frameworks provide building blocks for the design, training, and validation of neural networks through high-level programming interfaces. In other words, deep learning frameworks provide the building blocks for specific neural network algorithms (e.g., Figure 1 The neural network structure shown in FIG.
[0055] With the development of deep learning and neural network algorithms, many top-level deep learning frameworks have emerged for researchers and developers, such as TensorFlow and PyTorch. Developers can use the DSLs and APIs of these frameworks to design different computational graph models and implement specific tasks, such as face recognition, image detection, and speech recognition.
[0056] The programming methods of these computing frameworks differ significantly. Whether compiled or scripting, they compute variables step by step to produce results. However, TensorFlow and Pytorch differ. First, a computational graph is constructed through programming. Data is then used as input, and the computations are performed using the operations specified in the graph, ultimately yielding the results. This approach overcomes the limitations of programming languages, facilitates decoupling of the frontend and backend, and enables more intuitive visualization.
[0057] The computational graph model consists of nodes and edges. Nodes represent operators, or operators, and edges represent dependencies between computations. Solid lines indicate data transfer dependencies, where the transferred data is tensors.
[0058] Figure 2 An example of a computational graph for training a CNN network built on a deep computing framework is shown. Figure 2 The computational graph shown can be obtained from the CNN network training code generated by programming with TensorFlow. The computational graph shown incorporates the Pad or BiasAdd (bias addition) operations adjacent to a Conv2d (two-dimensional convolution) into the edges represented by that Conv2d, and all constant nodes into the attributes of the edges corresponding to the corresponding operators, thus constructing a directed acyclic graph.
[0059] When training a neural network, a batch of training samples needs to be provided each time. If the data selected in each iteration is represented by a constant, the computational graph of TensorFlow will become very large. Because each time a constant is added, TensorFlow will add a node to the computational graph. However, placeholders can solve this problem. It only has one node, the placeholder. The value of each placeholder used can be given by feed_dict (dictionary).
[0060] To this end, in the data input stage on the far left of the computational graph, the dictionary implements the input of the initial training image by feeding the training image (train_img) into the placeholder. In the prediction (predict) box, the two-dimensional convolution (conv2d) operation uses the convolution kernel obtained by get_varible (get variable) to convolve the input data, and through the reshaped (reconstruction) operation, the convolution is made to meet the shape requirements of various API calls in the deep learning framework. In multiple convolutional layers (the convolution operation in the prediction box can be repeated, that is, there can be multiple such as Figure 1After the convolution + pooling operation shown in the figure, the eigenvalues are fed into the first fully connected layer, and the vector obtained by get_varible is used to implement the fully connected calculation. The obtained eigenvalues are then fed into the second fully connected layer, and the vector obtained by get_varible is used to implement the fully connected calculation. The calculation results obtained can be used by the dictionary to perform cross-entropy backpropagation (softmax_cross_entropy_with_logits) based on softmax classification by feeding the training label (train_lab) into the placeholder. After multiple trainings with batches of training images, the various parameters (for example, the value of the convolution kernel) will converge to a relatively fixed distribution, so that the above parameters can be used as the trained neural network parameters to predict new images.
[0061] Figure 2 The figure shows a dynamic computation graph. That is, each time an operator is used, the operator is dynamically added to the implicit default computation graph and immediately executed to obtain the result, which facilitates debugging and use.
[0062] Tensors and shapes
[0063] In CNN calculations (both training and prediction) based on deep frameworks, the eigenvalues are calculated along Figure 2 The node operations in the function flow one by one, following the path shown in the diagram: conv2d → reshaped → matmul → matmul → softmax_cross_entropy_with_logits . Here, the eigenvalues are typically multidimensional matrices and are referred to as tensors. This is also the origin of the name of the deep learning framework TensorFlow.
[0064] Tensors have a shape. Shape refers to the length (number of elements) of each axis of the tensor. Rank refers to the number of axes in the tensor. A scalar has rank 0, a vector has rank 1, and a matrix has rank 2. Axis or dimension can be used to refer to a specific dimension of a tensor. Size or dimension can refer to the total number of entries in a tensor, i.e., the product shape vector.
[0065] Axes are usually referred to by their indices. Axes are usually ordered from global to local: first the batch, then the spatial dimensions, and finally the features at each position. This allows the feature vectors to be located in a contiguous region in memory.
[0066] Figure 3An example of an input feature map as a tensor is shown. The tensor shown is a rank 4 tensor with a shape of [2, 4, 5, 3] and a size of 60, i.e., it contains 60 elements. Specifically, the tensor can include four dimensions: batch, width, height, and feature. For example, when a 4x5 pixel RGB image is used as a training image and two images are trained as a batch each time, the following can be obtained: Figure 3 The input tensor shown in Figure 1 is a 3D image. At this point, each cube in the figure can represent the R, G, or B value of a pixel in a training image.
[0067] It should be understood that Figure 3 This example uses a smaller dataset for ease of illustration. In real-world training, each training image can have a larger pixel size, such as 40x60, and each batch can train more images, such as 512. Furthermore, the training images do not need to be RGB. Furthermore, although the figure shows a four-dimensional tensor in three-dimensional space for ease of understanding, this representation is not generally used to describe space.
[0068] Tensors flow in a single direction in a computational graph implemented as a directed acyclic graph, for example from Figure 2 The left side of the network flows to the right side and its shape changes due to the operations of the node operators. For example, it may be convolved with different convolution kernels, padded with different padding strategies, or restructured to meet API call requirements.
[0069] A ShapeError occurs when a TensorFlow operator is called with incompatible shape parameters (incompatible rank or dimension). A ShapeError manifests itself in the computational graph as a mismatch between the shape of the tensor flowing into a node and the shape expected by that node after the operation. This error frequently occurs in practice because developers struggle to understand the complex semantics of the thousands of APIs in deep learning frameworks.
[0070] For example, many TensorFlow operators (e.g., Figure 2 softmax_cross_entropy_with_logits in __init__.py ) supports NumPy "broadcasting" semantics (which "broadcasts" a small array within a relatively large array by copying the leading dimension of the higher-rank argument and padding any dimension of size 1 from that other argument to match the size of the argument), which often leads to incorrect results.
[0071] In the prior art, it is impossible to detect shape errors, especially parameters with partially or completely unknown shapes, through pre-checking (e.g., debugging) before actual neural network training.
[0072] As mentioned above, since ordinary computers are usually unable to provide the high computing power required by CNN calculations and are not equipped with sufficient memory, more and more neural network training and derivation are performed on dedicated neural network platforms. In order to perform calculations on the neural network platform, users need to transfer data that can generate calculation graphs, training data, etc. to the platform. In actual neural network calculations, eigenvalues (in the form of tensors) are passed layer by layer between multiple layers of the neural network. If the shape of the transferred tensor is incorrect, the calculation will be terminated. In the prior art, the above-mentioned shape mismatch cannot be detected in the pre-check operation, but will only cause the calculation to be terminated during the actual calculation process, which will cause a significant waste of transmission, storage and computing resources.
[0073] To this end, the present invention proposes an improved neural network calculation scheme, which constructs a shape flow graph based on the existing calculation graph, obtains the constraints of the eigenvalue shape flow and discovers shape errors based on this. In this way, shape errors that could only be discovered in actual calculation operations due to termination of calculations can be found through pre-check, thereby avoiding unnecessary transmission, storage and computing resource consumption.
[0074] Figure 4 FIG. 1 shows a schematic diagram of a neural network computing system using the present invention. Figure 4 As shown, the neural network computing system 400 includes a neural network platform 410 and a client 420. The client 420 can operate on the neural network platform to obtain computational graph construction information, and can also transmit the completed computational graph construction information together with training data (for example, in image classification tasks, the training data includes training images and training labels) to the platform 410. The platform 410 can be a large-scale neural network computing service platform (i.e., a server).
[0075] The server 410 can pre-check (e.g., debug) the computational graph construction information uploaded by the user to find out potential errors, and perform neural network calculations (e.g., neural network training, or image classification tasks based on the trained network) when there are no errors in the pre-check.
[0076] In the prior art, the pre-check module on server 410 cannot detect shape errors. Therefore, the computational graph construction information containing shape errors, along with the training data, is loaded into server 410's memory for computation, and the computation terminates due to the error. In other words, the inability to detect shape errors in advance results in inefficient data transmission, memory loading, and computation, resulting in a significant waste of resources.
[0077] To this end, the neural network computing platform of this application will be equipped with additional tensor shape tracking modules, which can be used to operate based on dynamic computational graphs, and thus can find shape errors through pre-checking before the neural network calculation is actually run, thereby avoiding the waste of resources caused by loading computational graph composition information containing shape errors for calculation.
[0078] This application can first be implemented as a neural network tensor shape tracking method. Figure 5 A schematic flowchart of a neural network tensor shape tracking method according to an embodiment of the present invention is shown.
[0079] In step S510, a shape flow graph is constructed based on the intermediate representation (IR) of the computational graph of the neural network. The shape flow graph includes nodes for performing neural network computation operations and edges representing dependencies between nodes. The edges are annotated with shape information of tensors to flow along the edges. Here, the shape flow graph can be an annotation of tensor shape information added to the computational graph, for example, as follows Figure 7 As shown in the figure, [batch_size, 28, 28 1] can indicate that the initial input to the convolutional layer is a four-rank tensor, where the input image is, for example, a grayscale image of 28x28 pixels, and the number indicated by batch_size is fed in at one time.
[0080] In step S520, constraints on the tensor shape information are extracted from the node operations and newly introduced operands of the shape flow graph. According to the basic knowledge of CNN mentioned above, the input data (for example, the input graph) is propagated layer by layer along the computational graph in the form of feature graphs. These feature graphs are in the form of 2-dimensional or higher tensors before being fed into the classifier. These feature graph tensors flow into the nodes along the shape flow graph and perform computational operations with the operands introduced by the node operations (for example, in the convolution node, convolution operations are performed with the newly introduced convolution kernel operands) or they operate on their own under the provisions of the node operations (for example, reconstruction with an unchanged number of elements). These node operations and operands can all trigger constraints. For example, the convolution operation with the convolution kernel requires that the rank of the input feature graph tensor is the same as that of the convolution kernel tensor, and the reconstruction operation requires that the number of tensor elements remain unchanged. These can all be used as constraints for the following determination.
[0081] In step S530, it is determined whether there are errors or restrictions in the shape information of the tensor based on the constraints. Here, the shape error can be a contradiction between the constraints. For example, in Example 2 below, the rank of the input tensor behavior_input needs to be equal to 4 and 3 at the same time. Since the constraint cannot be realized, the error is determined. Shape restriction can refer to the strict restrictions encountered when the user inputs the shape parameters. For example, the constraint extracted in Example 1 below is batch_size*32==batch_size∨batch_size==1, which means that the constraint condition is established only when batch_size is 1. This limits the tensor shape batch_size input by the user to be strictly equal to 1. Although this restriction can be realized, it is obviously not feasible in training. Therefore, it is also necessary to make a judgment and make the following report.
[0082] In step S540 , when it is determined that the shape information of the tensor has an error or limitation, a report is made.
[0083] The tensor shape tracking method of the present invention performs the determination in the absence of eigenvalue flow. In other words, the method is performed in a pre-check step before the actual execution of neural network calculations (e.g., training). For example, the method can be incorporated into the compilation step of generating IR, and is particularly applicable to situations where unknown information exists in the shape information of the tensor. Therefore, extracting constraints on the tensor shape information from the node operations and newly introduced operands of the shape flow graph can include: listing multiple constraints on the unknown information based on the relationship between each node operation and operand of the shape flow graph and the unknown information; and solving the unknown information based on the multiple constraints: if the unknown information has no solution, determining that a shape error exists; if the unknown information has a specific solution, determining that a shape restriction exists.
[0084] In one embodiment, the shape flow graph is constructed by slicing backward starting from the input tensor represented in the middle of the computation graph, and the operands required for node operations are collected during the backward slicing process, and the operands include constants and scalar variables. When a call is encountered, the return value of the call node operation can be added to the shape flow graph as a new node in the backward slicing, and the backward slicing continues after the newly added node. When an execution branch is encountered, all previous slices of the confluence node with n incoming values can be copied into n pieces; and each copied previous slice selects a different incoming value from the n incoming values to continue slicing, thereby constructing a shape flow graph for each path in the computation graph.
[0085] In one embodiment, reporting when it is determined that there is an error or limitation in the shape information of the tensor may include reporting the node where the error occurs, and to this end may include deleting the node transformation operations and the constraints introduced by the node transformation operations one by one; and locating the node operation that caused the error or limitation when the remaining constraints are met.
[0086] The above neural network tensor shape tracking method of the present invention can be implemented as a neural network tensor shape tracking module, which can be incorporated into the neural network compiler and located after the intermediate information generation module. To this end, Figure 6 This section shows an implementation example of a neural network tensor shape tracker according to one embodiment of the present invention. In the prior art, an intermediate representation generation module processes information constructed based on the computational graph from the application and generates an IR representation of this information. For example, the present invention can use Ariadane as its front-end to parse Python programs into WALAIR.
[0087] In order to decouple from the deep learning computing framework, it is necessary to construct a computational graph structure corresponding to the neural network processor. To this end, the compiler needs to abstract an intermediate representation (Intermediate Representation, or IR) that is independent of the framework, which can not only represent all the information of the algorithm, but also facilitate the subsequent compiler to optimize it based on the hardware platform. The shape tracker of the present invention can be located after the intermediate representation generation module, and includes the above-mentioned construction module for converting the IR representation into a shape flow graph, a constraint determination module for generating constraints and determining shape errors accordingly, and a reporting module for reporting when shape errors are determined.
[0088] In order to deepen the understanding of the principles of the present invention, the following will combine three code examples and Figure 7 and 8 The following describes an example of a shape flow graph shown in FIG. Figure 7 An example of a shape flow graph constructed according to the present invention is shown. Figure 8 Another example of a shape flow graph constructed according to the present invention is shown. Figure 7 The shape flow graph shown can be based on Figure 2 The computational graph shown is obtained. The computational graph can be described in various ways, and the above description can constitute the computational graph composition information of the present invention. In one embodiment, the computational graph can be described using the Python language under the deep learning framework TensorFlow.
[0089] Code Example 1
[0090]
[0091] The above example shows how to describe a computational graph using Python in the TensorFlow framework. TensorFlow programs typically written in Python consist of two phases: construction and execution. In Example 1, during the construction phase (lines 1-14), a computational graph is configured: each operator (e.g., tf.matmal on line 3) generates some nodes and edges that connect the data between the nodes. During the execution phase (lines 15-18), a session object is created to instantiate the graph, which is executed multiple times (sess.run on line 18), and data is input into placeholders (e.g., in_x and in_y on lines 10 and 11, respectively).
[0092] Figure 2 and Figure 7 can be considered as being based on the code example 1 above. Figure 2 The difference is that in Figure 7 In the depicted graph, each edge is annotated with the shape information of the tensor it propagates. The shape of a tensor can depend on the input: a placeholder tensor can have some dimensions (or the entire shape) set to none, and its shape will be instantiated when the graph is executed by feeding data to the placeholder (using the feed_dict operator, e.g., line 18).
[0093] The computation graph can be executed by calling Session.run(), Tensor.eval(), or Operation.run(). At line 18 as described above, the graph is executed to obtain the result of the operator train_step (line 14). The input data train_image and train_lab (line 15) are input to the placeholders in_x (line 10) and in_y (line 11), respectively. The first dimension of the input data is configured by the input parameter batch_size shown in the box in line 15. Therefore, the shapes of the tensors in_x and in_y are [batchsize, 28, 28, 1] and [batchsize, 10], respectively.
[0094] The tensor in_x is passed as an actual parameter to the function predict() in line 12 and processed by the convolution kernel operator conv2d (line 5). The conv2d operator is often used to extract intermediate features in complex neural networks. It requires a 4-dimensional input tensor, a 4-dimensional filter tensor, and a stride vector with 4 elements as input. Using the "same" padding strategy, cond2d(x, f, s,"same") will produce a tensor with a shape of [x[0], x[1] / s1, x[2] / s2, f[3]. Henceforth, the symbol x[i] is used to represent the i-th dimension of the shape of the tensor x, and the symbol s is used. iRepresents the i-th element of vector s. Figure 7 In the example, the conv2d operator generates a tensor mp of shape [batchsize, 28, 28, 32].
[0095] The reshape operator on line 6 changes the shape of the incoming tensor mp to the specified shape [-1, 28*28], which is a two-dimensional array. Here, the special dimension size -1 indicates that the size of the corresponding dimension needs to be dynamically calculated. If the size of the tensor (the total number of items in the tensor) is the same as the size of the specified shape, the tensor can be reshaped correctly. On line 6, after reshaping, we have a new tensor with a reshaped shape of [batchsize*32, 28*28].
[0096] On line 7, the full_connect function is called with the reshaped tensor as its actual parameter. Therefore, the operators get_variable (line 2) and matmul (line 3) are included in the computational graph. The matmul operator multiplies reshape([batchsize*32, 28*28]) with fc_w([28*28, 128]) to obtain a new tensor fc with a shape of [batchsize*32, 128]. Next, the tensor fc is processed again by the same function on line 8. Finally, logit([batchsize*32, 10]) is generated and returned as y.
[0097] The operator softmax_cross_entrophy_with_logits (line 13) generates normalized probabilities from the input tensors in_y([batchsize, 10]) and y([batchsize*32, 10]). It supports the "broadcasting" rule: the sizes of the matching dimensions must be the same, or one of them is 1 (in this case, the generated tensor adopts the other size in its corresponding shape dimension). Therefore, although the batch_size is unknown, the constraint on batch_size in the example can be solved, that is, the operator can only succeed when the size of the first shape dimension of the two tensors is the same or one of them is 1 (that is, batch_size*32 == batch_size ∨batch_size == 1). During subsequent actual training, because the user is unaware of the error (or strict limit) in the tensor shape and configures the input batch size to 200 (a typical value), the application fails with a runtime exception.
[0098] Code Example 2
[0099] 1.user_profile_cnn=tf.reshape(tmp_user_profile_cnn, shape=[-1, num_behavior_max[behavior_cnt], n_output_behavior, 1])
[0100] 2.attention_layer_input=tf.matmul(behavior_input, user_profile_cnn)
[0101] …
[0102] 3.tmp_attention_weights=tf.reshape(attention_weights, shape=[-1, num_behavior_max[behavior_cnt], 1])
[0103] 4.behavior_output=tf.matmul(tmp_attention_weights, behavior_input)
[0104] In this example, the tensor behavior_input comes from user input, and its shape (i.e., its rank and dimension size) is completely unknown. In line 1, the tensor tmp_user_profile_cnn is reconstructed into a 4-dimensional tensor user_profile_cnn, which is then multiplied with behavior_input through the operator matmul, indicating that the rank of behavior_input is also 4, otherwise the operator will fail. In line 3, the tensor attention_weights is reconstructed into a 3-dimensional tensor tmp_attention_weights, which is also multiplied with behavior_input. Since the input tensor behavior_input cannot meet these two constraints, the application will always fail on one of the matmul operators (either line 2 or line 4) during subsequent operations.
[0105] Code Example 3
[0106]
[0107] In this example, the tensor labels is a two-dimensional array and the tensor pred is a one-dimensional array. All their dimensions are unknown. When the condition loss_type == "mae" holds, the bug is triggered in line 2 because the operator absolute_difference expects the input tensors to have the same shape (i.e., the same rank and dimension size). However, if the other branch is taken (when the input parameter loss_type is "logloss"), the error will not be triggered. The operator sparse_softmax_cross_entropy_with_logits allows the rank of the input parameter pred to be 1 less than the rank of labels. Therefore, it will not trigger an error.
[0108] In the above code examples 1-3, there are tensors with completely unknown shapes (Example 2) or partially unknown shapes (Examples 1 and 3). Tensors with unknown or partially unknown shapes are often found in real-world applications. They can come from command line input, files, or unsupported library functions. It is difficult to write rules and infer these unknown shapes into a limited set of specific shapes. Therefore, the present invention proposes a shape error tracking method that is particularly suitable for unknown tensor shapes in computational graphs. The method symbolically represents the shape of the tensor by introducing symbolic values (i.e., shape value symbols) for the shape of tensors of unknown rank or unknown dimension size. In other words, the unknown shape information about the tensor can be set to an unknown number represented by a symbol. Constraints can be introduced from tensor operators, scalar variables, and conditional branches. Finally, a constraint solver is applied to check the satisfiability of these constraints.
[0109] For example, for Example 1 ( Figure 7 ), the value of the input variable batch_size is symbolically represented, for example, as X. Here, batch_size can be regarded as Figure 7 The unknown information (or unknowns) in the shape flow graph can be represented by X. The computational graph will generate the constraint X*32==X∨X==1, as well as other constraints. The solution to the constraint is X=1, which can be provided to the user as a warning. In Example 2, the rank of the tensor behavior_input is symbolized as The two matmul operators in lines 2 and 4 will introduce their respective constraints. and Since both cannot be satisfied simultaneously, an error is found.
[0110] like Figure 6As shown, the shape tracker of the present invention may include a construction module, a constraint determination module, and a reporting module. In actual operation, the construction module first traverses the path given by the IR and builds a shape flow graph (an abstract computational graph) for each path. Next, the constraint determination module formulates the shape flow graph into a constraint list, which is then solved using a constraint solver. Finally, if the constraint is unsatisfactory (if the user input is constrained), an error (warning) is issued. In order to accurately report the line number where the error / warning occurs, the reporting module can search for the first operator that introduces an unsatisfactory constraint and report it to the user.
[0111] Specifically, the Builder module builds a shape flow graph for each program path. Since the control flow structure of TensorFlow programs is usually simple, the number of program paths is mostly 2 and rarely reaches 8.
[0112] Specifically, the shape flow graph of the present invention can be an abstract computational graph annotated with shape information. To construct the shape flow graph, it is possible to slice backward from the call to session.run() (i.e., from the output tensor). In other words, Figure 6 As shown, the computational graph construction information generated by the application can be the program code shown in Example 1, and IR is obtained by compiling the compiler. The IR (or the use-def chain obtained therefrom) can be used by the construction module to construct the shape flow graph of the present invention. Since TensorFlow programs usually propagate values directly through assignment or parameter passing, they can be sliced along the use-def chain represented by WALA's SSA (single static assignment). During backward slicing, function calls can be inlined: when a function call is encountered, the return value of the called function is added to the graph (as a new node), and then the slicing can continue backward from the newly added return value. Finally, all operators (i.e., TensorFlow API calls) that the output tensor transfer depends on, and the operands introduced by the operator (e.g., tensors and scalars (e.g., actual parameters of the operator)) are included in the graph.
[0113] The following will refer to Example 1 and Figure 7The shape flow graph construction describing the building block. The output tensor train_step can be sliced from sess.run() line 18. Since train_step is returned from the operator minimize, the operator and its operands (i.e. cross_entropy) are added to the graph. Similarly, cross_entropy is produced by the operator softmax_cross_entropy_with_logits (line 13). Therefore, the operator and its operands (in_y and y) are included. Starting from y, the function call is inlined to predict (line 12) and continues to be sliced from its return value logit (line 9). Next, the function def_fully_connect is inlined twice in that order on lines 8 and 7. The final shape flow graph is as follows Figure 7 shown.
[0114] When there is a phi node (control flow confluence point in SSA) in the computation graph, the shape flow graph needs to be copied. When a phi node with n incoming values is encountered, the graph is copied n times, and each graph selects a different incoming value to continue slicing. Figure 8 The shape flow graph for Example 2 is given. In SSA, there is a phi node at the confluence of different branches of the if statement (lines 1-4). Therefore, two shape flow graphs are obtained, one for each branch, namely branch 0 and branch 1 in the figure.
[0115] Loops, although rare during the graph construction phase, are handled by unrolling the loop twice.
[0116] When collecting shape information, constants and scalar variables that are propagated along the use-def chain can be directly recorded. An attempt can be made to infer as much specific information as possible by applying constant propagation and computing specific shape information based on the documented semantics of the TensorFlow API. In some embodiments, the following two special cases can also be considered. First, the shape of a tensor can be set using the tf.setshape() function, so for each tensor, its use of the object's tf.setshape() call is checked and the shape of the tensor is updated accordingly. Second, in most cases, values are propagated directly. However, when a tensor is initialized with a given shape, the value is passed as an argument to the shape's constructor and stored in its corresponding field. Typically, pointer analysis is required to calculate field-related dependencies. However, such fields of a shape object are only stored once in its constructor (during initialization). Therefore, when a field load is encountered, only the unique storage for the corresponding field needs to be searched.
[0117] The constraint judgment module formulates the shape flow graph into a list of constraints and then performs shape error judgment based on these constraints. Constraints can be collected from tensor operators and scalar instructions in the shape flow graph. Since shape-related values rarely depend on conditions in real-world TensorFlow applications, branch conditions can be ignored.
[0118] For ease of understanding, some symbolic representations are defined below to represent the shape of the tensor T, the value of the vector V, and the shape of the scalar X. Specifically, T[0], T[-1], T[-], ... can represent the dimensions of T; |T| represents the total size of T (i.e., the number of elements); Indicates the rank (number of dimensions) of T, V0, V1, ..., V |V|-1 represents the element value of V; X represents the value of X.
[0119] Specifically, T[-1] represents the size of the last shape dimension of T. This variable is particularly useful when the rank of T is unknown, i.e. By default, all variables are assumed to be represented symbolically unless otherwise specified. If C is a constant value, then a variable is introduced for each dimension of T and T is materialized by applying the following function:
[0120]
[0121] This function sets the rank of T to C, sets T[-1] to the size of the last dimension of T (T[-1] == T[C-1]), and materializes the size of T to the product of the sizes of all its dimensions.
[0122] Introduce constraints for operators according to their documented semantics. For example, the operator C=reshape(A,B) reshapes tensor A into a tensor C of the same size, with the shape specified by vector B. Therefore, we can get:
[0123]
[0124] Here, the constraint |C| == |A| states that tensors C and A have the same size (as required by reshape ), and the remaining constraints specify the shape of C in terms of vector B: the rank of C is defined by the size of B The dimensions of C are defined by the elements of B (∧ 0≤i≤|B| C[i]==B i). Note that the size of B, i.e., |B|, is a constant. Therefore, the tensor C is materialized (Concretize(C,|B|)). Except for one element value (e.g., -1), all other element values are constants. The same list of constraints applies when the rank of tensor A is unchanged, that is, when A has been materialized.
[0125] We can then examine the operator softmax_cross_entropy_with_logits, abbreviated as C = logits(A,B), which supports Numpy "broadcasting". Before delving into the details of the tricky broadcasting semantics, we first introduce another auxiliary function Broadcast(A,B,C,i,j). This function expresses a constraint on the i-th dimension of a higher-rank input tensor A, the i-th dimension of the output tensor C, and the matching j-th dimension of another input tensor B, where j ≤ i:
[0126] Broadcast((A, B, C, i, j): ((A[i]==B[j]∧C[i]==A[i])
[0127] ∨(A[i]==]∧C[i]==R[j])V(B[j]==1∧C[i]==A[i]))
[0128] Broadcast((A,B,C,i)) holds if any of the following three cases holds: 1) A[i] matches B[j], yielding the same size for the i-th dimension of C, 2) A[i] is 1, in which case the size of the i-th dimension of C is taken from the j-th dimension of B, and 3) B[j] is 1, in which case the size of the i-th dimension of C is taken from the size of A. These three cases reproduce the semantics of broadcasting a pair of matching dimensions (A[i] and B[j]).
[0129] The constraint list for C = logits(A,B) is given by:
[0130]
[0131] when or When is a symbol, the above constraints apply. In this case, only constraints on the sizes of the first and last dimensions of the input and output tensors are introduced. In the case of , the constraints apply to the first and last dimensions of all three tensors A, B, and C (Broadcast<A,B,C,0> ∧Broadcast<A,B,C,-1> In the other two cases, the constraints are applied to the last dimension of the three tensors, and the output tensor C takes its size from the higher-rank tensor (e.g., A[0] == C[0] when ).
[0132] when and When both are constants (i.e., A and B are concretized), C can be concretized as follows:
[0133]
[0134] exist In the case of , the output tensor C is materialized, with each dimension defined by the Broadcast rule. Otherwise, C is materialized with a higher rank (e.g., when ), higher dimensions are directly copied from higher rank tensors and the matching dimensions are broadcast
[0135] Similarly, appropriate constraints can be introduced for other operators, such as conv2d and matmul. Finally, the constraints of the shape flow graph are fed into the constraint solver. For example, according to the constraint batch_size*32==batch_size∨batch_size==1, it is concluded that Example 1 will not have shape errors only when batch_size==1; and according to the constraints of the two matmul operators, and It is concluded that Example 2 has a shape error.
[0136] If the constraint solver of the constraint evaluator cannot solve a given set of constraints (for example, if the user input is constrained by constant values), the reporting module issues an error (warning). To accurately report the error location, the reporting module searches for the operator that introduced the discovered unsatisfiable constraints. Specifically, this is achieved by removing each operator (more precisely, by removing the constraints introduced by each operator) one by one, in the reverse order in which they were added to the underlying dataflow graph, until the constraints become satisfiable. In practice, this process can be accelerated using binary search.
[0137] As above, combined with Examples 1-3 and the attached Figure 7-8 The application of the present invention in actual operation is described. Furthermore, the present invention can also be implemented as a neural network calculation method. Figure 9 FIG. 9 is a schematic diagram showing the composition of a neural network computing platform according to an embodiment of the present invention. The computing platform 900 may include the aforementioned Figure 6The illustrated implementation is a tensor shape tracker that is part of a compiler, or a separate tensor shape tracking device. The tracker or device 910 includes: a construction module 911 for constructing a shape flow graph based on an intermediate representation (IR) of a computational graph of a neural network, wherein the shape flow graph includes nodes for performing neural network computation operations and edges representing dependencies between nodes, wherein the edges are labeled with shape information of tensors to flow along the edges; a constraint determination module 912 for extracting constraints on tensor shape information from node operations and newly introduced operands in the shape flow graph, and determining whether the shape information of the tensor has errors or restrictions based on the constraints; and a reporting module 913 for reporting when it is determined that the shape information of the tensor has errors or restrictions.
[0138] The computing platform 900 may also include a neural network training module 920 for performing neural network training calculations based on the intermediate representation of the computational graph in which the errors or limitations are eliminated according to the operation of the tensor shape tracking module. Although not shown in the figure, the computing platform may also include: a training data acquisition module for acquiring training data and training labels, wherein the training data serves as the initial tensor of the flow tensor in the computational graph, and the training labels are used to reversely adjust the classification results output by the computational graph, wherein, in response to the tensor shape tracking module determining that there are no errors or limitations, the neural network training module loads the training data and the training labels to perform neural network training calculations. In other words, the shape tracking solution of the present invention can only load training data and perform actual training calculations after confirming that there are no shape errors.
[0139] Thus, the neural network computing platform of the present invention traverses the computational graph paths and constructs a shape flow graph (an abstract data flow computational graph) for each path. It then uses a constraint to solve shape-related constraints (introduced by shape operators) for each shape flow graph and reports an error when the constraint cannot find a feasible solution. Unlike existing techniques based on static analysis, the present invention's constraint-based shape error determination scheme can detect subtle shape-related errors even when the shape's rank (number of dimensions) or dimensionality (size of the dimensions) is completely unknown.
[0140] It should be understood that Figure 9 The computing platform shown can be a distributed computing platform composed of large servers including powerful CPUs and even dedicated GPUs. In some embodiments, it can also be implemented as a powerful local device. Furthermore, because the constraint-based shape error determination scheme of the present invention does not need to be operated when the neural network calculation is actually performed, it can be implemented as part of the pre-check function of the computing platform.
[0141] Figure 10A schematic structural diagram of a computing device that can be used to implement the above-mentioned shape tracking method according to an embodiment of the present invention is shown.
[0142] See also Figure 10 , the computing device 1000 includes a memory 1010 and a processor 1020 .
[0143] Processor 1020 may be a multi-core processor or may include multiple processors. In some embodiments, processor 1020 may include a general-purpose main processor and one or more special coprocessors, such as a graphics processing unit (GPU), a digital signal processor (DSP), etc. In some embodiments, processor 1020 may be implemented using customized circuits, such as application-specific integrated circuits (ASICs) or field-programmable gate arrays (FPGAs).
[0144] The memory 1010 may include various types of storage units, such as system memory, read-only memory (ROM), and permanent storage devices. ROM may store static data or instructions required by the processor 1020 or other modules of the computer. The permanent storage device may be a readable and writable storage device. The permanent storage device may be a non-volatile storage device that retains stored instructions and data even when the computer is powered off. In some embodiments, the permanent storage device uses a large-capacity storage device (such as a magnetic or optical disk, flash memory) as the permanent storage device. In other embodiments, the permanent storage device may be a removable storage device (such as a floppy disk, optical drive). The system memory may be a readable and writable storage device or a volatile readable and writable storage device, such as dynamic random access memory. The system memory may store some or all instructions and data required by the processor during operation. In addition, the memory 1010 may include any combination of computer-readable storage media, including various types of semiconductor memory chips (DRAM, SRAM, SDRAM, flash memory, programmable read-only memory), and magnetic disks and / or optical disks may also be used. In some embodiments, the memory 1010 may include a readable and / or writable removable storage device, such as a compact disc (CD), a read-only digital versatile disc (e.g., DVD-ROM, double-layer DVD-ROM), a read-only Blu-ray disc, an ultra-density optical disc, a flash memory card (e.g., SD card, miniSD card, Micro-SD card, etc.), a magnetic floppy disk, etc. Computer-readable storage media do not include carrier waves and transient electronic signals transmitted wirelessly or wired.
[0145] The memory 1010 stores executable codes. When the executable codes are processed by the processor 1020 , the processor 1020 can execute the shape tracking method described above.
[0146] The neural network computing platform and method according to the present invention have been described in detail above with reference to the accompanying drawings. The shape tracker of the present invention traverses program paths and constructs a shape flow graph (an abstract data flow computation graph) for each path. A constraint solver is then used to solve shape-related constraints (introduced by shape operators) for each shape flow graph. If the constraint solver cannot find a feasible solution, an error is reported. If the user input is subject to constraints, suggestions are provided as warnings.
[0147] In addition, the method according to the present invention may also be implemented as a computer program or a computer program product, which includes computer program code instructions for executing the above steps defined in the above method of the present invention.
[0148] Alternatively, the present invention can also be implemented as a non-transitory machine-readable storage medium (or computer-readable storage medium, or machine-readable storage medium) on which executable code (or computer program, or computer instruction code) is stored. When the executable code (or computer program, or computer instruction code) is executed by a processor of an electronic device (or computing device, server, etc.), the processor executes the various steps of the above-mentioned method according to the present invention.
[0149] Those skilled in the art will further appreciate that the various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the disclosure herein may be implemented as electronic hardware, computer software, or combinations of both.
[0150] The flowcharts and block diagrams in the accompanying drawings show the possible implementation architecture, functions and operations of the systems and methods according to multiple embodiments of the present invention. In this regard, each box in the flowchart or block diagram can represent a module, program segment or part of code, and the part of the module, program segment or code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0151] While various embodiments of the present invention have been described above, the foregoing description is intended to be illustrative, non-exhaustive, and not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. A pair of neural network tensor shape tracking methods, including: Constructing a shape flow graph based on an intermediate representation (IR) of a computational graph of a neural network, the shape flow graph including nodes for performing neural network computation operations and edges representing dependencies between the nodes, the edges being annotated with shape information of tensors to flow along the edges; wherein the shape information includes a symbolically represented shape of a tensor of unknown rank or unknown dimension; Extracting constraints on shape information of tensors from node operations of the shape flow graph and newly introduced operands; Based on the constraint condition, in the absence of actual eigenvalue flow, determining whether shape information of the tensor has errors or limitations; and reporting if it is determined that shape information of the tensor is erroneous or limited; The constructing of the shape flow graph comprises: Starting from the input tensor of the intermediate representation of the computation graph, slicing backwards to build the shape flow graph, and The operands required for node operations are collected during the backward slicing process, where the operands include constants and scalar variables.
2. The method according to claim 1, wherein The shape information of the added corresponding tensor contains unknown information, and The constraints for tensor shape information extracted from the node operations and newly introduced operands of the shape flow graph include: Based on the relationship between the node operations and operands of the shape flow graph and the unknown information, a plurality of constraints for the unknown information are listed. The unknown information is solved based on the multiple constraint conditions: if the unknown information has no solution, it is determined that a shape error exists; if the unknown information has a specific solution, it is determined that a shape restriction exists.
3. The method according to claim 1, wherein The constructing of the shape flow graph comprises: In the backward slicing, the return value of the calling node operation is added to the shape flow graph as a new node, and the backward slicing is continued after the newly added node.
4. The method according to claim 1, wherein The constructing of the shape flow graph comprises: Duplicate all preceding slices of a join node with n incoming values into n copies; and Each copied previous slice selects a different one of the n incoming values to continue slicing, and a shape flow graph is constructed for each path in the computation graph.
5. The method according to claim 1, wherein The reporting when determining that the shape information of the tensor has an error or limitation includes: Deleting node transformation operations and constraints introduced by the node transformation operations one by one; and The node operation that caused the error or restriction is located when the remaining constraints are satisfied.
6. A neural network computing platform comprising: Tensor shape tracking module, including: A construction submodule is configured to construct a shape flow graph based on an intermediate representation (IR) of a computational graph of a neural network, wherein the shape flow graph includes nodes for performing neural network computation operations and edges representing dependencies between the nodes, wherein the edges are annotated with shape information of tensors to flow along the edges; wherein the shape information includes a symbolically represented shape of a tensor of unknown rank or unknown dimension size; a constraint determination submodule, configured to extract constraints on tensor shape information from node operations and newly introduced operands of the shape flow graph, and determine whether the shape information of the tensor has errors or limitations based on the constraints in the absence of actual eigenvalue flow; and A reporting submodule, configured to report when it is determined that the shape information of the tensor has errors or limitations, a neural network training module for performing neural network training calculations based on the intermediate representation of the computational graph in which the errors or limitations are eliminated according to the operation of the tensor shape tracking module; The construction submodule is specifically used to slice backward from the input tensor of the intermediate representation of the computational graph to construct the shape flow graph, and collect the operands required for node operations during the backward slicing process, wherein the operands include constants and scalar variables.
7. The computing platform of claim 6, further comprising: The training data acquisition module is used to obtain training data and training labels. The training data is used as the initial tensor of the flow tensor in the calculation graph, and the training label is used to reversely adjust the classification results output by the calculation graph, wherein: In response to the tensor shape tracking module determining that there are no errors or limitations, the neural network training module loads the training data and the training labels to perform neural network training calculations.
8. A non-transitory machine-readable storage medium having executable code stored thereon, wherein when the executable code is executed by a processor of an electronic device, the processor is caused to perform the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
TensorFlow program vulnerability detection method and device and electronic equipment
CN113221126A