Method for data processing, electronic device and storage medium
By segmenting the high-order information matrix of the deep learning model, the Cronek factor matrix is converted into multiple small square matrixes in parallel processing, the problem of high computational complexity of high-order optimization algorithms is solved, and faster model training is achieved.
Patent Information
- Application Number
- CN202010480676.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-05-30
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2040-05-30
AI Technical Summary
The existing high-order optimization algorithms have high computational complexity in deep learning model training, making it difficult to reduce time costs while ensuring training accuracy.
By segmenting the Cronek factor matrix of the higher-order information matrix of the neural network model, converting it into multiple small square matrixes, reducing the computation time complexity, and adjusting the model parameters by processing the inverse matrixes of these square matrixes in parallel.
It significantly reduces the time cost of model training and improves computing efficiency, especially during training of large-scale models such as resnet50, reducing the time of a single iteration.
Smart Images

Figure CN113743571B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure mainly relate to the field of computer technology, and more specifically, to methods, devices, electronic devices, and computer-readable storage media for data processing. Background Art
[0002] With the continuous development of computer technology, deep learning technology has been applied to various fields. Currently, deep learning has shown excellent performance in many application fields, such as image recognition, object detection, and natural language processing.
[0003] The training of deep learning models has become the focus of current attention. Common optimization algorithms for deep learning models include first-order optimization algorithms (e.g., gradient descent algorithm) and higher-order optimization algorithms (e.g., natural gradient algorithm). The first-order optimization has a poor convergence speed. In contrast, higher-order optimization algorithms usually can bring better training accuracy, but require a greater time cost. Summary of the Invention
[0004] Embodiments of the present disclosure provide a solution for data processing.
[0005] In a first aspect of the present disclosure, a method for a data processing device is provided. The method includes: obtaining a Kronecker factor matrix for indicating a higher-order information matrix of a neural network model, where the higher-order information matrix is used to correct a first-order gradient of the neural network model; slicing the Kronecker factor matrix to obtain a plurality of square matrices, such that the plurality of square matrices are sub-matrices of the Kronecker factor matrix, and the main diagonals of the plurality of square matrices respectively correspond to a part of the main diagonal of the Kronecker factor matrix; and adjusting parameters of the neural network model based on the plurality of square matrices.
[0006] In the present disclosure, the higher-order information matrix refers to a matrix generated during the training process of the neural network model for correcting the first-order derivative of the model, such as the Hessian matrix, Fisher matrix, and second-moment matrix. The Kronecker factor matrix refers to a pair of matrices obtained by performing Kronecker decomposition on the higher-order information matrix, and the Kronecker product of the pair of symmetric matrices is equal to the higher-order information matrix.
[0007] In the present disclosure, the slicing can be performed along the main diagonal of the Kronecker factor matrix according to a predetermined slicing dimension to obtain a plurality of square matrices with a rank equal to the slicing dimension. In some cases, when the rank of the Kronecker factor matrix is divisible by the slicing dimension, the Kronecker factor will be sliced into an integer number of square matrices with the same rank. In other cases, when the rank of the Kronecker factor matrix is not divisible by the slicing dimension, the Kronecker factor will be sliced into an integer number of square matrices with the same rank and one or more square matrices with a rank less than the slicing dimension.
[0008] Based on such partitioning, the main diagonals of multiple square matrices respectively correspond to a part of the main diagonal of the Kronecker factor matrix. That is, all the elements on the main diagonals of the multiple square matrices will include all the elements on the main diagonal of the Kronecker factor matrix. Additionally, the square matrices with rank greater than 1 among the multiple square matrices also include other elements in the Kronecker factor matrix, and these elements are the elements adjacent to the main diagonal in the Kronecker factor matrix.
[0009] Additionally, the data processing device in this disclosure is a device with AI data computing capabilities, which can be a terminal device or a network device.
[0010] In this way, the embodiments of this disclosure approximate the Kronecker factor matrix into multiple small square matrices, thereby greatly reducing the time complexity of the operation and thus reducing the time cost of model training.
[0011] In some implementation manners of the first aspect, the data processing device includes computing resources for performing the adjustment, and the partitioning of the Kronecker factor matrix includes: partitioning the Kronecker factor matrix based on the resource identifier of the computing resources and the model identifier of the neural network model.
[0012] In some embodiments, the model identifier may be identification information indicating the type of the neural network model. For example, "resnet50" can indicate that the neural network model is a 50-layer deep residual network (ResNet). Examples of the model identifier may also include "resnet18", "resnet101", "VGG16", "LeNet5", etc. Additionally, the resource identifier may be identification information indicating the model of the computing resources. For example, the chip name, chip code name, chip type number, or even the identifier of the device where the chip is located can be used as the resource identifier. As an example, the resource identifier may be the specific model of the AI chip used to train the neural network model, such as "GPU V100". Alternatively, the resource identifier may also be the unique identifier of the computing resources. For example, the MAC address of the computing resources, etc.
[0013] In some embodiments, the partitioning dimension for partitioning the Kronecker factor matrix can be determined based on the resource identifier and the model identifier, and the Kronecker factor is partitioned based on this partitioning dimension. For example, the corresponding historical optimization strategy can be queried according to the resource identifier and the model identifier to determine the partitioning dimension for partitioning the Kronecker factor matrix. In this way, the embodiments of this disclosure can use the historical optimization strategy corresponding to the model and the computing resources to quickly determine how to partition the Kronecker factor matrix, thereby further reducing the computing cost.
[0014] In certain implementations of the first aspect, a correspondence relationship among a resource identifier, a model identifier, and a dimension is stored in a data processing device, where the dimension indicates the rank of at least one square matrix obtained by partitioning a Kronecker factor matrix.
[0015] In some embodiments, the correspondence relationship among the resource identifier, the model identifier, and the dimension may be stored in a configuration file, for example. The correspondence relationship may be the correspondence relationship among the resource identifier, the model identifier, and the dimension, or the correspondence relationship between any two of them. By maintaining the correspondence relationship among the resource identifier, the model identifier, and the dimension, embodiments of the present disclosure can quickly implement lookups based on the resource identifier and the model identifier, and thus can efficiently determine available optimization strategies from historical optimization strategies.
[0016] In certain implementations of the first aspect, the data processing device includes computing resources for performing adjustments, and partitioning the Kronecker factor matrix includes: selecting a target dimension from multiple dimensions to partition the Kronecker factor matrix based on the performance information corresponding to the multiple dimensions, where the performance information corresponding to one dimension indicates the efficiency of the computing resources in processing the square matrix corresponding to the dimension, and the target dimension indicates the rank of at least one square matrix obtained by partitioning the Kronecker factor matrix.
[0017] In this application, the target dimension can be selected according to the performance information corresponding to multiple dimensions, and the Kronecker factor matrix can be partitioned based on this dimension. For example, the multiple dimensions may be a preset set of candidate dimensions. A set of sample square matrices corresponding to a set of candidate dimensions can be constructed, and the performance information corresponding to different dimensions can be determined by obtaining the efficiency of the computing resources in processing these sample square matrices.
[0018] In some embodiments, the performance information corresponding to each dimension can be determined, and the dimension whose performance information meets a predetermined requirement (for example, the performance is better than a specific threshold or is optimal) can be selected as the target dimension. In this way, embodiments of the present disclosure can ensure that the computing resources can efficiently process the multiple square matrices obtained by partitioning, thereby improving the computing efficiency.
[0019] In certain implementations of the first aspect, the value of the performance information corresponding to one dimension is related to the time required for the data processing device to calculate the inverse matrix of the square matrix corresponding to the dimension.
[0020] In this application, since the main time overhead of the computing resources comes from the inverse matrix operation of the square matrix, therefore, by considering the efficiency of the computing resources in solving the inverse matrix of the corresponding square matrix when selecting the target partitioning dimension, embodiments of the present disclosure can further reduce the time overhead required for solving the inverse matrix.
[0021] In certain implementations of the first aspect, partitioning the Kronecker factor matrix includes: selecting a target dimension from multiple dimensions to partition the Kronecker factor matrix based on the information loss corresponding to the multiple dimensions, where the information loss corresponding to one dimension indicates the information loss caused by partitioning the reference Kronecker factor matrix using that dimension.
[0022] In some embodiments, for example, each dimension in a set of candidate dimensions can be used to partition the reference Kronecker factor matrix to determine the information loss corresponding to each dimension. To more accurately evaluate the information loss during the actual training process, the reference Kronecker factor matrix can be, for example, the Kronecker factor matrix obtained after a predetermined number of iterations of the neural network model using the same training data.
[0023] In some embodiments, the information loss corresponding to each dimension can be determined, and the dimension whose information loss meets a predetermined requirement (for example, the information loss is less than a specific threshold or the least information loss) can be selected as the target dimension. In this application, by considering the information loss when selecting the target partitioning dimension, the embodiments of the present disclosure can reduce the time complexity of matrix inversion while ensuring the calculation accuracy of matrix inversion.
[0024] In certain implementations of the first aspect, the information loss corresponding to one dimension is the difference between the spectral norm of the joint matrix of the multiple square matrices obtained by partitioning the reference Kronecker factor matrix using that dimension and the spectral norm of the reference Kronecker factor matrix.
[0025] In this application, quantifying the information loss by the spectral norm enables the data processing device to more efficiently determine the target partitioning dimension.
[0026] In certain implementations of the first aspect, the data processing device includes computing resources for performing the adjustment, where partitioning the Kronecker factor matrix includes: selecting a target dimension from multiple dimensions to partition the Kronecker factor matrix based on the performance information corresponding to the multiple dimensions and the information loss corresponding to the multiple dimensions, where the performance information corresponding to the dimension indicates the efficiency of the computing resources in processing the square matrix corresponding to the dimension, the information loss corresponding to the dimension indicates the information loss caused by partitioning the reference Kronecker factor matrix using that dimension, and the dimension indicates the rank of at least one of the multiple square matrices obtained by partitioning the Kronecker factor matrix.
[0027] In some embodiments, the performance information and information loss corresponding to each dimension can be determined, and the target dimension can be determined based on the two. For example, the target dimension can be such that both the corresponding performance information and information loss meet predetermined conditions (for example, the performance is better than a specific threshold and the information loss is less than a specific threshold).
[0028] In some embodiments, the target dimension can also be determined, for example, based on the normalization results of performance information and information loss. For example, by constructing a fitting function between different candidate dimensions and performance information and / or information loss, the target dimension is determined by finding the intersection point of the fitting function. In this application, by considering both performance information and information loss, the embodiments of the present disclosure can achieve a balance between the time cost of solving the inverse matrix and the calculation accuracy, avoiding overly poor calculation accuracy or overly low calculation efficiency.
[0029] In some implementations of the first aspect, adjusting the parameters of the neural network model includes: processing multiple square matrices in parallel to determine the inverse matrices of the multiple square matrices; and adjusting the parameters of the neural network model based on the combination of the inverse matrices of the multiple square matrices.
[0030] In this application, since at least some of the square matrices obtained by partitioning will have the same rank, this provides good support for parallel processing. By processing these square matrices in parallel, the time cost of solving the inverse matrix can be further reduced.
[0031] In some implementations of the first aspect, where the neural network model is an image processing model, and where obtaining the Kronecker factor matrix includes: obtaining image training data; and applying the image training data to the image processing model to obtain the Kronecker factor matrix.
[0032] In this application, the object processed by the neural network model can be image data. Through the method of the present disclosure, the training of the image analysis model can be achieved more quickly, enabling the image analysis model to be put into use more quickly.
[0033] In some implementations of the first aspect, the neural network model is a text processing model, and where obtaining the Kronecker factor matrix includes: obtaining text training data; and applying the text training data to the text processing model to obtain the Kronecker factor matrix.
[0034] In this application, the object processed by the neural network model can be text data. Through the method of the present disclosure, the training of the text processing model can be achieved more quickly, enabling the text processing model to be put into use more quickly.
[0035] In a second aspect of the present disclosure, there is provided a data processing apparatus, including an acquisition unit configured to acquire a Kronecker factor matrix for indicating a high-order information matrix of a neural network model, where the high-order information matrix is used to correct the first-order gradient of the neural network model. The data processing apparatus further includes a slicing unit configured to slice the Kronecker factor matrix to obtain a plurality of square matrices, such that the plurality of square matrices are sub-matrices of the Kronecker factor matrix, and the main diagonals of the plurality of square matrices respectively correspond to a part of the main diagonal of the Kronecker factor matrix. In addition, the data processing apparatus further includes an adjustment unit configured to adjust the parameters of the neural network model based on the plurality of square matrices.
[0036] In the present disclosure, the high-order information matrix refers to a matrix generated during the training process of the neural network model for correcting the first-order derivative of the model, such as the Hessian matrix, Fisher matrix, and second moment matrix, etc. The Kronecker factor matrix refers to a pair of matrices obtained by performing Kronecker decomposition on the high-order information matrix, and the Kronecker product of the pair of symmetric matrices is equal to the high-order information matrix.
[0037] In the present disclosure, the slicing unit may slice along the main diagonal of the Kronecker factor matrix according to a predetermined slicing dimension to first obtain a plurality of square matrices with ranks equal to the slicing dimension. In some cases, when the rank of the Kronecker factor matrix is divisible by the slicing dimension, the Kronecker factor will be sliced into an integer number of square matrices with the same rank. In other cases, when the rank of the Kronecker factor matrix is not divisible by the slicing dimension, the Kronecker factor will be sliced into an integer number of square matrices with the same rank and a square matrix with a rank less than the slicing dimension.
[0038] Based on such slicing, the main diagonals of the plurality of square matrices respectively correspond to a part of the main diagonal of the Kronecker factor matrix, that is, all the elements on the main diagonals of the plurality of square matrices will include all the elements on the main diagonal of the Kronecker factor matrix.
[0039] In this way, the embodiments of the present disclosure convert the Kronecker factor matrix into an approximation of a plurality of small square matrices, thereby greatly reducing the time complexity of the operation and thus reducing the time cost of model training.
[0040] In certain implementations of the second aspect, the slicing unit is further configured to slice the Kronecker factor matrix based on the resource identifier of the computing resource and the model identifier of the neural network model.
[0041] In some embodiments, the model identifier may be identifier information indicating the type of the neural network model. For example, "resnet50" may indicate that the neural network model is a 50-layer deep residual network (ResNet). Examples of the model identifier may also include "resnet18", "resnet101", "VGG16", "LeNet5", and so on. Additionally, the resource identifier may be identifier information indicating the model number of the computing resource. For example, the chip name, chip code name, chip type number, or even the identifier of the device where the chip is located can be used as the resource identifier. As an example, the resource identifier may be the specific model of the AI chip used to train the neural network model, such as "GPU V100". Alternatively, the resource identifier may also be the unique identifier of the computing resource. For example, the MAC address of the computing resource, etc.
[0042] In some embodiments, the splitting unit may determine the splitting dimension for splitting the Kronecker factor matrix based on the resource identifier and the model identifier, and split the Kronecker factor based on this splitting dimension. For example, the splitting unit may query the corresponding historical optimization strategy according to the resource identifier and the model identifier to determine the splitting dimension for splitting the Kronecker factor matrix. Based on this manner, the embodiments of the present disclosure can utilize the historical optimization strategy corresponding to the model and the computing resource to quickly determine how to split the Kronecker factor matrix, thereby further reducing the computing cost.
[0043] In some implementations of the second aspect, a correspondence relationship among the resource identifier, the model identifier, and the dimension is stored in the data processing device, and the dimension indicates the rank of at least one of the square matrices obtained by splitting the Kronecker factor matrix.
[0044] In some embodiments, the correspondence relationship among the resource identifier, the model identifier, and the dimension may be stored in a configuration file, for example. The correspondence relationship may be the correspondence relationship among the resource identifier, the model identifier, and the dimension, or the correspondence relationship between any two of the three. By maintaining the correspondence relationship among the resource identifier, the model identifier, and the dimension, the embodiments of the present disclosure can quickly implement the lookup based on the resource identifier and the model identifier, and thus can efficiently determine the available optimization strategy from the historical optimization strategy.
[0045] In some implementations of the second aspect, the splitting unit is further configured to: select a target dimension from multiple dimensions to split the Kronecker factor matrix based on the performance information corresponding to the multiple dimensions, where the performance information corresponding to one dimension indicates the efficiency of the computing resource in processing the square matrix corresponding to this dimension, and the target dimension indicates the rank of at least one of the square matrices obtained by splitting the Kronecker factor matrix.
[0046] In some embodiments, the splitting unit may select a target dimension based on performance information corresponding to multiple dimensions, and split the Kronecker factor matrix based on this dimension. For example, the multiple dimensions may be a preset set of candidate dimensions. A set of sample square matrices corresponding to the set of candidate dimensions may be constructed, and the performance information corresponding to different dimensions may be determined by obtaining the efficiency of processing these sample square matrices using computing resources.
[0047] In some embodiments, the splitting unit may determine the corresponding performance information for each dimension, and select the dimension whose performance information meets a predetermined requirement (for example, the performance is better than a specific threshold or the optimal performance) as the target dimension. In this way, the embodiments of the present disclosure can ensure that the computing resources can efficiently process the multiple square matrices obtained by splitting, thereby improving the computing efficiency.
[0048] In certain implementations of the second aspect, the value of the performance information corresponding to a dimension is related to the time required for the data processing device to calculate the inverse matrix of the square matrix corresponding to the dimension.
[0049] In this application, since the main time overhead of the computing resources comes from the inverse matrix operation of the square matrix, therefore, by considering the efficiency of the computing resources in solving the inverse matrix of the corresponding square matrix when selecting the target splitting dimension, the embodiments of the present disclosure can further reduce the time overhead required for solving the inverse matrix.
[0050] In certain implementations of the second aspect, the splitting unit is further configured to: select a target dimension from multiple dimensions based on the information loss corresponding to the multiple dimensions, where the information loss corresponding to a dimension indicates the information loss caused by splitting the reference Kronecker factor matrix using this dimension, and the target dimension indicates the rank of at least one of the multiple square matrices obtained by splitting the Kronecker factor matrix.
[0051] In some embodiments, the splitting unit may, for example, use each dimension in a set of candidate dimensions to split the reference Kronecker factor matrix to determine the information loss corresponding to each dimension. To more accurately evaluate the information loss in the actual training process, the reference Kronecker factor matrix may be, for example, the Kronecker factor matrix obtained after a predetermined number of iterations of the neural network model using the same training data.
[0052] In some embodiments, the splitting unit may determine the corresponding information loss for each dimension, and select the dimension whose information loss meets a predetermined requirement (for example, the information loss is less than a specific threshold or the least information loss) as the target dimension. In this application, by considering the information loss when selecting the target splitting dimension, the embodiments of the present disclosure can ensure the calculation accuracy of the inverse matrix while reducing the time complexity of the inverse matrix calculation.
[0053] In certain implementations of the second aspect, the information loss corresponding to one dimension is the difference between the spectral norm of the joint matrix of a plurality of square matrices obtained by partitioning the reference Kronecker factor matrix according to the dimension and the spectral norm of the reference Kronecker factor matrix.
[0054] In the present application, quantifying the information loss by the spectral norm enables the data processing device to more efficiently determine the target partitioning dimension.
[0055] In certain implementations of the second aspect, the partitioning unit is further configured to: based on the performance information corresponding to a plurality of dimensions and the information loss corresponding to the plurality of dimensions, select a target dimension from the plurality of dimensions to partition the Kronecker factor matrix, where the performance information corresponding to one dimension indicates the efficiency of the computing resource in processing the square matrix corresponding to the dimension, the information loss corresponding to one dimension indicates the information loss caused by partitioning the reference Kronecker factor matrix using the dimension, and the target dimension indicates the rank of at least one of the plurality of square matrices obtained by partitioning the Kronecker factor matrix.
[0056] In some embodiments, the partitioning unit may determine the corresponding performance information and information loss for each dimension, and determine the target dimension based on the two. For example, the target dimension may be such that both the corresponding performance information and information loss satisfy predetermined conditions (e.g., the performance is better than a specific threshold and the information loss is less than a specific threshold).
[0057] In some embodiments, the target dimension may also be determined, for example, based on the normalization results of the performance information and the information loss. For example, by constructing a fitting function between different candidate dimensions and the performance information and / or information loss, the target dimension is determined by determining the intersection point of the fitting function. In the present application, by considering both the performance information and the information loss, the embodiments of the present disclosure can achieve a balance between the time cost of solving the inverse matrix and the calculation accuracy, and avoid too poor calculation accuracy or too low calculation efficiency.
[0058] In certain implementations of the second aspect, the adjustment module is configured to: first process a plurality of square matrices in parallel to determine a plurality of inverse matrices of the plurality of square matrices. Subsequently, based on the combination of the plurality of inverse matrices of the plurality of square matrices, adjust the parameters of the neural network model.
[0059] In the present application, since at least some of the square matrices obtained by partitioning will have the same rank, this provides good support for parallel processing. By processing these square matrices in parallel, the time cost of solving the inverse matrix can be further reduced.
[0060] In certain implementations of the second aspect, the neural network model is an image processing model. The obtaining module is configured to: obtain image training data; and then, apply the image training data to the image processing model to obtain a Kronecker factor matrix.
[0061] In this application, the object processed by the neural network model can be image data. Through the method of the present disclosure, the training of the image analysis model can be achieved more quickly, enabling the image analysis model to be put into use more quickly.
[0062] In certain implementations of the second aspect, the neural network model is a text processing model. The obtaining module is configured to: obtain text training data; and then, apply the text training data to the text processing model to obtain a Kronecker factor matrix.
[0063] In this application, the object processed by the neural network model can be text data. Through the method of the present disclosure, the training of the text processing model can be achieved more quickly, enabling the text processing model to be put into use more quickly.
[0064] In a third aspect of the present disclosure, there is provided an electronic device, including: at least one computing unit; at least one memory, the at least one memory being coupled to the at least one computing unit and storing instructions for execution by the at least one computing unit, the instructions when executed by the at least one computing unit causing the device to perform the method in the first aspect or any one of the implementations in the first aspect.
[0065] In a fourth aspect of the present disclosure, there is provided a computer-readable storage medium having one or more computer instructions stored thereon, wherein the one or more computer instructions are executed by a processor to implement the method in the first aspect or any one of the implementations in the first aspect.
[0066] In a fifth aspect of the present disclosure, there is provided a computer program product, when the computer program product runs on a computer, causing the computer to execute instructions for performing some or all of the steps of the method in the first aspect or any one of the implementations in the first aspect.
[0067] It can be understood that the electronic device described in the third aspect, the computer storage medium described in the fourth aspect, or the computer program product described in the fifth aspect provided above are all used to execute the method provided in the first aspect. Therefore, the explanations or descriptions regarding the first aspect also apply to the third aspect, the fourth aspect, and the fifth aspect. In addition, the beneficial effects that can be achieved by the third aspect, the fourth aspect, and the fifth aspect can refer to the beneficial effects in the corresponding methods, which will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] In conjunction with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages, and aspects of the embodiments of the present disclosure will become more apparent. In the drawings, the same or similar reference numerals denote the same or similar elements, where:
[0069] Figure 1 A schematic diagram showing the decomposition process of the high-order information matrix of a deep learning model;
[0070] Figure 2 A schematic diagram showing an example environment in which multiple embodiments of the present disclosure can be implemented;
[0071] Figure 3 A schematic diagram showing the structure of an example data processing module according to some embodiments of the present disclosure;
[0072] Figure 4 A schematic diagram showing the structure of an example Kronecker factor matrix slicing module according to some embodiments of the present disclosure;
[0073] Figure 5 A schematic diagram showing the structure of an example slicing dimension calculation module according to some embodiments of the present disclosure;
[0074] Figure 6 A schematic diagram showing an example fitting according to some embodiments of the present disclosure;
[0075] Figures 7A to 7C A schematic diagram of slicing a Kronecker factor matrix according to some embodiments of the present disclosure;
[0076] Figure 8 A schematic diagram showing an example data processing system according to some embodiments of the present disclosure;
[0077] Figure 9 A schematic diagram showing an example data processing system according to still some other embodiments of the present disclosure;
[0078] Figure 10 A flowchart showing the data processing process according to some embodiments of the present disclosure;
[0079] Figure 11 A schematic block diagram of a data processing device according to an embodiment of the present application; and
[0080] Figure 12 A block diagram showing a computing device capable of implementing multiple embodiments of the present disclosure. Detailed Description of Specific Embodiments
[0081] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0082] In the description of the embodiments of the present disclosure, the term "including" and its like should be understood as an open inclusion, that is, "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "an embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc. may refer to different or the same objects. There may also be other explicit and implicit definitions hereinafter.
[0083] As discussed above, with the development of computer technology, machine learning technology has been applied in many fields, such as image recognition, natural language processing, audio analysis, video analysis, etc. Along with the wide application of machine learning technology, people are also increasingly concerned about the training process of machine learning models. On the one hand, people hope that the trained model can be more robust. On the other hand, people also hope to reduce the time required for the model to converge. Therefore, the optimization of machine learning models has become an important part of machine learning.
[0084] Common optimization algorithms include first-order optimization algorithms and higher-order optimization algorithms. The gradient descent algorithm is a common first-order optimization algorithm, and its parameter update process can be expressed as formula (1):
[0085]
[0086] where θ is the parameter to be updated in the machine learning model, η is the learning rate, is the first-order gradient of the loss function with respect to the parameter. In addition, there are some optimization algorithms that introduce strategies such as momentum and adaptive learning rate decay on the basis of the gradient descent algorithm, such as Momentum, Nesterov, AdaGrad, RMSprop, Adadelta, and Adam, etc. These improved optimization algorithms can, for example, adaptively update the step size by using the historical information of the stochastic gradient, making them easier to tune the parameters. However, in dealing with most problems, they do not have much performance improvement compared with the stochastic gradient descent algorithm with precisely tuned parameters. In addition, considering that the loss function of the neural network is a highly non-convex function and the curvature of the surface is unbalanced, the performance of the first-order optimization algorithm is not particularly ideal.
[0087] Based on first-order optimization algorithms, some methods propose using high-order gradient information to guide parameter updates. Such optimization algorithms are also known as high-order optimization algorithms. High-order optimization algorithms correct the direction and step size of the first-order gradient by utilizing the high-order derivatives of the objective function, thereby accelerating their convergence rate. During the training process, high-order optimization algorithms can closely approximate the optimal value, and geometrically, the descent path is also more in line with the true optimal descent path.
[0088] Compared with first-order optimization algorithms, the parameter update of high-order optimization algorithms can be expressed as formula (2):
[0089]
[0090] where θ is the parameter to be updated, η is the learning rate, is the first-order gradient of the loss function with respect to the parameter, and IM is the high-order information matrix. Different high-order optimization algorithms can have different definitions for the high-order information matrix. For example, the high-order information matrix corresponding to the Newton optimization algorithm is the Hessian matrix, which is a square matrix composed of all second-order partial derivatives of a multivariable real-valued function. The high-order information matrix corresponding to the natural gradient method is the Fisher matrix, which is the covariance matrix of the derivative of the maximum likelihood function.
[0091] Although high-order optimization algorithms have a fast convergence rate, the time complexity of calculating the inverse of the high-order information matrix is O(n 3 ). Therefore, how to use an efficient high-order optimization algorithm to reduce the time complexity of the optimization algorithm while ensuring the training accuracy has become the current focus of attention.
[0092] Some existing algorithms reduce the time complexity of calculating the inverse of the high-order information matrix through approximation. For example, the KFAC (Kronecker-factored Approximate Curvature) algorithm decomposes the Fisher matrix into block matrices according to the network layers, and each block matrix is approximated as the Kronecker product of two much smaller matrices. Since the inverse of the Kronecker product of two matrices is equal to the Kronecker product of their inverses, and these two smaller matrices are easier to calculate and invert than the entire block matrix, the KFAC optimization algorithm greatly simplifies the calculation of the Fisher matrix through such decomposition and approximation.
[0093] Figure 1 Illustrates an example process 100 of the KAFC algorithm. As Figure 1As shown, the high-order information matrix 110 of the model is grouped according to the weight of a given layer and diagonally approximated. For example, taking a 3-4-2 neural network (i.e., including 1 input layer, 1 hidden layer and 1 output layer, the input layer includes 3 nodes, the hidden layer includes 4 nodes, and the input layer includes 2 nodes) as an example, its high-order information matrix is a 20*20 matrix. Figure 1 As shown, at 115 , the high-order information matrix 110 may be approximated as, for example, a 12*12 matrix 120 - 1 and an 8*8 matrix 120 - 2 .
[0094] Further, at 130-1 and 130-2, the matrix 120-1 and the matrix 120-2 are respectively decomposed by Kronecker into corresponding Kronecker factor matrices. For example, the 12*12 matrix 120-1 can be decomposed into a 3*3 Kronecker factor matrix 140-1 and a 4*4 Kronecker factor matrix 140-2, and the 8*8 matrix 120-4 can be decomposed into a 4*4 Kronecker factor matrix 140-1 and a 2*2 Kronecker factor matrix 140-4.
[0095] Such optimization algorithms have good effects on neural networks with simple structures. However, for larger models, such optimization algorithms still have the disadvantage of high computational complexity. Figure 1 When the KFAC algorithm shown in the figure optimizes the training of the resnet50 neural network, since the dimension of its fully connected layer is 2048*1000, after optimization, a 2048-dimensional matrix and an inverse matrix of a 1000-dimensional matrix still need to be calculated. Although the inverse time is greatly reduced, such a speed is still difficult to meet the requirements of current applications. Through experiments, it is found that for the resnet50 model, the time for a single iteration of the KAFC optimization algorithm on the entire imagenet set is about 33 times that of the first-order optimization algorithm.
[0096] The data processing scheme according to the embodiment of the present disclosure will be described below. For the convenience of description, the various aspects of the present disclosure are described below through different subsections. These subsections are not intended to be different embodiments of the present disclosure. On the contrary, the technical features described in different subsections can be combined without violating the spirit of the present disclosure.
[0097] Basic Principle and Example Environment
[0098] In order to at least solve the problem of high time complexity of high-order optimization algorithms, according to various embodiments of the present disclosure, a data processing solution is provided. In an embodiment of the present disclosure, after obtaining the Kronecker factor matrix indicating the high-order information matrix of the neural network model, the Kronecker factor matrix is partitioned to obtain a plurality of square matrices, so that each element on the main diagonal of the Kronecker factor matrix is included in a single square matrix of the plurality of square matrices and is located on the main diagonal of the single square matrix. Subsequently, the parameters of the neural network model are adjusted based on the inverse operations of the obtained plurality of square matrices. In this way, the embodiments of the present disclosure can further reduce the dimension of the matrix that needs to perform inverse calculations, thereby reducing the time complexity of the optimization algorithm and further accelerating the training process of the neural network model.
[0099] Figure 2 FIG. shows a schematic diagram of an example environment 200 in which multiple embodiments of the present disclosure can be implemented. As Figure 2 shown, the environment 200 includes a data processing device 230, and the data processing device 230 can receive training data 210 and a neural network model 220. The model 220 can learn certain knowledge and capabilities from existing data for processing new data. The model 220 can be designed to perform various tasks, such as image classification, object detection, speech recognition, machine translation, content filtering, and so on. Examples of the model 220 include but are not limited to various types of deep neural networks (DNNs) and recurrent data networks (RNNs), etc. In an embodiment of the present disclosure, the model 220 can also be referred to as a "machine learning model". Hereinafter, the terms "neural network", "learning model", "learning network", "model", and "network" can be used interchangeably. The data processing device 230 should have a certain computing power to meet the computing resource requirements for implementing the method of the present application. The present disclosure does not limit the specific form of the data processing device 230. For example, it can be a network device or a terminal device.
[0100] Figure 2 The model 220 is shown as a type of deep neural network in the figure. A deep neural network has a hierarchical architecture, and each processing layer (also called a network layer) has one or more processing units (also called processing nodes, neurons, or filters), which process the input based on corresponding parameters. In a deep neural network, the output after processing by the previous layer is the input of the next layer, where the first layer in the architecture receives the network input for processing, and the output of the last layer is provided as the network output. The parameters used by all processing units of the model 220 constitute the parameter set of the model 220. The specific values of such a parameter set need to be determined through the training process.
[0101] It should be understood that Figure 2The architecture of the illustrated model 220 and the number of processing layers and processing units therein are schematic and not restrictive. In different applications, according to requirements, the prediction model can be designed to have other appropriate architectures and / or an appropriate number of processing layers, and each processing layer can have an appropriate number of processing units.
[0102] In some embodiments, the training data 210 received by the data processing device 230 is data corresponding to the model 220. For example, when the model 220 is an image processing model, the training data 210 can be image training data. When the model 220 is a text processing model, the training data 210 can be text training data. In some other application scenarios, the training data can also be appropriate types of training data such as speech data, video data, medical data, business data, etc.
[0103] As Figure 2 shown, a data processing module 235 is implemented on the data processing device 230. The data processing module 235 can be implemented as one or more software engines, hardware components, or a combination thereof, etc., which are configured with logic for implementing the functions of the corresponding module. When executed, the data processing module 235 causes the data processing device 230 to perform the following operations: obtaining Kronecker factor matrices 245-1 and 245-2 (collectively referred to as Kronecker factor matrix 245 alone or together) indicating the high-order information matrix 240; partitioning the obtained Kronecker factor matrix 245 to obtain a plurality of square matrices 250; and adjusting the parameters of the model 220 based on performing an inverse operation on the plurality of square matrices 250. It should be understood that for ease of description, although the process of determining the square matrix 250 from the Kronecker factor matrix 245 is shown outside the data processing device 230 in Figure 2 it can still be implemented by the data processing device 230. Based on such a manner, embodiments of the present disclosure can further reduce the time complexity of the high-order optimization algorithm, thereby reducing the time cost of training the model.
[0104] Example Architecture of Data Processing Module
[0105] First, referring to Figure 3 , Figure 3 shows a block diagram of an example data processing module 235 according to some embodiments of the present disclosure. As Figure 3 shown, the data processing module 235 includes a plurality of modules for implementing an example data processing process according to some embodiments of the present disclosure. As Figure 3 shown, the data processing module 235 includes a Kronecker factor matrix obtaining module 310, a Kronecker factor matrix partitioning module 320, and a model parameter adjusting module 330.
[0106] In some embodiments, the training data 210 and the model 220 can be provided as inputs to the Kronecker factor matrix obtaining module 310. The Kronecker factor matrix obtaining module 310 can utilize the training data 210 to iteratively adjust the parameters of the model 220, and obtain the Kronecker factor matrix 245 for indicating the high-order information matrix 240 of the model 220 based on the intermediate results of the forward calculation and the backpropagation during the training process.
[0107] In the present disclosure, the high-order information matrix 240 refers to a matrix for correcting the first-order gradient of the model 220, and its examples include but are not limited to the Hessian matrix, the Fisher matrix, and the second moment matrix mentioned above. It should be understood that Figure 2 The shown high-order information matrix 240 is only for indicating that a pair of Kronecker factor matrices 245 is used to indicate a high-order information matrix, and it is not required to calculate the high-order information matrix 240 itself during the training process. Some existing optimization algorithms can directly calculate the Kronecker factor matrix 245 of the high-order information matrix 240.
[0108] In some embodiments, the obtained Kronecker factor matrix 245 can be the Kronecker factor matrix 245 generated during any iteration process. For example, for a model 220 that has not been trained previously, the obtained Kronecker factor matrix 245 can be the Kronecker factor matrix generated based on the initial parameters of the model 220. Alternatively, the obtained Kronecker factor matrix 245 can also be the Kronecker factor matrix generated after several iterations. In some other alternative embodiments, the model 220 may have been pre-trained previously. At this time, the obtained Kronecker factor matrix 245 can also be the Kronecker factor matrix generated by continuing to iterate based on the pre-trained parameters.
[0109] As discussed above, the model 220 may generate multiple pairs of Kronecker factor matrices 245 during the iteration process. For example, the resnet50 model includes 53 convolutional layers and 1 fully connected layer, and it will generate 54 high-order information matrices during the iteration process. Each high-order information matrix will be approximated as a pair of Kronecker factor matrices. That is, the resenet50 model will generate a total of 108 Kronecker factor matrices during the parameter iteration process. For the convenience of description, the matrix in front of the Kronecker product operator in a pair of Kronecker factor matrices is called the left factor matrix, and the matrix behind the Kronecker product operator is called the right factor matrix.
[0110] In some embodiments, the Kronecker factor matrix obtaining module 310 can obtain all the Kronecker factor matrices generated by the model 220 during the iteration process, and provide them as outputs to the Kronecker factor matrix splitting module 320.
[0111] Alternatively, the Kronecker factor matrix obtaining module 310 may also obtain only any one, any pair, or any other number of matrices from all the Kronecker factor matrices, and provide it as an output to the Kronecker factor matrix splitting module 320. For example, considering that the dimensions of some Kronecker factor matrices are relatively small and the time complexity of their inversion operations is low, the Kronecker factor matrix obtaining module 310 may only select the Kronecker factor matrices with dimensions exceeding a predetermined threshold as inputs to be provided to the Kronecker factor matrix splitting module 320.
[0112] As Figure 3 shown, the data processing module 235 further includes a Kronecker factor matrix splitting module 320. After receiving the Kronecker factor matrix 245 provided by the Kronecker factor matrix obtaining module 310, the Kronecker factor matrix splitting module 320 may determine a corresponding splitting strategy, and split the Kronecker factor matrix 245 based on this splitting strategy to obtain a plurality of square matrices 250, where the plurality of square matrices 250 correspond to the elements located on the diagonal of the Kronecker factor matrix 245. Based on such a split, the inverse matrix of the Kronecker factor matrix 245 can be approximated as a combination of the inverses of the obtained plurality of square matrices 250.
[0113] In some embodiments, the splitting strategy adopted by the Kronecker factor matrix splitting module 320 may be pre-stored. The Kronecker factor matrix splitting module 320 may, for example, obtain the pre-stored splitting strategy based on the model identifier of the model 220 and the resource identifier of the target device for optimizing the model 220. Alternatively, the splitting strategy may also be determined in real time by the Kronecker factor matrix obtaining module 310 based on the model 220 and the target device for optimizing the model 220. The process of splitting the Kronecker factor matrix will be described in detail below.
[0114] The plurality of square matrices 250 obtained by splitting the Kronecker factor matrix splitting module 320 are further provided to Figure 3 the parameter adjustment module 330 therein. The parameter adjustment module 330 may, for example, calculate the inverse matrices of the obtained plurality of square matrices 250 respectively, and use their combination (whose value is approximately equal to the inverse matrix of the Kronecker factor matrix 245) to determine the parameter adjustment of the model 220, thereby guiding the training of the model 220. The process of adjusting the model parameters will be described in detail below.
[0115] It should be understood that Figure 3The modules in can be implemented as one or more software engines, hardware components, or combinations thereof, etc., which are configured with logic for implementing the functions of the corresponding modules. The software engines, etc. are executed on one or more processors of one or more computing systems or devices, and utilize or operate on data stored in one or more storage devices or memories, etc. on one or more computing systems. In some embodiments, Figure 3 the different modules in can be implemented as a single module, and Figure 3 the single module in can be separated into multiple modules. In some embodiments, the data processing module 235 may further include one or more additional modules.
[0116] Partitioning of Kronecker Factor Matrix
[0117] As discussed above, the Kronecker factor matrix slicing module 320 is configured to slice the received Kronecker factor matrix 245 according to a specific slicing strategy to obtain a plurality of square matrices 250. In some embodiments, the Kronecker factor matrix slicing module 320 may determine whether a slicing strategy corresponding to the model 220 and the computing resources used to train the model 220 has been stored by querying a configuration file.
[0118] Figure 4 A block diagram of an exemplary Kronecker factor matrix slicing module 320 according to some embodiments of the present disclosure is shown. In some embodiments, the Kronecker factor matrix slicing module 320 may include a slicing strategy query module 430. The slicing strategy query model 430 may be configured to obtain a resource identifier 410 of the computing resources and a model identifier 420 of the model 220, where the computing resources may be, for example, an AI chip for training a machine learning model, including but not limited to a GPU (Graphics Processing Unit), an FPGA (Field Programmable Gate Array), and an ASIC (Application Specific Integrated Circuit), etc.
[0119] Further, the segmentation strategy query module 430 may query the configuration file 440 based on the resource identifier 410 and the model identifier 420 to determine whether there is a maintained segmentation strategy corresponding to the model 220 and the computing resources. In some embodiments, the configuration file 440 may include the correspondence between the resource identifier 410, the model identifier 420, and the corresponding segmentation strategy. This correspondence indicates the existence of an optimization strategy previously determined for the computing resources corresponding to the resource identifier 410 and the model corresponding to the model identifier 420. Such an optimization strategy may be determined by the same data processing device 230, or may be determined by different computing devices, or may also be manually configured by a developer based on mathematical calculations, experiments, or experience. It should be understood that the configuration file 440 is only an example form of maintaining such a mapping relationship, and any appropriate manner may be adopted to maintain such a mapping, and the present disclosure is not intended to limit this.
[0120] The model identifier may be identification information indicating the type of the neural network model. For example, "resnet50" may indicate that the neural network model is a 50-layer deep residual network (ResNet). Examples of the model identifier may also include "resnet18", "resnet101", "VGG16", "LeNet5", and so on. Additionally, the resource identifier may be identification information indicating the model number of the computing resources. For example, the chip name, chip code name, chip type number, or even the identifier of the device where the chip is located can all be used as the resource identifier. Computing resources of the same model often have the same or similar performance. As an example, the resource identifier may be the specific model number of the AI chip used to train the model 220. In this case, the segmentation strategy query module 430 will query whether there is an optimization strategy corresponding to the computing resources of the same model and the same type of model in the configuration file 440, without requiring it to be the same computing resource. For example, the segmentation strategy query module 430 may query whether there is an optimization strategy corresponding to the model number "GPUV100" and the model "resnet50".
[0121] Alternatively, the resource identifier may also be the unique identifier of the computing resource. For example, the MAC address of the computing resource, etc. In this case, the segmentation strategy query module 430 will query whether there is an optimization strategy for this computing resource and the model 220 in the configuration file 440. For example, the segmentation strategy query module 430 may query whether there is an optimization strategy corresponding to the GPU with the MAC address "abcd" and the model "resnet50".
[0122] As Figure 4As shown, when it is determined that there is a corresponding optimization strategy in the configuration file 440, the slicing strategy query module 430 may determine the slicing dimension 450 from the optimization strategy and send it to the matrix slicing module 480. In some embodiments, the optimization strategy may include, for example, a slicing dimension for the left factor matrix and a slicing dimension for the right factor matrix. The slicing strategy query module 430 may send at least one of the slicing dimension for the left factor matrix and the slicing dimension for the right factor matrix to the matrix slicing module 480.
[0123] As an example, the configuration file 440 may include, for example, a mapping between the resource identifier "chip ABC", the model identifier "resnet50", the left factor matrix slicing dimension "128", and the right factor matrix slicing dimension "1". When it is determined that the resource identifier of the computing resource is "chip ABC" and the model identifier is "resenet50", the slicing strategy query module 430 may obtain the corresponding slicing strategy by looking up the configuration file 440 as follows: the left factor matrix slicing dimension "128" and the right factor matrix slicing dimension "1". Subsequently, the slicing strategy query module 430 may send both the left factor matrix slicing dimension "128" and the right factor matrix slicing dimension "1" to the matrix slicing module 480, for example.
[0124] In some embodiments, when it is determined that there is no corresponding slicing strategy in the configuration file 440, the slicing strategy query module 430 may provide an indication 460 that there is no existing slicing strategy to the slicing dimension calculation module 470, so that the slicing dimension calculation module 470 calculates the slicing dimension 450 corresponding to the model 220 and the computing resource. The following will refer to Figure 5 to describe the slicing dimension calculation module 470, Figure 5 shows a block diagram of an example slicing dimension calculation module 470 according to some embodiments of the present disclosure.
[0125] As Figure 5 shown, in some embodiments, the slicing dimension calculation module 470 may include a reference Kronecker factor matrix acquisition module 510. To avoid excessive variation in the Kronecker factor matrix used to determine the slicing dimension, the reference Kronecker factor matrix acquisition module 510 may receive the training data 210 and the model 220 to perform a predetermined number of parameter iterations on the model 220. After the Kronecker factor matrix becomes relatively stable, the reference Kronecker factor matrix 520 is extracted as the input to the information loss confirmation module 530.
[0126] To more accurately estimate the information loss caused by the partitioning to the Kronecker factor matrix generated during the actual training process, the reference Kronecker factor matrix 520 obtained by the reference Kronecker factor matrix acquisition module 510 can be all the Kronecker factor matrices generated during the process of iterating the parameters of the model 220 using the training data 210. Taking the resnet50 model as an example, the reference Kronecker factor matrix acquisition module 510 can, for example, provide the 108 reference Kronecker factor matrices 520 generated after iterating the parameters of the model 220 20 times based on the training data 210 as input to the information loss confirmation module 530.
[0127] In some embodiments, the information loss confirmation module 530 can receive multiple reference Kronecker factor matrices 520 and multiple candidate partitioning dimensions 550, and calculate the information loss of each reference Kronecker factor matrix 520 under each candidate partitioning dimension 550. As an example, the candidate partitioning dimensions 550 can be, for example, [1, 16, 32, 64, 128, 256, 512, 1024, 2048], and the information loss confirmation module 530 can use each dimension in the list to perform diagonal partitioning on the reference Kronecker factor matrix 520 to obtain multiple candidate square matrices, and calculate the corresponding information loss 535.
[0128] For example, if the selected partitioning dimension is 16 and the reference Kronecker factor matrix 520 to be partitioned is a 64*64 matrix, then the information loss confirmation module 530 will partition the reference Kronecker factor matrix 520 into 4 16*16 candidate square matrices along the diagonal.
[0129] As another example, if the selected partitioning dimension is 16 and the reference Kronecker factor matrix 520 to be partitioned is a 70*70 matrix, then the information loss confirmation module 530 can partition the reference Kronecker factor matrix 520 into 4 16*16 candidate square matrices and 1 6*6 candidate square matrix along the diagonal. The partitioning process of the reference Kronecker factor matrix 520 can be similar to the process of partitioning the Kronecker factor matrix 245 described in detail below for the reference matrix partitioning module 480.
[0130] In some embodiments, the information loss 535 can indicate the information difference between the joint matrix of multiple candidate square matrices and the partitioned reference Kronecker factor matrix 520. In some embodiments, the information loss confirmation module 530 can select the spectral norm of the matrix as the information metric, and determine the difference between the spectral norm of the reference Kronecker factor matrix 520 and the spectral norm of the joint matrix of multiple candidate square matrices. It should be understood that other appropriate information metrics can also be used to determine the information loss 535.
[0131] In some embodiments, the segmentation dimension selection module 560 may determine the segmentation dimension 450 from the candidate segmentation dimensions 550 based on the received information loss 535. For example, the segmentation dimension selection module 560 may compare the information loss 535 corresponding to different candidate dimensions with a loss threshold. Such a loss threshold can be manually configured by the developer as needed. For example, taking resnet50 as an example, the segmentation dimension selection module 560 may compare the total loss of the information loss of the left factor matrix corresponding to each candidate segmentation dimension 550 with the loss threshold, and only retain the candidate segmentation dimensions 550 whose total loss is less than the loss threshold.
[0132] When the number of candidate segmentation dimensions retained based on the loss threshold is still greater than one, the segmentation dimension selection module 560 may, for example, select the smallest segmentation dimension among them to reduce the time complexity. Additionally, the segmentation dimension selection module 560 may also introduce other additional factors (such as the performance information introduced below) to select from the retained candidate segmentation dimensions.
[0133] In some embodiments, the segmentation dimension selection module 560 may also receive a specification regarding the tolerable information loss. For example, the segmentation dimension selection module 560 may receive an input regarding the tolerable information loss, the range of which may be set to 0% to 100%, where the larger the value, the greater the degree of information loss that can be tolerated. Further, the segmentation dimension selection module 560 may calculate the ratio of the information difference between the joint matrix of multiple candidate square matrices and the segmented reference Kronecker factor matrix 520 to the amount of information of the Ronnecker factor matrix 520, and compare this ratio with the tolerable information loss.
[0134] Subsequently, the segmentation dimension selection module 560 may calculate the proportion of the number of Kronecker factor matrices that satisfy the tolerable information loss in all the left factor matrices and all the right factor matrices, respectively. For example, Table 1 and Table 2 respectively give examples of the proportion of the left factor matrices and the right factor matrices that satisfy the loss when the tolerable information loss is 1%.
[0135] Table 1 Information Loss Statistics of Left Factor Matrices
[0136] Candidate Partitioning Dimension 1 16 32 64 128 256 512 1024 2048 Number of matrices with information loss < 1% 4 4 5 9 13 23 32 46 48
[0137] Table 2 Information Loss Statistics of Right Factor Matrices
[0138] Candidate Partitioning Dimension 1 16 32 64 128 256 512 1024 2048 Number of matrices with information loss < 1% 54 54 54 54 54 54 54 54 54
[0139] As can be seen from Table 1, the change in the size of the candidate segmentation dimension has a greater impact on the information loss of the left factor matrix. In contrast, based on the data in Table 2, it can be seen that all candidate segmentation dimensions enable all right factor matrices to meet the required 1% information loss.
[0140] In some embodiments, the segmentation dimension selection module 560 can, for example, compare the number of matrices that meet the tolerable information loss or its proportion in the total number with a predetermined threshold, and select the candidate segmentation dimension that can meet the threshold requirement. Taking the examples in Table 1 and Table 2, for example, the segmentation dimension selection module 560 can select the candidate segmentation dimension for which the proportion of matrices that meet the tolerable information loss exceeds 50%. In the example of Table 1, after the candidate segmentation dimension is greater than or equal to 256, the matrix proportion will exceed 50%. In the example of Table 2, all candidate segmentation dimensions meet the requirement that the proportion is greater than 50%.
[0141] The segmentation dimension selection module 560 can, for example, further select the candidate segmentation dimension with the smallest segmentation dimension from the candidate segmentation dimensions that meet the threshold requirement, so as to not only meet the requirement of information loss, but also reduce the computational complexity as much as possible. For the examples in Table 1 and Table 2, the segmentation dimension selection module 560 can, for example, select "256" as the segmentation dimension of the left factor matrix and select "1" as the segmentation dimension of the right factor matrix.
[0142] In some embodiments, as Figure 5 shown, the segmentation dimension calculation module 470 further includes a performance information calculation module 540, which is used to calculate the performance information 545 corresponding to the candidate segmentation dimension 550. The performance information 545 can, for example, indicate the efficiency of the computing resource for optimizing the model 220 to process the square matrix corresponding to the candidate segmentation dimension 550. In some embodiments, the performance information calculation module 540 can set a set of predetermined candidate segmentation dimensions 550, and construct a set of sample square matrices corresponding to the candidate segmentation dimensions 550. Each square matrix in the sample square matrix has a rank corresponding to the segmentation dimension. For example, when the candidate segmentation dimension 550 is "5", the corresponding sample square matrix can be a 5*5 square matrix. The performance information calculation module 540 can obtain the performance information of the computing resource for processing the sample square matrix as the performance information corresponding to the candidate dimension 550.
[0143] Taking the candidate segmentation dimensions as [1, 16, 32, 64, 128, 256, 512, 1024, 2048] as an example, the performance information calculation module 540 can, for example, cause the computing resources to calculate the inverse matrices of the sample square matrices with corresponding ranks respectively, and obtain their processing durations as the corresponding performance information 545. It should be understood that other appropriate processing procedures (such as matrix multiplication) can also be selected for evaluating the performance information of the computing resources for processing the corresponding sample square matrices. For example, Table 3 shows the example performance information of the computing resources calculating the inverse matrices of the sample square matrices corresponding to the segmentation dimensions.
[0144] Table 3 Performance Information of Computing Resources
[0145] Candidate Partitioning Dimension 1 16 32 64 128 256 512 1024 2048 Performance Information (us) 35 59 83 121 212 534 1418 3738 9824
[0146] As can be seen in Table 3, the larger the segmentation dimension, the longer the time taken for the computing resources to process the corresponding square matrix. In some embodiments, the segmentation dimension selection module 560 can determine the segmentation dimension 450 based on the received performance information 545. For example, the segmentation dimension selection module 560 can compare the performance information 545 with a predetermined performance threshold, and retain the candidate segmentation dimensions with processing durations less than the predetermined performance threshold. For example, the performance threshold can be specified as 300 us, and thus the segmentation dimension selection module 560 can determine that the candidate segmentation dimensions 1 to 128 all meet the requirements of the performance threshold. In some examples, in order to minimize the information loss caused by segmentation as much as possible, the segmentation dimension selection module 560 can, for example, select the largest value among the candidate segmentation dimensions that meet the threshold requirements as the segmentation dimension 450.
[0147] In some embodiments, the segmentation dimension selection module 560 can also select the segmentation dimension 450 from the candidate segmentation dimensions that meet the threshold requirements based on other factors (such as the information loss discussed above). For example, the segmentation dimension selection module 560 can select the dimension whose information loss also meets the corresponding requirements from the candidate segmentation dimensions that meet the threshold requirements as the segmentation dimension 450.
[0148] In some embodiments, in order to more accurately quantify the information loss and the performance information 545, the segmentation dimension selection module 560 can also normalize the information loss 535 and the performance information 545, and select the segmentation dimension 450 from the candidate segmentation dimensions 550 based on the normalized results. For example, for the examples in Table 1, Table 2, and Table 3, the segmentation dimension selection module 560 can normalize by converting the number of information losses that meet the tolerable information loss in Table 1 and Table 2 into a proportion, and can normalize by dividing each performance information in Table 3 by the performance information corresponding to the smallest candidate segmentation dimension. The normalized results of Table 1 - Table 3 can be represented as Table 4, for example.
[0149] Table 4 Normalization of Performance Information and Information Loss
[0150] Candidate Partitioning Dimension Table 1 Normalized Data Table 2 Normalized Data Table 3 Normalized Data 1 0.074074 1 1 16 0.074074 1 0.601374 32 0.092593 1 0.427578 64 0.166667 1 0.293913 128 0.240741 1 0.166927 256 0.425926 1 0.06641 512 0.592593 1 0.025 1024 0.851852 1 0.009486 2048 0.888889 1 0.00361
[0151] In some embodiments, the segmentation dimension selection module 560 may fit the normalized performance information and information loss, and the fitting function may be determined, for example, as a combination of piecewise linear functions. Taking the fitting of the information loss of the left factor matrix as an example, its fitting function can be represented, for example, as a linear function connecting the points (1, 0.074074) and (16, 0.074074), a linear function connecting the points (16, 0.074074) and (32, 0.092593), a linear function connecting the points (32, 0.092593) and (64, 0.166667), a linear function connecting the points (64, 0.166667) and (128, 0.240741), a linear function connecting the points (128, 0.240741) and (256, 0.425926), a linear function connecting the points (256, 0.425926) and (512, 0.592593), a linear function connecting the points (512, 0.592593) and (1024, 0.851852), and a combination of a linear function connecting the points (1024, 0.851852) and (2048, 0.888889). The fitting functions for the information loss and performance information of the right factor matrix can be determined similarly.
[0152] Figure 6 Fig. 600 shows a schematic diagram of an example fitting according to some embodiments disclosed in the version. As Figure 6 shown, the fitting function 610 represents the fitting function for the information loss of the left factor matrix, the fitting function 620 represents the fitting function for the information loss of the right factor matrix, and the fitting function 630 represents the fitting function for the performance information.
[0153] In some embodiments, the segmentation dimension selection module 560 may further determine the segmentation dimension based on the intersection points of different fitting functions. For example, the segmentation dimension selection module 560 may determine that the intersection point of the fitting function 630 of the performance information and the fitting function 620 of the information loss for the right factor matrix is the point (1, 1), so the segmentation selection module 560 may determine the segmentation dimension 450 for the right factor matrix as "1".
[0154] For Figure 6For example, the segmentation dimension selection module 560 can also determine that the intersection point of the fitting function 630 of the performance information and the fitting function 610 of the information loss for the right factor matrix is (104.5, 0.213547). Since 104.5 is not a candidate segmentation dimension 550, the segmentation dimension selection module 560 can select the candidate segmentation dimension closest to the value of the abscissa of this intersection point as the segmentation dimension 450. For example, the segmentation dimension selection module 560 can select "128" as the segmentation dimension 450 for the left factor matrix. In this way, the segmentation dimension selection module 560 can achieve a balance between low information loss and high computational efficiency.
[0155] Continue to refer to Figure 4 , after determining the segmentation dimension 450, the matrix segmentation module 480 can segment the Kronecker factor matrix 245 according to the segmentation dimension 450 to obtain a plurality of square matrices 250. As Figure 2 shown for the Kronecker factor matrix 245 and the square matrix 250, the matrix segmentation module 480 segments along the main diagonal of the Kronecker factor matrix 245, which makes the plurality of square matrices 250 correspond to the elements located on the diagonal of the Kronecker factor matrix 245.
[0156] The following will be described in conjunction with Figures 7A to 7C to describe the process of segmenting the Kronecker factor matrix. Figure 7A FIG. 700A shows a schematic diagram of segmenting a Kronecker factor matrix according to an embodiment of the present disclosure. As Figure 7AAs shown, the Kronecker factor matrix 245 is, for example, a 12×12 square matrix 710. When the partitioning dimension 450 is determined to be "4", the matrix partitioning module 480 can construct a square matrix 720-1 of rank "4" starting from the starting element X0Y0 on the main diagonal of the Kronecker factor matrix 710 according to the partitioning dimension. The elements on the main diagonal of the square matrix 720-1 are a subset of the elements on the main diagonal of the Kronecker factor matrix 710. Subsequently, the matrix partitioning module 480 can use the first element on the main diagonal of the Kronecker factor matrix 710 after the square matrix 720-1 (in this example, X4Y4) as the starting element, and construct a square matrix 720-2 of rank "4" along the main diagonal of the Kronecker factor matrix 710. Based on a similar method, the matrix partitioning module 480 can partition the Kronecker factor matrix 710 into three square matrices 720-1, 720-2, and 720-3 of the same size and rank "4", where the square matrices 720-1, 720-2, and 720-3 are sub-matrices of the Kronecker factor matrix 710. Based on such a partitioning method, the main diagonals of multiple square matrices (square matrices 720-1, 720-2, and 720-3) correspond one by one to a part of the main diagonal of the Kronecker factor matrix 710, and all the elements on the main diagonals of the multiple square matrices (square matrices 720-1, 720-2, and 720-3) include all the elements on the main diagonal of the Kronecker factor matrix 720.
[0157] In some other embodiments, the Kronecker factor matrix may not be completely partitioned into an integer number of square matrices with the same rank. Figure 7B FIG. 700B shows a schematic diagram of partitioning a Kronecker factor matrix according to another embodiment of the present disclosure. As Figure 7B shown, the Kronecker factor matrix 245 is, for example, a 12×12 square matrix 710. When the partitioning dimension 450 is determined to be "5", based on the method of continuous partitioning along the main diagonal, the matrix partitioning module 480 will first partition two square matrices 740-1 and 740-2 of rank "5" from the Kronecker factor matrix 730. When it is determined that the rank of the square matrix constructed with the remaining elements on the main diagonal as the main diagonal will be less than the partitioning dimension, the matrix partitioning module 480 can use the square matrix constructed with the remaining elements on the main diagonal as the main diagonal as one of the multiple square matrices partitioned from the Kronecker factor matrix. For example, for Figure 7B the example, the square matrix 740-3 of rank "2" will be determined as one of the square matrices partitioned from the Kronecker factor matrix 730. Based on such a method, the Kronecker factor matrix 730 will be partitioned into a square matrix 740-1 and 740-2 of rank "5", and a square matrix 740-3 of rank "2".
[0158] In still other embodiments, to facilitate batch processing of the split square matrices 250, the matrix splitting module 480 may pad the square matrices obtained by splitting with a rank less than the splitting dimension 450 to square matrices with a rank equal to the splitting dimension 450. Figure 7C FIG. 700C shows a schematic diagram of splitting a Kronecker factor matrix according to another embodiment of the present disclosure. As Figure 7C shown, the Kronecker factor matrix to be split is matrix 750, and the splitting dimension 450 is "5". Based on the method of continuous splitting along the main diagonal, the matrix splitting module 480 will first split out two square matrices 770-1 and 770-2 with a rank of "5" from the Kronecker factor matrix 750.
[0159] Compared with Figure 7B In the example of 7C, before sending the square matrix to different model parameter adjustment modules 330, the matrix splitting module 480 may expand the remaining square matrix to a square matrix with a rank of "5" by filling in specified values (e.g., 1). In some embodiments, the matrix splitting module 480 may determine before splitting whether the rank of the Kronecker factor matrix 240 to be split is divisible by the splitting dimension 450. If it is determined that the rank of the Kronecker factor matrix 240 is not divisible by the splitting dimension 450, the matrix splitting module 480 may expand the Kronecker factor matrix 240 by filling in specified values.
[0160] Taking Figure 7C as an example, the matrix splitting module 480 may, for example, determine that the rank "12" of the Kronecker factor matrix 750 is not divisible by the splitting dimension "5". Therefore, the matrix splitting module 480 may, for example, expand the Kronecker factor matrix 750 to an intermediate matrix 760, which has a rank "15" that is divisible by the splitting dimension "5". After obtaining the intermediate matrix 760, the matrix splitting module 480 may, for example, refer to the method described in Figure 7A to split the intermediate matrix 760 into three square matrices 770-1, 770-2, and 770-3 with a rank of "5".
[0161] In yet another embodiment, the matrix splitting module 480 may also refer to the method described in Figure 7B to obtain a square matrix with a rank less than the splitting dimension, and then expand the square matrix to a square matrix with a rank equal to the splitting dimension by filling in specified values. Combining Figure 7B and Figure 7CFor example, the matrix splitting module 480 can first split to obtain the square matrix 740-3, and then expand the square matrix 740-3 into a square matrix 770-3 with a rank of "5" by filling in the specified value "1". Since the filled value "1" will be converted to "0" during the inverse operation, such filling will not introduce additional errors to the inverse operation of the Kronecker factor matrix. On the contrary, by making the multiple square matrices obtained by splitting have the same rank, the embodiments of the present disclosure can better support the batch processing process for multiple square matrices.
[0162] Based on the matrix splitting method discussed above, the matrix splitting module 480 can make the main diagonals of multiple square matrices correspond one by one to a part of the main diagonal of the Kronecker factor matrix, and all the elements on the main diagonals of the multiple square matrices include all the elements on the main diagonal of the Kronecker factor matrix. In this way, the matrix splitting module 480 enables the inverse matrix operation of the Kronecker factor matrix 245 to be approximated as the inverse matrix operations of multiple square matrices 250, thereby reducing the time cost of the inverse matrix operation.
[0163] In some embodiments, the matrix splitting module 480 can also continuously store the square matrices 250 with the same dimension obtained by splitting in the memory for subsequent inverse matrix operation processes. Based on this way, the efficiency of batch processing can be improved.
[0164] It should be understood that the specific values involved in the above discussion are only illustrative and are not intended to limit the present disclosure.
[0165] Adjustment of Model Parameters
[0166] Continuing to refer to Figure 3 , the data processing module 235 further includes a model parameter adjustment module 330, which is configured to receive the multiple square matrices 250 output by the Kronecker factor matrix splitting module 320 and adjust the parameters of the model 220 according to the square matrices 250.
[0167] Specifically, the model parameter adjustment module 330 can calculate the inverse matrices of the multiple square matrices 250 and use the combination of the inverse matrices of these square matrices as an approximation of the inverse matrix of the Kronecker factor matrix 245 obtained by decomposing these square matrices 250. Subsequently, the model parameter adjustment module 330 can guide the optimization of the model 220 according to any appropriate process of adjusting parameters based on high-order optimization information, which will not be elaborated here. In this way, the embodiments of the present disclosure can convert the original large matrix inverse operation into the inverse operations of multiple small matrices, thereby greatly reducing the time complexity.
[0168] In some embodiments, the model parameter adjustment module 330 may also determine the inverse matrices of multiple square matrices 250 in parallel in a batch processing manner. Through the segmentation method discussed above, most of the segmented square matrices 250 have the same dimension, which provides a basis for parallel computing of the inverse matrices of these square matrices. By batch processing the square matrices with the same dimension among the obtained multiple square matrices 250, the embodiments of the present disclosure can further reduce the time overhead of computing the inverse matrix, thereby reducing the time cost required for model convergence.
[0169] It is found through experiments that the method of the present disclosure has obvious performance improvement compared with traditional high-order optimization algorithms (e.g., KFAC). For example, the time for computing the inverse matrix of the Kronecker factor matrix of the high-order information matrix is reduced from 445 ms of the KFAC algorithm to 28 ms of the method according to the present disclosure, and the performance is improved by 16 times.
[0170] Example Data Processing System
[0171] Figure 8 An example data processing system 800 according to an embodiment of the present disclosure is shown. The example data processing system 800 may be implemented as one or more software engines, hardware components or a combination thereof, etc., which are configured with logic for implementing the functions of the corresponding modules.
[0172] As Figure 8 shown, the data processing system 800 may include an offline module 810 and an online module 830. As discussed above with reference to Figures 4 to 5 For example, when it is determined that there is no segmentation dimension corresponding to the model identifier and the resource identifier, the data processing system 800 may activate the logic of the offline module 810. Specifically, the offline module 810 may include a judgment module 818, which is configured to receive a reference Kronecker factor matrix 812, a tolerable information loss 814, and a list of candidate segmentation dimensions 816, and determine a segmentation dimension 820 for segmenting the Kronecker factor matrix 832. In some embodiments, the tolerable information loss 814 is, for example, configurable by the user. In some embodiments, the reference Kronecker factor matrix 812 may refer to Figure 5 the reference Kronecker factor matrix 520 in
[0173] In some embodiments, the judgment model 818 may, for example, implement the same as Figure 5The same logic as that of the segmentation dimension selection module 560 in []. The evaluation model 818 may refer to the method discussed above to determine the information loss and performance information corresponding to each candidate segmentation dimension based on the reference Kronecker factor matrix 812, the tolerable information loss 814, and the candidate segmentation dimension list 816. The evaluation model 818 may, for example, further determine the segmentation dimension 820 based on both the information loss and the performance information.
[0174] After the determination of the segmentation dimension 820 is completed, the matrix segmentation module 834 included in the online module 820 may segment the Kronecker factor matrix 832 according to the segmentation dimension 820 to obtain a plurality of square matrices 836. In some embodiments, the matrix segmentation module 834 may implement the same logic as that of Figure 4 the matrix segmentation module 480 in []. By segmenting the Kronecker factor matrix 832 into a plurality of square matrices 836, the data processing system 800 can reduce the complexity of the inverse matrix operation required for adjusting the parameters of the model.
[0175] In some embodiments, the plurality of square matrices 836 obtained by segmentation may be provided to the memory integration module 838 to store the square matrices 836 with the same rank in a continuous manner in the memory 840. Based on this way, the efficiency of subsequent batch calculation of the inverse matrix of the square matrix 836 can be improved. In some embodiments, the online module 830 in the data processing system 800 may also run independently. For example, if it is determined that there is a segmentation dimension corresponding to the model identifier and the resource identifier, the matrix segmentation module 834 may not receive the segmentation dimension 820 from the offline module. Instead, the matrix segmentation module 834 may segment the Kronecker factor matrix 832 based on the segmentation dimension determined based on the model identifier and the resource identifier. Based on this way, the data processing system 800 can effectively utilize the historical optimization strategy, thereby avoiding an additional segmentation dimension calculation process.
[0176] Figure 9 FIG. shows an example data processing system 900 according to still some embodiments of the present disclosure. The example data processing system 900 may be implemented as one or more software engines, hardware components, or a combination thereof, etc., which are configured with logic for implementing the functions of the corresponding modules.
[0177] As Figure 9 shown, the data processing system 900 includes an offline module 910 and an online module 938. As referred to above with reference to Figures 4 to 5As discussed, for example, when it is determined that there is no segmentation dimension corresponding to the model identifier and the resource identifier, the data processing system 900 may activate the logic of the offline module 910. Specifically, the offline module 910 may obtain the tolerable information loss 912, the reference Kronecker factor matrix 914, and the candidate segmentation dimension list 916. In some embodiments, the tolerable information loss 912 is, for example, configurable by the user. In some embodiments, the reference Kronecker factor matrix 914 may refer to Figure 5 the reference Kronecker factor matrix 520 in
[0178] In some embodiments, as Figure 9 shown, the offline module 910 may include a matrix segmentation module 918, which is configured to segment the reference Kronecker factor matrix 914 according to the segmentation dimensions in the candidate segmentation dimension list 916. Further, the matrix segmentation module 918 may also determine the matrix information metric 920 of the combined matrix of the plurality of square matrices obtained by segmentation. In some embodiments, the matrix information metric 920 may be, for example, the spectral norm of the matrix.
[0179] In some embodiments, the offline module 910 may also determine the performance data 922 corresponding to the candidate segmentation dimensions based on the candidate segmentation dimension list 916. The offline module 910 may, for example, implement the same logic as the Figure 5 performance information calculation module 540 in
[0180] to determine the corresponding performance data. Additionally, the offline module 910 may include an evaluation module 924. The evaluation module 924 may be configured to determine the spectral norm ratio 926 for the left factor matrix and the spectral norm ratio 928 for the right factor matrix according to the tolerable information loss 912 and the matrix information metric 920. The evaluation module 924 may refer to the processes described above with respect to Tables 1 and 2 to determine the spectral norm ratio 926 and the spectral norm ratio 928. The spectral norm ratio 926 may represent the number of left factor matrices that use the corresponding segmentation dimension to segment the left factor matrix such that the information loss satisfies the tolerable information loss 912. Similarly, the spectral norm ratio 928 may represent the number of right factor matrices that use the corresponding segmentation dimension to segment the right factor matrix such that the information loss satisfies the tolerable information loss 912.
[0181] Further, as Figure 9 shown, after determining the spectral norm ratio 926 and the spectral norm ratio 928, the evaluation module 924 may also refer to the performance data 922 to determine the left factor matrix segmentation dimension 930 and the right factor matrix segmentation dimension 932. It should be understood that the evaluation module 924 may refer to the above with respect toFigure 6 The described process is used to determine the partitioning dimension 930 of the left factor matrix and the partitioning dimension 932 for the right factor matrix.
[0182] In some embodiments, the data processing system 900 further includes an online module 938. As Figure 9 shown, the matrix partitioning module 940 in the online module 938 can receive the left factor matrix partitioning dimension 930 and the right factor matrix partitioning dimension 932 determined by the evaluation module 924. Additionally, the matrix partitioning module 940 can also receive the input value gradient covariance matrix 934 (corresponding to the left factor matrix) and the feature map covariance matrix 936 (corresponding to the right factor matrix) generated during the training model process. The matrix partitioning module 940 can partition the input value gradient covariance matrix 934 and the feature map covariance matrix 936 according to the left factor matrix partitioning dimension 930 and the right factor matrix partitioning dimension 932 respectively. It should be understood that the matrix partitioning module 940 can partition the input value gradient covariance matrix 934 and the feature map covariance matrix 936 according to the partitioning process described with reference to Figures 7A to 7C the described partitioning process.
[0183] Subsequently, the multiple square matrices obtained by partitioning through the matrix partitioning module 940 can be provided to the memory integration module 942. The memory integration module 942 can store the square matrices with the same rank among the multiple square matrices in a continuous manner, thereby forming a batch matrix 944. Based on this way, the efficiency of subsequent batch calculation of the inverse matrix of the square matrix can be improved.
[0184] Based on this way, the data processing system 900 can use the offline module 910 to determine the corresponding partitioning dimensions for the left factor matrix and the right factor matrix, and perform corresponding partitioning in the online module 938, thereby reducing the complexity of the inverse matrix operation required to adjust the model parameters. In addition, the offline module 910 comprehensively considers performance factors and information loss factors when determining the left factor matrix partitioning dimension 930 and the right factor matrix partitioning dimension 932, which can enable the matrix partitioning performed by the online module 938 to meet both the predetermined performance requirements and the requirements of information loss.
[0185] Example Process and Example Device
[0186] Figure 10 shows a flowchart of an example data processing process 1000 according to an embodiment of the present disclosure. The process 1000 can be implemented, for example, by Figure 2 the data processing device 230 in. For ease of description, the process 1000 is described below with reference to Figure 2 this.
[0187] At block 1002, the data processing device 230 obtains a Kronecker factor matrix 245 of a high-order information matrix 240 for indicating a neural network model 220, where the high-order information matrix 240 is used to correct the first-order gradient of the neural network model 220. At block 1004, the data processing device 230 partitions the Kronecker factor matrix 245 to obtain a plurality of square matrices 250, where the plurality of square matrices 250 are sub-matrices of the Kronecker factor matrix, the main diagonals of the plurality of square matrices 250 respectively correspond to a part of the main diagonal of the Kronecker factor matrix 245, and all the elements on the main diagonals of the plurality of square matrices 250 include all the elements on the main diagonal of the Kronecker factor matrix 250. At block 1006, the data processing device 230 adjusts the parameters of the neural network model based on the plurality of square matrices 250.
[0188] Figure 10 Further shown is a block diagram of a data processing apparatus 1100 according to an embodiment of the present disclosure. The data processing apparatus 1100 may include a plurality of modules for performing corresponding steps in the process 1000 as Figure 10 discussed. As Figure 11 shown, the data processing apparatus 1100 includes an obtaining unit 1110 configured to obtain a Kronecker factor matrix for indicating a high-order information matrix of a neural network model, where the high-order information matrix is used to correct the first-order gradient of the neural network model. The data processing apparatus 1100 further includes a partitioning unit 1120 configured to partition the Kronecker factor matrix to obtain a plurality of square matrices such that the plurality of square matrices are sub-matrices of the Kronecker factor matrix, and the main diagonals of the plurality of square matrices respectively correspond to a part of the main diagonal of the Kronecker factor matrix. In addition, the data processing apparatus 1100 further includes an adjusting unit 1130 configured to adjust the parameters of the neural network model based on the plurality of square matrices
[0189] In some embodiments, the partitioning unit 1120 is further configured to: partition the Kronecker factor matrix based on a resource identifier of computing resources and a model identifier of the neural network model.
[0190] In some embodiments, a correspondence relationship between a resource identifier, a model identifier, and a dimension is stored in the data processing device, and the dimension indicates the rank of at least one of the plurality of square matrices obtained by partitioning the Kronecker factor matrix.
[0191] In some embodiments, the partitioning unit 1120 is further configured to: select a target dimension from a plurality of dimensions to partition the Kronecker factor matrix based on the performance information corresponding to the plurality of dimensions, where the performance information corresponding to one dimension indicates the efficiency of the computing resources in processing the square matrix corresponding to the dimension, and the target dimension indicates the rank of at least one of the plurality of square matrices obtained by partitioning the Kronecker factor matrix.
[0192] In some embodiments, the value of the performance information corresponding to a dimension is related to the time required for the data processing device to calculate the inverse matrix of the square matrix corresponding to the dimension.
[0193] In some embodiments, the splitting unit 1120 is further configured to: based on the information loss corresponding to multiple dimensions, select a target dimension from the multiple dimensions to split the Kronecker factor matrix, where the information loss corresponding to a dimension indicates the information loss caused by using the dimension to split the reference Kronecker factor matrix, and the target dimension indicates the rank of at least one of the multiple square matrices obtained by splitting the Kronecker factor matrix.
[0194] In some embodiments, the information loss corresponding to a dimension is the difference between the spectral norm of the joint matrix of the multiple square matrices obtained by splitting the reference Kronecker factor matrix using the dimension and the spectral norm of the reference Kronecker factor matrix.
[0195] In some embodiments, the splitting unit is further configured to: based on the performance information corresponding to multiple dimensions and the information loss corresponding to multiple dimensions, select a target dimension from the multiple dimensions to split the Kronecker factor matrix, where the performance information corresponding to a dimension indicates the efficiency of the computing resource in processing the square matrix corresponding to the dimension, the information loss corresponding to a dimension indicates the information loss caused by using the dimension to split the reference Kronecker factor matrix, and the target dimension indicates the rank of at least one of the multiple square matrices obtained by splitting the Kronecker factor matrix.
[0196] In some embodiments, the adjustment module 1130 is configured to: first, process multiple square matrices in parallel to determine the inverse matrices of the multiple square matrices. Subsequently, based on the combination of the inverse matrices of the multiple square matrices, adjust the parameters of the neural network model.
[0197] In some embodiments, the neural network model is an image processing model. The acquisition module 1110 is configured to: acquire image training data; subsequently, apply the image training data to the image processing model to acquire the Kronecker factor matrix.
[0198] In some embodiments, the neural network model is a text processing model. The acquisition module 1110 is configured to: acquire text training data; subsequently, apply the text training data to the text processing model to acquire the Kronecker factor matrix.
[0199] Example Device
[0200] Figure 12FIG. shows a schematic block diagram of an example device 1200 that can be used to implement embodiments of the present disclosure. Device 1200 can be used to implement data processing device 230. As shown, device 1200 includes a computing unit 1201 that can perform various appropriate actions and processes according to computer program instructions stored in random access memory (RAM) and / or read-only memory (ROM) 1202 or computer program instructions loaded from storage unit 1207 into RAM and / or ROM 1202. In RAM and / or ROM 1202, various programs and data required for the operation of device 1200 can also be stored. The computing unit 1201 and RAM and / or ROM 1202 are connected to each other via a bus 1203. Input / output (I / O) interface 1204 is also connected to bus 1203.
[0201] A plurality of components in device 1200 are connected to I / O interface 1204, including: an input unit 1205, such as a keyboard, mouse, etc.; an output unit 1206, such as various types of displays, speakers, etc.; a storage unit 1207, such as a magnetic disk, optical disc, etc.; and a communication unit 1208, such as a network card, modem, wireless communication transceiver, etc. Communication unit 1208 allows device 1200 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0202] The computing unit 1201 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1201 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1201 executes the various methods and processes described above, such as process 1000. For example, in some embodiments, process 1000 can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as storage unit 1207. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 1200 via RAM and / or ROM and / or communication unit 1208. When the computer program is loaded into RAM and / or ROM and executed by the computing unit 1201, one or more steps of process 1000 described above can be executed. Alternatively, in other embodiments, the computing unit 1201 can be configured to execute process 1000 in any other suitable manner (e.g., by means of firmware).
[0203] The program code for implementing the methods of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general purpose computer, a special purpose computer, or other programmable data processing device, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0204] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0205] Moreover, although the operations are depicted in a particular order, this should be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed to achieve the desired result. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments may also be implemented in combination in a single implementation. Conversely, various features described in the context of a single implementation may also be implemented separately or in any suitable sub-combination in multiple implementations.
[0206] Although the subject matter has been described in language specific to structural features and / or methodological acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. A method for data processing, the method being used for a data processing device, the method comprising: Obtaining a Kronecker factor matrix for indicating a high-order information matrix of a neural network model, the high-order information matrix being used for correcting a first-order gradient of the neural network model, wherein the neural network model is a text processing model or an image processing model; Partitioning the Kronecker factor matrix to obtain a plurality of square matrices, the plurality of square matrices being sub-matrices of the Kronecker factor matrix, and main diagonals of the plurality of square matrices corresponding to a part of the main diagonal of the Kronecker factor matrix one by one; and Adjusting parameters of the neural network model based on the plurality of square matrices, wherein the data processing device includes computing resources for performing the adjustment, and wherein partitioning the Kronecker factor matrix includes: based on information corresponding to a plurality of dimensions, selecting a target dimension from the plurality of dimensions to partition the Kronecker factor matrix, the information corresponding to one dimension being performance information or information loss, wherein the performance information indicates the efficiency of the computing resources in processing the square matrix corresponding to the dimension, and the information loss indicates the information loss caused by partitioning a reference Kronecker factor matrix using the dimension.
2. The method according to claim 1, wherein partitioning the Kronecker factor matrix includes: Partitioning the Kronecker factor matrix based on a resource identifier of the computing resources and a model identifier of the neural network model.
3. The method according to claim 2, wherein a correspondence relationship between the resource identifier, the model identifier, and a dimension is stored in the data processing device, the dimension indicating the rank of at least one square matrix obtained by partitioning the Kronecker factor matrix.
4. The method according to claim 1, wherein the target dimension indicates the rank of at least one square matrix obtained by partitioning the Kronecker factor matrix.
5. The method according to claim 1, wherein the value of the performance information corresponding to the dimension is related to the time required for the data processing device to calculate the inverse matrix of the square matrix corresponding to the dimension.
6. The method according to claim 1, wherein the information loss corresponding to the dimension is the difference between the spectral norm of the joint matrix of the plurality of square matrices obtained by partitioning the reference Kronecker factor matrix using the dimension and the spectral norm of the reference Kronecker factor matrix.
7. The method according to claim 1, wherein partitioning the Kronecker factor matrix includes: Based on the performance information corresponding to a plurality of dimensions and the information loss corresponding to the plurality of dimensions, selecting a target dimension from the plurality of dimensions to partition the Kronecker factor matrix, wherein the performance information corresponding to one dimension indicates the efficiency of the computing resources in processing the square matrix corresponding to the dimension, the information loss corresponding to one dimension indicates the information loss caused by partitioning a reference Kronecker factor matrix using the dimension, and the target dimension indicates the rank of at least one square matrix obtained by partitioning the Kronecker factor matrix.
8. The method according to any one of claims 1-7, wherein adjusting the parameters of the neural network model comprises: processing the plurality of square matrices in parallel to determine a plurality of inverse matrices of the plurality of square matrices; and adjusting the parameters of the neural network model based on a combination of the plurality of inverse matrices of the plurality of square matrices.
9. The method according to any one of claims 1-7, wherein the neural network model is an image processing model, and wherein obtaining the Kronecker factor matrix comprises: obtaining image training data; and applying the image training data to the image processing model to obtain the Kronecker factor matrix.
10. The method according to any one of claims 1-7, wherein the neural network model is a text processing model, and wherein obtaining the Kronecker factor matrix comprises: obtaining text training data; and applying the text training data to the text processing model to obtain the Kronecker factor matrix.
11. An electronic device, comprising: at least one computing unit; at least one memory coupled to the at least one computing unit and storing instructions for execution by the at least one computing unit, the instructions when executed by the at least one computing unit cause the device to perform the method according to any one of claims 1-10.
12. A computer-readable storage medium having stored thereon a computer program, which when executed by a processor implements the method according to any one of claims 1-10.
Citation Information
Patent Citations
Hydrogen internal combustion engine ignition correct time calibration optimization system based on L-M neural network and optimization method thereof
CN106066606A
Acceleration method of convolutional neural network parallelization training
CN108090565A