An image processing method, a data processing method, a device, a medium and a product
By scale-up, rearrange elements and tensor column decomposition of tensor operators, the third tensor operator is generated, which solves the stress problems of the computational complexity and storage requirements of artificial intelligence models, and achieves more efficient image processing task performance.
Patent Information
- Application Number
- CN202411471758.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-22
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2044-10-22
AI Technical Summary
The computational complexity and storage requirements of artificial intelligence models grow with the increase in model size, resulting in huge pressure on computing and storage resources, especially in image processing tasks.
By scale-up, rearrange element processing and tensor column decomposition of the tensor operator to be processed, a third tensor operator is generated, and tensor calculation is used to reduce the computational complexity and storage requirements.
It significantly reduces the complexity and computational complexity of tensor operators, alleviates the pressure on computing resources and storage resources, and improves the performance of image processing tasks.
Smart Images

Figure CN118982723B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to an image processing method, a data processing method, a device, a medium, and a product. Background Art
[0002] With the development of artificial intelligence (AI) technology, the scale and complexity of the artificial intelligence models adopted have been continuously increasing, and the computational complexity and storage requirements involved will also increase with the increase in the model scale, resulting in huge pressure on both computational resources and storage resources to maintain high-performance model calculations. For example, in the field of image processing, when using a vision model to process input image data, due to the complexity of the image data itself, the complexity of the image processing task, and the sharp increase in the number of image processing tasks, etc., it brings huge computational pressure and storage pressure to the computing device for performing image processing.
[0003] How to relieve the pressure on the demand for computational resources and storage resources in artificial intelligence model calculations is a technical problem that those skilled in the art need to solve. Summary of the Invention
[0004] The purpose of the present invention is to provide an image processing method, a data processing method, a device, a medium, and a product for relieving the pressure on the demand for computational resources and storage resources in artificial intelligence models.
[0005] To solve the above technical problem, the present invention provides an image processing method, including:
[0006] Receiving an input image;
[0007] After vectorizing the input image by using an image processing model, performing tensor calculations according to the obtained sequence data;
[0008] During the process of performing tensor calculations, expanding the scale of the tensor operator to be subjected to the tensor calculations so that the number of elements in each dimension is the product of z positive integers to obtain a first tensor operator; rearranging the elements of the first tensor operator to increase the dimension of the first tensor operator while reducing the number of elements in each dimension of the first tensor operator to obtain a second tensor operator, and the dimension increase amount of each dimension of the second tensor operator corresponding to the first tensor operator is the same; performing tensor train decomposition processing on the second tensor operator to obtain a third tensor operator; using the third tensor operator to perform the tensor calculations;
[0009] Outputting an image processing result;
[0010] Wherein, z is a positive integer greater than 1.
[0011] On the one hand, processing the tensor operator into the third tensor operator includes:
[0012] Determining a conversion coefficient for processing the tensor operator into the third tensor operator according to the model calculation accuracy requirement parameter of the image processing model.
[0013] On the other hand, determining a conversion coefficient for processing the tensor operator into the third tensor operator according to the model calculation accuracy requirement parameter of the image processing model includes:
[0014] Determining the rank number of the third tensor operator according to the model calculation accuracy requirement parameter.
[0015] On the other hand, processing the tensor operator into the third tensor operator includes:
[0016] Determining a conversion coefficient for processing the tensor operator into the third tensor operator according to the model calculation accuracy requirement parameter of the image processing model and the size of the calculation parameter storage space allocated to the tensor calculation.
[0017] On the other hand, determining a conversion coefficient for processing the tensor operator into the third tensor operator according to the model calculation accuracy requirement parameter of the image processing model and the size of the calculation parameter storage space allocated to the tensor calculation includes:
[0018] Determining the minimum rank number allowed for the third tensor operator according to the model calculation accuracy requirement parameter;
[0019] Generating the third tensor operator according to the minimum rank number;
[0020] Determining the number of parameters required to perform the tensor calculation using the third tensor operator. If it is not sufficient to occupy the calculation parameter storage space allocated to the tensor calculation, increase the rank number of the third tensor operator, regenerate the third tensor operator, and re-determine the number of parameters required to perform the tensor calculation using the third tensor operator.
[0021] On the other hand, determining a conversion coefficient for processing the tensor operator into the third tensor operator according to the model calculation accuracy requirement parameter of the image processing model and the size of the calculation parameter storage space allocated to the tensor calculation includes:
[0022] Determining the minimum rank number allowed for the third tensor operator according to the model calculation accuracy requirement parameter;
[0023] Generating the third tensor operator according to the minimum rank number;
[0024] The tensor operator is scaled up and its elements are rearranged to obtain the second tensor operator. The second tensor operator rearranges the elements of each dimension corresponding to the tensor operator into k dimensions, and the number of elements in each dimension of the k-dimensional elements can be expressed in exponential form, and both the base and the exponent are positive integers;
[0025] Determine the number of parameters of a single third tensor operator according to the product of the bases at the corresponding positions of each dimension in the second tensor operator and the rank of the third tensor operator. If the base can be square-rooted to obtain a non-1 positive integer, then the base is square-rooted before performing the product calculation;
[0026] If the maximum value of the number of parameters of a group of the third tensor operators is not sufficient to occupy the calculation parameter storage space allocated for a single group of calculations of the tensor calculation, then change the scaling factor of the tensor operator or the element rearrangement method of the first tensor operator. After regenerating the second tensor operator, return to determine the number of parameters of a single third tensor operator according to the product of the bases at the corresponding positions of each dimension in the second tensor operator and the rank of the third tensor operator.
[0027] On the other hand, according to the model calculation accuracy requirement parameters of the image processing model and the size of the calculation parameter storage space allocated for the tensor calculation, determine the conversion coefficient for processing the tensor operator into the third tensor operator, including:
[0028] The tensor operator is scaled up and its elements are rearranged to obtain the second tensor operator. The second tensor operator rearranges the elements of each dimension corresponding to the tensor operator into k dimensions, and the number of elements in each dimension of the k-dimensional elements can be expressed in exponential form, and both the base and the exponent are positive integers;
[0029] Determine the number of parameters of a single third tensor operator according to the product of the bases at the corresponding positions of each dimension in the second tensor operator and the rank of the third tensor operator. If the base can be square-rooted to obtain a non-1 positive integer, then the base is square-rooted before performing the product calculation;
[0030] Determine the number of the third tensor operators according to the sum of the exponents of a single dimension in the second tensor operator;
[0031] Determine the number of parameters of the tensor calculation according to the number of parameters of a single third tensor operator and the number of the third tensor operators;
[0032] If the number of parameters of the tensor calculation is not sufficient to occupy the storage space for the calculation parameters allocated to the tensor calculation, then change the scale expansion factor of the tensor operator or the way of rearranging elements of the first tensor operator. After regenerating the second tensor operator, return the number of parameters of a single third tensor operator determined according to the product of the bases at the corresponding positions of each dimension in the second tensor operator and the rank number of the third tensor operator, and determine the number of the third tensor operators according to the sum of the exponents of a single dimension in the second tensor operator.
[0033] On the other hand, processing the tensor operator into the third tensor operator includes:
[0034] Determine the rank number of the third tensor operator according to the model calculation accuracy requirement parameters;
[0035] Expand the scale of the tensor operator and perform element rearrangement processing to obtain the second tensor operator. The second tensor operator is obtained by rearranging the elements of each dimension corresponding to the tensor operator into k dimensions, and the number of elements in each dimension of the k-dimensional elements can be expressed in exponential form, and both the base and the exponent are positive integers;
[0036] Determine the number of parameters of a single third tensor operator according to the product of the bases at the corresponding positions of each dimension in the second tensor operator and the rank number of the third tensor operator, where if the base can be square-rooted to obtain a non-1 positive integer, then perform square-root calculation on the base and then perform product calculation;
[0037] Determine the number of the third tensor operators according to the sum of the exponents of a single dimension in the second tensor operator;
[0038] Determine the number of parameters of the tensor calculation according to the rank number of the third tensor operator, the number of parameters of a single third tensor operator, and the number of the third tensor operators;
[0039] Taking the minimization of the number of parameters of the tensor calculation as the optimization goal, solve to obtain the way to expand the tensor operator into the first tensor operator and the way to rearrange the second tensor operator into the third tensor operator.
[0040] On the other hand, the tensor calculation is a multiplication calculation of two tensor operators;
[0041] Expanding the scale of the tensor operator includes:
[0042] Expand both of the two tensor operators so that the number of elements in each dimension is the product of z positive integers and the number of elements in each dimension meets the requirements for multiplying the two tensor operators.
[0043] On the other hand, both of the two tensor operators performing multiplication calculation are two-dimensional tensor operators, and the number of columns of the multiplicand tensor operator among the two tensor operators is the same as the number of rows of the multiplier tensor operator among the two tensor operators.
[0044] On the other hand, performing the tensor calculation by using the third tensor operator includes:
[0045] After storing the third tensor operator in the cache, reading the third tensor operator from the cache to perform the tensor calculation.
[0046] To solve the above technical problems, the present invention further provides a data processing method, including:
[0047] After determining the information of the storage device and the information of the computing device according to the model calculation task, reading the data to be processed corresponding to the model calculation task from the storage device;
[0048] After vectorizing the data to be processed by using the target model, performing tensor calculation according to the obtained sequence data;
[0049] During the process of performing tensor calculation, expanding the scale of the tensor operator to be subjected to the tensor calculation so that the number of elements in each dimension is the product of z positive integers to obtain a first tensor operator; performing rearrangement of elements on the first tensor operator to increase the dimension of the first tensor operator while reducing the number of elements in each dimension of the first tensor operator to obtain a second tensor operator, and the dimension increasing quantity of each dimension of the second tensor operator corresponding to the first tensor operator is the same; performing tensor train decomposition processing on the second tensor operator to obtain a third tensor operator; writing the third tensor operator into the calculation parameter storage space of the computing device to perform the tensor calculation by using the third tensor operator;
[0050] Outputting a processing result corresponding to the data to be processed;
[0051] Wherein, z is a positive integer greater than 1.
[0052] To solve the above technical problems, the present invention further provides a computing device, including:
[0053] A memory for storing a computer program;
[0054] A processor for executing the computer program, and when the computer program is executed by the processor, implementing the steps of the image processing method described in any one of the above or the steps of the data processing method described above.
[0055] To solve the above technical problems, the present invention further provides a non-volatile storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the image processing method or the steps of the data processing method described in any one of the above are implemented.
[0056] To solve the above technical problems, the present invention further provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the image processing method or the steps of the data processing method described in any one of the above are implemented.
[0057] The beneficial effect of the image processing method provided by the present invention is that when using an image processing model to process an input image, when using the image processing model to vectorize the input image and perform tensor calculations, the tensor operator to be subjected to tensor calculations is scaled up so that the number of elements in each dimension is in the form of a product of z positive integers, and then dimensionality increase processing is performed and the number of dimensionality increases in each dimension is the same, and then tensor train decomposition processing is performed. While using tensor train decomposition to decompose the tensor operator and reduce the number of operator parameters, it is adapted to perform tensor train decomposition into a suitable form when the computing device executes model calculations, so that when the computing device executes an image processing task, the complexity of the tensor operator can be significantly reduced, the complexity of tensor calculations can be reduced, the pressure on computing resources can be alleviated, at the same time the number of parameters of the tensor operator is reduced, the pressure on storage resources can be alleviated, and a single tensor calculation can be converted into parallel calculations of multiple groups of small-scale tensor operators, and the performance of the image processing task can be improved when computing resources permit.
[0058] In the data processing method provided by the present invention, when determining the information of the storage device and the information of the computing device according to a model calculation task and reading the data to be processed corresponding to the model calculation task from the storage device into the calculation parameter storage space of the computing device to execute the model calculation task, when using a target model to vectorize the data to be processed and perform tensor calculations, the tensor operator to be subjected to tensor calculations is scaled up so that the number of elements in each dimension is in the form of a product of z positive integers, and then dimensionality increase processing is performed and the number of dimensionality increases in each dimension is the same, and then tensor train decomposition processing is performed. While using tensor train decomposition to decompose the tensor operator and reduce the number of operator parameters, it is adapted to perform tensor train decomposition into a suitable form when the computing device executes model calculations, so that when the computing device executes a model calculation task, the complexity of the tensor operator can be significantly reduced, the complexity of tensor calculations can be reduced, the pressure on computing resources can be alleviated, at the same time the number of parameters of the tensor operator is reduced, the pressure on storage resources can be alleviated, and a single tensor calculation can be converted into parallel calculations of multiple groups of third tensor operators, and the performance of the model calculation task can be improved when computing resources permit.
[0059] The computing device, non-volatile storage medium, and computer program product provided by the present invention have the above-mentioned beneficial effects, which will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] In order to more clearly illustrate the technical solutions of the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0061] Figure 1 It is a flowchart of a data processing method provided by an embodiment of the present invention;
[0062] Figure 2 It is a schematic structural diagram of a data processing system provided by an embodiment of the present invention;
[0063] Figure 3 It is a schematic structural diagram of another data processing system provided by an embodiment of the present invention;
[0064] Figure 4 It is a flowchart of an image processing method provided by an embodiment of the present invention;
[0065] Figure 5 It is a schematic structural diagram of a computing device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0066] The core of the present invention is to provide an image processing method, a data processing method, a device, a medium, and a product, which are used to relieve the pressure on the demand for computing resources and storage resources by artificial intelligence models.
[0067] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0068] With the rapid rise of large models and the new development of artificial intelligence (AI) technology, a large number of model calculation tasks are involved in the process of training artificial intelligence models and using artificial intelligence models for inference calculation to solve practical problems. In model calculation tasks, tensor calculation is particularly resource-consuming in terms of computing resources and storage resources, and the complexity of model calculation and the huge number of parameters bring great pressure on computing resources and storage resources.
[0069] In model calculations, intermediate calculation results are represented by tensors. A tensor is a multilinear mapping defined on the Cartesian product of some vector spaces and some dual spaces. Its coordinates are a quantity with |n| components in an |n|-dimensional space, where each component is a function of the coordinates, and when the coordinates are transformed, these components also undergo linear transformations according to certain rules. r is called the rank or order of the tensor.
[0070] Taking the Vision Transformer model (hereinafter referred to as the Transformer model) in the field of computer vision as an example, it enables computer systems to understand and interpret image data or video data, and has made remarkable progress and been widely used in many fields such as industry, healthcare, security monitoring, autonomous driving, and virtual reality. The Transformer model treats image data as sequence data, uses the self-attention mechanism to capture spatial and temporal information, and captures the position information of the image through position encoding, becoming one of the mainstream methods for tasks such as image classification, object detection, and image generation. When processing image data, the Transformer model first performs an image segmentation (flatten patches) operation on the image. For example, an image with a resolution of 224*224 is processed into 196 small patches through a 16*16 convolutional kernel, and then flattened into a one-dimensional 196, which becomes the sequence data used in natural language processing. Then, a class token is added to each small patch. This information is the classification record added to the image and is used as a parameter to obtain the best parameter data after learning through subsequent learning for predicting the final classification result of the image. Subsequently, the data with the added class token is added to the position encoding matrix. The position information reflects the spatial position information of the segmented data in the original image, and adding this information helps to obtain better parameter values in the subsequent learning process. Then, it enters the Transformer encoder process, which obtains the processing result of this process through a combination of steps such as data normalization, multi-head attention mechanism, and multi-layer perceptron (MLP) calculation for the data processed by the multi-layer perceptron head (MLP Head), thereby obtaining the final classification result. And almost in every link of the Transformer model processing image data, there is tensor calculation, and its computational complexity increases with the growth of tensor elements, especially more significantly in the multiplication calculation of tensors.
[0071] It can be seen that when facing large-scale model calculations, the required computational complexity is very high. Therefore, it is difficult to meet the needs of high-performance model calculations by directly using tensor operators for tensor calculations in the current model calculation scheme.
[0072] To alleviate the computational pressure and storage pressure brought about by large-scale model calculations, embodiments of the present invention provide an image processing method and a data processing method. By expanding the scale of tensor operators to be tensor-calculated in model calculations to a form where the number of elements in each dimension is the product of z positive integers, then performing dimension-increasing processing with the same dimension-increasing quantity for each dimension, and then performing tensor train decomposition processing, while using tensor train decomposition to achieve the decomposition of tensor operators and reduce the number of operator parameters, it is adapted to perform tensor train decomposition into a suitable form when the computing device executes model calculations. Thus, when the computing device executes an image processing task, it can significantly reduce the complexity of tensor operators, reduce the complexity of tensor calculations, alleviate the pressure on computing resources, and at the same time reduce the number of parameters of tensor operators, alleviate the pressure on storage resources, and can convert a single tensor calculation into parallel calculations of multiple groups of small-scale tensor operators, and can improve the performance of the image processing task when computing resources permit.
[0073] Embodiments of the present invention provide an image processing method and a data processing method, which can be applied to a single computing device or in a cluster composed of multiple computing devices. If a cluster is adopted, each computing device can use the same type of computing device or heterogeneous computing devices. The types of computing devices can include, but are not limited to, Graphics Processing Unit (GPU), Field Programmable Gate Array (FPGA), Application Specific Integrated Circuit (ASIC), and Data Processing Unit (DPU).
[0074] Whether performing a model training task or a model inference task, a large amount of storage space is required to store model parameters and data to be processed. This storage space can be provided by the computing device, so the computing device can be a single host, or a host + accelerator architecture. This storage space can also be provided by another storage device, or by a storage pool composed of multiple storage devices in a cluster. To improve the model calculation performance, after decomposing large-scale tensor operators into small-scale tensor operators, the high-speed cache carried by the computing device can be fully utilized to further improve the calculation efficiency.
[0075] Based on the above architecture, the data processing method provided by embodiments of the present invention will be described below with reference to the accompanying drawings.
[0076] Figure 1 It is a flowchart of a data processing method provided by embodiments of the present invention; Figure 2 It is a schematic structural diagram of a data processing system provided by embodiments of the present invention.
[0077] As Figure 1 shown, the data processing method provided by the embodiment of the present invention includes:
[0078] S101: After determining the information of the storage device and the information of the computing device according to the model calculation task, read the data to be processed corresponding to the model calculation task from the storage device;
[0079] S102: After vectorizing the data to be processed by using the target model, perform tensor calculation according to the obtained sequence data;
[0080] S103: During the process of performing tensor calculation, expand the scale of the tensor operator to be calculated to the number of elements in each dimension being the product of z positive integers to obtain the first tensor operator; rearrange the elements of the first tensor operator to increase the dimension of the first tensor operator while reducing the number of elements in each dimension of the first tensor operator to obtain the second tensor operator, and the dimension increase amount of each dimension of the second tensor operator corresponding to the first tensor operator is the same; perform tensor train decomposition processing on the second tensor operator to obtain the third tensor operator; write the third tensor operator into the calculation parameter storage space of the computing device to perform tensor calculation by using the third tensor operator;
[0081] S104: Output the processing result corresponding to the data to be processed.
[0082] Wherein, z is a positive integer greater than 1.
[0083] In a specific implementation, the data processing method provided by the embodiment of the present invention can be applied to a data processing system as Figure 2 shown. The data processing system includes one or more computing devices 201. The storage space for storing model calculation related parameters (model parameters and data to be processed) can be provided by the computing device 201 or the storage device 202. All parameters can be stored in the storage device 202, and the parameters required for the current calculation are moved to the memory or cache of the computing device 201.
[0084] For S101, determine the information of the storage device and the information of the computing device according to the model calculation task, that is, determine the read and write positions of the parameters required for the model calculation according to the information of the data processing system where it is located, obtain the computing resources and storage resources for the intermediate calculation process, and can determine the degree of scale reduction required for the tensor operator according to the size of the calculation parameter storage space allocated to the model calculation task.
[0085] For S102, according to the type of the model calculation task, the target model can be a model to be trained or a pre-trained model, and can be an image processing model, a video processing model, a language model, etc. Correspondingly, the data to be processed can be training samples or input data in model inference calculation, and can be image data, video data, text data, etc.
[0086] Input the data to be processed into the target model. First, perform vectorization processing to obtain sequence data, and then perform one or more tensor calculations according to the calculation type included in the target model.
[0087] For S103, during the process of performing tensor calculations, perform transformation processing on the tensor operator to be calculated instead of directly performing tensor calculations. This transformation processing process includes first expanding the scale of the tensor operator to obtain a first tensor operator, then rearranging the elements of the first tensor operator to obtain a second tensor operator, then performing tensor train decomposition on the second tensor operator to obtain a third tensor operator, and then using the third tensor operator to perform the originally required tensor calculations. After this transformation processing, the tensor operator is decomposed into intermediate results with very small storage complexity, which can reduce the occupation of data on the storage space, thereby weakening the pressure on the large storage space requirement of the computing device during the model calculation process. At the same time, the tensor calculation of this small-scale third tensor operator can consume less computing resources, thereby reducing the overall model calculation amount, further accelerating the model calculation rate, and improving the model calculation performance. It can also further use the cache to store the parameters required for tensor calculations, thereby further improving the model calculation efficiency.
[0088] It should be noted that since a large number of accelerators such as GPUs are currently introduced in model calculations, the pressure brought by model calculations to computing resources is not as obvious as the pressure brought to storage resources. And due to the reduction in the scale of the tensor operator, although the transformation processing of the tensor operator before tensor calculation in the embodiments of the present invention increases the calculation steps, it will not additionally increase the calculation pressure and instead reduce the model calculation performance, but significantly reduces the storage pressure and can bring a great improvement in the model calculation performance.
[0089] It should be noted that the first tensor operator, the second tensor operator, and the third tensor operator in the embodiments of the present invention only represent the results obtained after different transformations of the tensor operator, and do not refer to a unique tensor, because for different tensor calculations or different tensor operators in the same tensor calculation, different transformation parameters can be used.
[0090] When expanding the scale of a tensor operator to obtain a first tensor operator, for the dimensions in the tensor operator that already satisfy the form of the product of z positive integers in terms of the element scale, no scale expansion is required, and only the dimensions that do not satisfy the form of the product of z positive integers in terms of the element scale are expanded. However, it is necessary to ensure that each dimension of the first tensor operator can be expressed as the product of the same number of positive integers, without restricting the elements in each dimension to be the same. For example, the scale of the first tensor operator can be (6, 35), that is, (2×3, 5×7), so that the scale of the second tensor operator obtained after rearranging the elements can be (2, 3; 5, 7).
[0091] The scale expansion can be achieved by padding with 0 elements. For example, a two-dimensional matrix with a scale of (768, 768) is padded with 0s to be expanded into a matrix with 784 rows and 784 columns as the second tensor operator. The first 768 rows and the first 768 columns belong to the elements of the original matrix, and the extra 16 rows and 16 columns are all 0 elements.
[0092] In some optional embodiments of the embodiments of the present invention, on the premise that the number of elements in each dimension of the first tensor operator is the product of z positive integers, the goal can be to minimize the number of elements added to each dimension of the tensor operator, thereby reducing the number of parameters of all the finally obtained third tensor operators.
[0093] In addition, when expanding the scale of a tensor operator to obtain a first tensor operator, the computational requirements of tensor calculations also need to be considered. For example, if the tensor calculation is a multiplication calculation of two tensor operators, then expanding the scale of the tensor operator can include: expanding both tensor operators to the number of elements in each dimension being the product of z positive integers and the number of elements in each dimension satisfying the requirements for the multiplication calculation of the two tensor operators.
[0094] Taking the example that both tensor operators for multiplication calculation are two-dimensional tensor operators, the number of columns of the multiplicand tensor operator in the two tensor operators should be the same as the number of rows of the multiplier tensor operator in the two tensor operators.
[0095] When performing element rearrangement processing on the first tensor operator to obtain a second tensor operator, to enable the computing device to generate a decomposition form with appropriate dimensions during tensor column decomposition processing, it is necessary to ensure that the dimension increase amounts of the second tensor operator corresponding to each dimension of the first tensor operator are the same. For example, if the first tensor operator is two-dimensional, the second tensor operator can be obtained by decomposing the rows and columns of the first tensor operator into two dimensions respectively to obtain a four-dimensional second tensor operator, or by decomposing the rows and columns of the first tensor operator into three dimensions respectively to obtain a six-dimensional second tensor operator.
[0096] To reduce the computational complexity, the first tensor operator is rearranged to obtain the second tensor operator. The rearrangement of the first tensor operator can be performed multiple times to gradually rearrange it to the second tensor operator.
[0097] After that, the second tensor operator is processed by tensor train decomposition to obtain multiple small-scale third tensor operators. The second tensor operator can be approximately equal to the product of the converted multiple third tensor operators and the rank number (boundary condition). Therefore, using the third tensor operator instead of the tensor operator for tensor calculation brings a loss of accuracy to the model calculation. The magnitude of this accuracy loss is mainly determined by the rank number selected when the second tensor operator is processed by tensor train decomposition to obtain the third tensor operator.
[0098] Then, in some optional embodiments of the present invention, in S103, processing the tensor operator into the third tensor operator may include: determining the conversion coefficient for processing the tensor operator into the third tensor operator according to the model calculation accuracy requirement parameter of the target model. Tensor train decomposition decomposes a large-scale high-dimensional tensor into the product of small-scale tensors and the rank number. By selecting a smaller rank number, a more significant scale reduction effect can be achieved, but this also brings a reduction in the accuracy of model calculation. Therefore, when processing the tensor operator into the third tensor operator, an appropriate rank number can be selected according to the model calculation accuracy requirement parameter to achieve the effect of reducing the number of parameters and computational complexity of tensor calculation while satisfying the model calculation accuracy.
[0099] In practical applications, according to the requirements, all tensor calculations in the model calculation can be processed by the transformation as introduced in S103, or only some tensor calculations can be processed.
[0100] To further improve the model calculation speed, in S103, using the third tensor operator to perform tensor calculation may include: after storing the third tensor operator in the cache, reading the third tensor operator from the cache to perform tensor calculation. Since the tensor operator is processed into the third tensor operator, the tensor calculation is also converted from the calculation between large-scale tensor operators to the calculation between multiple small-scale tensor operators. Therefore, the tensor calculation that could not use the cache with a small storage space due to too many calculation parameters can also use the cache to store the calculation parameters required for one or more small-scale tensor calculations, making full use of the read and write performance of the cache.
[0101] For S104, after completing the model calculation, the processing result of the data to be processed is output.
[0102] The data processing method provided by the embodiments of the present invention, when determining the information of the storage device and the information of the computing device according to the model calculation task and reading the data to be processed corresponding to the model calculation task from the storage device into the calculation parameter storage space of the computing device to execute the model calculation task, when using the target model to perform vectorization processing on the data to be processed and perform tensor calculation, the tensor operator to be subjected to tensor calculation is scaled up so that the number of elements in each dimension is in the form of a product of z positive integers, then dimensionality increase processing is performed and the number of dimensionality increases in each dimension is the same, and then tensor train decomposition processing is performed. While using tensor train decomposition to achieve the decomposition of the tensor operator and reduce the number of operator parameters, it is adapted to the tensor train decomposition into a suitable form when the computing device performs model calculation, so that when the computing device executes the model calculation task, the complexity of the tensor operator can be significantly reduced, the complexity of tensor calculation can be reduced, the pressure on the computing resources can be alleviated, at the same time the number of parameters of the tensor operator is reduced, the pressure on the storage resources can be alleviated, and a single tensor calculation can be converted into parallel calculations of multiple groups of third tensor operators, and the performance of the model calculation task can be improved when the computing resources permit.
[0103] On the basis of the above embodiments, the embodiments of the present invention continue to describe the steps of processing the tensor operator into a third tensor operator.
[0104] In the above embodiments, it is introduced that since the expression of the third tensor operator after tensor train decomposition is approximately equal to that of the second tensor operator, the data processing method provided by the embodiments of the present invention will lose the model calculation accuracy to a certain extent. Then, in S103, processing the tensor operator into a third tensor operator includes: determining a conversion coefficient for processing the tensor operator into a third tensor operator according to the model calculation accuracy requirement parameter of the target model.
[0105] Among them, determining a conversion coefficient for processing the tensor operator into a third tensor operator according to the model calculation accuracy requirement parameter of the target model may include: determining the rank number of the third tensor operator according to the model calculation accuracy requirement parameter. According to the relationship between the expression of the third tensor operator after tensor train decomposition and the second tensor operator and the type of tensor calculation, for a given tensor calculation, a relational expression between the rank number selected for tensor train decomposition and the model calculation accuracy requirement parameter can be listed, so as to determine the rank number of the third tensor operator according to the model calculation accuracy required by the user for the model calculation. The model calculation accuracy requirement parameter is determined according to the type of the target model. For example, if the target model is a classification model, the model calculation accuracy requirement parameter may be the maximum classification error, and the model calculation accuracy requirement is that the classification result does not exceed the maximum classification error.
[0106] Alternatively, processing the tensor operator as a third tensor operator may further include: determining a conversion coefficient for processing the tensor operator as the third tensor operator according to the model calculation accuracy requirement parameter of the target model and the size of the calculation parameter storage space allocated for tensor calculation. That is, through the selection of the rank number, while significantly reducing the overall calculation parameter storage space required for tensor calculation on the premise of meeting the model calculation accuracy requirement, it is also possible to consider the size of the calculation parameter storage space allocated for tensor calculation by the computing device, with the goal of making full use of the calculation parameter storage space, to reduce the model calculation loss and improve the model calculation accuracy.
[0107] Among them, the calculation parameter storage space can be the memory of the computing device or the cache of the computing device.
[0108] Among them, determining the conversion coefficient for processing the tensor operator as the third tensor operator according to the model calculation accuracy requirement parameter of the target model and the size of the calculation parameter storage space allocated for tensor calculation may include: determining the minimum rank number allowed for the third tensor operator according to the model calculation accuracy requirement parameter; generating the third tensor operator according to the minimum rank number; determining the number of parameters required to perform tensor calculation using the third tensor operator. If it is not enough to occupy the calculation parameter storage space allocated for tensor calculation, increase the rank number of the third tensor operator and then regenerate the third tensor operator, and re-determine the number of parameters required to perform tensor calculation using the third tensor operator.
[0109] That is to say, it is possible to first determine the minimum rank number according to the relationship between the rank number selected by tensor train decomposition and the model calculation accuracy requirement parameter and the user's model calculation accuracy requirement parameter, and then gradually increase the rank number to fill up the calculation parameter storage space to make full use of the storage resources.
[0110] In some alternative embodiments of the embodiments of the present invention, according to the model calculation accuracy requirement parameter of the target model and the size of the calculation parameter storage space allocated to tensor calculation, determining the conversion coefficient for processing the tensor operator into a third tensor operator may further include: determining the minimum rank number allowed for the third tensor operator according to the model calculation accuracy requirement parameter; generating the third tensor operator according to the minimum rank number; expanding the scale of the tensor operator and performing a rearrangement element process to obtain a second tensor operator, where the second tensor operator rearranges the elements of each dimension of the corresponding tensor operator into k dimensions, and the number of elements in each dimension of the k-dimensional elements can be expressed in exponential form, and both the base and the exponent are positive integers; determining the number of parameters of a single third tensor operator according to the product of the bases at the corresponding positions of each dimension in the second tensor operator and the rank number of the third tensor operator, where if the base can be square-rooted to obtain a non-1 positive integer, then the base is square-rooted before the product calculation; if the maximum value of the number of parameters of a group of third tensor operators is not sufficient to occupy the calculation parameter storage space allocated to the single-group calculation of tensor calculation, then change the scale expansion scale of the tensor operator or the rearrangement element method of the first tensor operator, and after regenerating the second tensor operator, return to determining the number of parameters of a single third tensor operator according to the product of the bases at the corresponding positions of each dimension in the second tensor operator and the rank number of the third tensor operator.
[0111] That is to say, through the data processing method provided by the embodiments of the present invention, a single tensor calculation can be converted into tensor calculations of multiple groups of small-scale tensor operators (third tensor operators). At this time, if parallel calculation of each group of third tensor operators is implemented, it can be considered that the number of parameters required for each group of tensor operators in the parallel calculation fully occupies the calculation parameter storage space allocated to the tensor operators of this group.
[0112] In some other alternative embodiments of the embodiments of the present invention, according to the model calculation accuracy requirement parameters of the target model and the size of the calculation parameter storage space allocated to tensor calculation, determining the conversion coefficient for processing the tensor operator into a third tensor operator may further include: expanding the scale of the tensor operator and performing element rearrangement processing to obtain a second tensor operator, where the second tensor operator rearranges the elements of each dimension of the corresponding tensor operator into k dimensions, and the number of elements in each dimension of the k-dimensional elements can be expressed in exponential form, and both the base and the exponent are positive integers; determining the number of parameters of a single third tensor operator according to the product of the bases at the corresponding positions of each dimension in the second tensor operator and the rank number of the third tensor operator, where if the base can be square-rooted to obtain a non-1 positive integer, then perform square-root calculation on the base and then perform product calculation; determining the number of third tensor operators according to the sum of the exponents of a single dimension in the second tensor operator; determining the number of parameters of tensor calculation according to the number of parameters of a single third tensor operator and the number of third tensor operators; if the number of parameters of tensor calculation is not sufficient to occupy the calculation parameter storage space allocated to tensor calculation, then change the scale expansion scale of the tensor operator or the element rearrangement method of the first tensor operator, regenerate the second tensor operator, and then return to determining the number of parameters of a single third tensor operator according to the product of the bases at the corresponding positions of each dimension in the second tensor operator and the rank number of the third tensor operator and determining the number of third tensor operators according to the sum of the exponents of a single dimension in the second tensor operator.
[0113] That is to say, if all the parameters required for the original tensor calculation are written into the calculation parameter storage space at one time, it is necessary to consider filling the calculation parameter storage space with the scale of all the calculation parameters after processing the tensor operator into a third tensor operator to fully utilize the storage resources.
[0114] In the above embodiments, the second tensor operator can be expressed as form, where the base can be the same or different positive integers, is the exponent. Then the number of the finally obtained third tensor operators is .
[0115] For example, if the scale of the second tensor operator is (2, 3; 5, 7), that is , then the number of the obtained third tensor operators is two, and their sizes are r×2×5×r and r×3×7×r respectively, where r represents the rank number (the two r's on the left and right represent different rank numbers).
[0116] Define to mean that if or can be square-rooted to obtain a non-1 positive integer, then or Continue with the square root calculation, and then multiply the results obtained respectively. Then the size of the th third tensor operator is . Then the number of parameters of all third tensor operators corresponding to the tensor calculation can be obtained according to and
[0117] Based on the above embodiments, in the embodiments of the present invention, processing the tensor operator into a third tensor operator may further include: determining the rank number of the third tensor operator according to the model calculation accuracy requirement parameter; expanding the scale of the tensor operator and performing element rearrangement processing to obtain a second tensor operator, where the second tensor operator rearranges the elements of each dimension of the corresponding tensor operator into k dimensions, and the number of elements in each dimension of the k-dimensional elements can be expressed in exponential form, and both the base and the exponent are positive integers; determining the number of parameters of a single third tensor operator according to the product of the bases at the corresponding positions of each dimension in the second tensor operator and the rank number of the third tensor operator, where if the base can be square-rooted to obtain a non-1 positive integer, then perform square-root calculation on the base and then perform product calculation; determining the number of third tensor operators according to the sum of the exponents of a single dimension in the second tensor operator; determining the number of parameters of the tensor calculation according to the rank number of the third tensor operator, the number of parameters of a single third tensor operator, and the number of third tensor operators; taking the minimization of the number of parameters of the tensor calculation as the optimization goal, and solving the method of expanding the tensor operator into the first tensor operator and the method of rearranging the second tensor operator into the third tensor operator.
[0118] That is to say, it is possible to solve for the conversion coefficient of processing the tensor operator into a third tensor operator on the condition of meeting the model calculation accuracy requirement and with the goal of minimizing the number of parameters of all calculation parameters required for the final tensor calculation.
[0119] Based on the above embodiments, in the embodiments of the present invention, in S103, processing the tensor operator into a third tensor operator may further include: determining the conversion coefficient of processing the tensor operator into a third tensor operator according to the model calculation accuracy requirement parameter of the target model; determining the number of parameters of performing tensor calculation using the third tensor operator according to the conversion coefficient; writing the parameters of the tensor calculation of multiple groups of third tensor operators into the calculation parameter storage space according to the number of parameters of the tensor calculation of a group of third tensor operators and the size of the calculation parameter storage space allocated to the tensor calculation, so as to perform the tensor calculation of multiple groups of third tensor operators using the parameters stored in the calculation parameter storage space.
[0120] Among them, the tensor calculation of a group of third tensor operators can be a tensor calculation based on the third tensor operator once, can also include multiple tensor calculations based on the third tensor operator, can also be the tensor calculation of the third tensor operator after decomposing the tensor operators of multiple tensor calculations in the model calculation task, and can also be the tensor calculation of the third tensor operator after decomposing the tensor operators of the tensor calculations in multiple model calculation tasks.
[0121] Applying the data processing method provided by the embodiments of the present invention, the parameters required for multiple tensor calculations can be stored in the calculation parameter storage space allocated to the tensor calculation by the computing device, which can not only make full use of the storage resources, but also further improve the model calculation rate.
[0122] Based on the above embodiments, a processing module corresponding to the type of tensor calculation involved in the model calculation task can be designed, and this processing module is used to process the tensor operator into a third tensor operator according to a preset conversion coefficient.
[0123] In the embodiments of the present invention, taking the example of obtaining the second tensor operator by performing two reordering element processes on the first tensor operator. Among them, the second reordering element process and the tensor train decomposition process can be collectively referred to as the quantization tensor train decomposition process. It should be noted that the quantization here does not mean converting the data type in the computer, but a reordering element process of increasing the dimension of the low-dimensional tensor and reducing the number of elements in each dimension.
[0124] In the Transformer model, the tensor calculations involved include six types: matrix multiply-add, matrix multiply, matrix add, vector-matrix multiply, scalar-matrix multiply, and vector add. Then, in the embodiments of the present invention, taking the tensor calculation as a matrix calculation as an example, six matrix calculation types such as matrix multiply-add, matrix multiply, matrix add, vector-matrix multiply, scalar-matrix multiply, and vector add are illustrated by examples respectively.
[0125] Type 1: Matrix multiply-add.
[0126] Table 1 shows an example of the scale of input and output parameters involved in five matrix multiply-add calculations.
[0127] Table 1
[0128]
[0129] The matrix multiply-add calculation processes involved in the five examples in Table 1 are all similar, only the matrix scales are different. The embodiments of the present invention only introduce the case of matrix A with a scale of (768, 768), matrix B with a scale of (768, 196), and bias with a scale of (1, 196).
[0130] Taking matrix A as the tensor operator to be processed, the following steps need to be taken:
[0131] Step 1: Scale expansion. Perform a zero-padding operation on matrix A to expand it into a matrix with 784 rows and 784 columns , which is the first tensor operator introduced in the above embodiment. The elements in the first 768 rows and columns are the same as those in matrix A, and the extra 16 rows and columns are all zero elements.
[0132] Step 2: First rearrangement of elements. Rearrange the elements of the matrix to obtain a 6D tensor operator with a size of (4, 4, 49; 4, 4, 49). The mathematical formula for this step is described as follows.
[0133] The matrix is composed of elements , that is:[[]]
[0134] ; ; . (1)
[0135] Among them, , represent the row index and column index of the element in the matrix .
[0136] The higher-dimensional tensor operator obtained after the first rearrangement of elements is composed of elements , that is:[[]]
[0137] ; (2)
[0138] Among them, ; ; ; ; ; . , , , , , represent the indexes of the element in the 6 different dimensions of the higher-dimensional tensor operator .
[0139] The element corresponds to the element of the matrix , and their corresponding relationship is as follows:[[]]
[0140] ; (3)
[0141] Among them, ; ; ; ;
[0142] ; ; ; .
[0143] Thus, the high-dimensional tensor operator is obtained.
[0144] Step 3: Use the quantized tensor train decomposition method provided in the embodiments of the present invention to perform quantized tensor train decomposition on the high-dimensional tensor operator .
[0145] First, it is necessary to perform quantization processing on the high-dimensional tensor operator , that is, the second rearrangement of elements mentioned above. After rearrangement, a 12-dimensional tensor operator with a size of (2, 2, 2, 2, 7, 7; 2, 2, 2, 2, 7, 7) is obtained, that is, the tensor operator is composed of elements .
[0146] The elements of the high-dimensional tensor operator and the elements of the 12-dimensional tensor operator have the following corresponding relationship:
[0147] = ; (4)
[0148] Among them, ; ; ; ;
[0149] ; ; ;
[0150] ; ; ; ;
[0151] ; ; ;
[0152] , , , , , , , , , , , represent elements in the tensor operator at the indices of 12 different dimensions.
[0153] Using the 12 - dimensional tensor operator as the second tensor operator, perform tensor train decomposition:
[0154] ; (5)
[0155] wherein, ( ; ) is a small - scale 4 - dimensional tensor operator of size (i.e., the third tensor operator), ( ) is a small - scale 4 - dimensional tensor operator of size (i.e., the third tensor operator), , represent the indices of this small - scale 4 - dimensional tensor operator among all the third tensor operators obtained by transforming with matrix A. represents the indices of the two - dimensional data remaining after excluding the dimension corresponding to the rank of this small - scale 4 - dimensional tensor operator, and so on.
[0156] is related to the exponents in the second tensor operator because in this example, the size of the high - dimensional tensor operator is (4, 4, 49; 4, 4, 49). If the high - dimensional tensor operator is directly used as the second tensor operator for tensor train decomposition to obtain the third tensor operator, then , then can take values of . And in a more general case, can be expressed as and so on.
[0157] Because , are very small numbers after tensor train decomposition. Additionally , , is the boundary condition. That is to say, if matrix A can be expressed as equation (5), then it is called the quantized tensor train decomposition representation of matrix A with rank . Through the quantized tensor train decomposition processing method provided by the embodiments of the present invention, the complexity of the tensor operator can be reduced from Reduce to , where ; , for example, the third tensor operator obtained in this example includes and two scales, and the corresponding is 7. is the maximum value among the number of elements in each dimension of the tensor operator. In this example, it is the larger value between the number of rows and the number of columns of matrix A.
[0158] For the matrix A with the original scale of (768, 768) exemplified in the embodiment of the present invention, the complexity is reduced , is , and is a very small number. Therefore, the storage complexity after quantized tensor train decomposition is logarithmically dependent on the original scale of each dimension of the tensor . Therefore, compared with the original data scale of the tensor, the storage complexity of the data after quantized tensor train decomposition is greatly reduced, so the occupied storage space will be very small, which will relieve the pressure on the demand for large storage space during the model calculation process.
[0159] For matrix B, in order to use the quantized tensor train decomposition method, the similar steps of the above matrix A are adopted. First, 0s are filled to form a matrix with 784 rows and 256 columns , and then it is rearranged into a 6D tensor operator with the scale of (4, 4, 49; 4, 4, 16) , and then the quantized tensor train decomposition step is performed. First, the quantization process is carried out and rearranged into a 12D tensor operator with the scale of (2, 2, 2, 2, 7, 7; 2, 2, 2, 2, 4, 4) , and then the tensor train decomposition is carried out to obtain the quantized tensor train decomposition of the 12D tensor operator :
[0160] ; (6)
[0161] Among them, ; ; ; ; ;
[0162] ; ; ; ; .
[0163] , , , , , , , , , , , represent elements in the tensor operator at the indices of 12 different dimensions.
[0164] Among them, ( ; ) is a small-scale 4D tensor operator (i.e., the third tensor operator) with a scale of , ( ) is a small-scale 4D tensor operator (i.e., the third tensor operator) with a scale of , , represent the indices of this small-scale 4D tensor operator among all the third tensor operators obtained by transforming with matrix B. represent the indices of the two-dimensional data remaining after excluding the dimensions corresponding to the rank of this small-scale 4D tensor operator, and so on.
[0165] Because , result in very small numbers after decomposition. Additionally, , , are boundary conditions, thus reducing the matrix B with a complexity of to a complexity of , where , and is a very small number. Therefore, the storage complexity after the quantized tensor train decomposition is much smaller than the original data scale of matrix B, thus occupying very little storage space, which will relieve the pressure on the large storage space requirement during the model calculation process.
[0166] Step 4: After obtaining the quantized tensor train decompositions of matrix A and matrix B, then the product C = A * B of matrix A and matrix B can be calculated from the decomposed form, i.e.:
[0167] .
[0168] Through the calculation of the tensor operator multiplication in this step, the calculation complexity of the matrix product can be reduced from the original to , where is the rank of the operator, which is a very small number after decomposition. Therefore, the computational complexity of matrix multiplication will also be significantly reduced due to tensor decomposition. The key is that the multiplication calculation after tensor decomposition is linearly dependent on the dimensions of the tensor operator, while the original matrix multiplication calculation is cubic-dependent on the dimensions of the matrix.
[0169] In the process of converting the matrix into the third tensor operator, the following points need to be noted:
[0170] (1) This type of calculation is matrix multiply-add. In addition to the product of matrix A and matrix B, each row of the result matrix C also needs to be added with a bias vector. However, essentially it is a vector addition operation, and this type of calculation will be introduced in detail in type 6 later.
[0171] (2) When using the method of rearranging elements twice, it is necessary to pay attention to the scale of the tensor after the first rearrangement of elements. It is necessary to ensure that the dimensions that can be quantized subsequently in the left half of the dimension of the rearranged tensor scale are the same as those in the right half. For example, for a tensor of size (768, 196) padded with 0s to (784, 256), it can be rearranged into a high-dimensional tensor operator of size (4, 4, 49; 4, 4, 16) because 4 = 2 * 2, 49 = 7 * 7, 16 = 4 * 4. After such rearrangement, they are all the squares of some integers, that is, the dimensions that can be quantized subsequently in the left half are 6, and the dimensions that can be quantized subsequently in the right half are also 6. However, it cannot be rearranged into a high-dimensional tensor operator of size (4, 4, 49; 8, 8, 4) because 4 = 2 * 2, 49 = 7 * 7, 8 = 2 * 2 * 2. Such rearrangement results in the dimensions that can be quantized subsequently in the left half being 6, and the dimensions that can be quantized subsequently in the right half being 8. This inconsistency will cause the subsequent quantized tensor column decomposition to be unable to generate a decomposition form with appropriate dimensions.
[0172] For this situation, the scale of the first rearrangement of elements corresponding to the five matrix multiply-add calculations shown in Table 1 can be adopted in the way shown in Table 2. Table 2 shows the scale of the first rearrangement of elements corresponding to the five matrix multiply-add calculations shown in Table 1.
[0173] Table 2
[0174]
[0175] (3) As shown in Table 2, the rearrangement scale of the right half of the tensor operator after rearranging matrix A should be the same as the rearrangement scale of the left half of matrix B, so as to meet the rules of matrix multiplication, similar to the number of columns of the original matrix A being equal to the number of rows of matrix B.
[0176] (4) To reduce the overall parameter scale after quantized Tensor-Train decomposition, the operation of padding zeros to the original matrix follows the principle of padding as few zeros as possible for calculation. For example, for the matrix multiplication of a matrix with size (768, 768) and a matrix with size (768, 196), in addition to the zero-padding operations listed in the above table, zeros can also be padded to multiply a matrix with size (1024, 1024) and a matrix with size (1024, 256). During subsequent quantization, it is rearranged into the product of tensor operators with sizes (8, 8, 16; 8, 8, 16) and (8, 8, 16; 8, 8, 4). Although the quantizable dimensions of the left part are equal to those of the right part, which is 8, this rearrangement method requires padding more zeros compared to the rearrangement methods listed in the table. Therefore, it increases the memory requirement and the computational amount. So, in actual calculation, under the condition of meeting the calculation requirements, as few zeros as possible should be padded for calculation.
[0177] Type 2: Matrix multiplication.
[0178] Table 3 shows examples of the input and output parameter scales involved in two matrix multiplication calculations.
[0179] Table 3
[0180]
[0181] In the calculation process of the Transformer model, the above two matrix multiplications refer to performing 12 matrix multiplication operations. Each time, a matrix with size (197, 64) is multiplied by a matrix with size (64, 197) or a matrix with size (197, 197) is multiplied by a matrix with size (197, 64). These two matrix multiplications continue to use the processing method of Type 1 to be converted into tensor operators for product operations, and the rearrangement method shown in Table 4 can be adopted. Table 4 shows the scales of the first rearrangement element processing corresponding to the two matrix multiplication and addition calculations shown in Table 3.
[0182] Table 4
[0183]
[0184] Type 3: Matrix addition.
[0185] Table 5 shows the matrix scale schematic of matrix addition.
[0186] Table 5
[0187]
[0188] For the matrix addition operation of the above scales, the quantized Tensor-Train decomposition method is used in the following steps:
[0189] Step 1: Perform zero-padding operations on both matrix A and matrix B to supplement them into matrices with a size of (256, 784). and matrix .
[0190] Step 2: Rearrange the zero-padded matrix into a 6D tensor operator with a size of (4, 4, 16; 4, 4, 49). and .
[0191] Step 3: Using the quantized tensor train decomposition method, first rearrange it into a tensor operator with a size of (2, 2, 2, 2, 4, 4; 2, 2, 2, 2, 7, 7). and , and then perform tensor train decomposition processing as shown in the following formula:
[0192] ,
[0193] ; (6)
[0194] where ; ; ; ; ;
[0195] ; ; ; ; .
[0196] Step 4: After obtaining the decomposed third tensor operator, the matrix addition calculation of C = A + B can be performed. The specific calculation of the tensor operator addition is as follows:
[0197] ,
[0198] ,
[0199] where ; ; ;
[0200] .
[0201] Thus, the calculation result of the tensor operator addition is obtained. Through this method of operation, the original matrix addition operation requires computational complexity, but the addition of tensor operators does not require calculation and only realizes the addition operation through element rearrangement, thus reducing the computational complexity.
[0202] For the symbol explanations in the tensor operator in this example, please refer to Type 1.
[0203] Type 4: Vector-matrix multiplication.
[0204] Table 6 is a schematic diagram of the parameter scale for vector-matrix multiplication.
[0205] Table 6
[0206]
[0207] The following steps are required to calculate the multiplication of a vector and a matrix using the quantized tensor train decomposition method:
[0208] Step 1: Pad the vector a and the matrix B with zeros to expand them into a vector of size (1, 784) and a matrix of size (784, 1024).
[0209] Step 2: Rearrange the elements of the zero-padded vector and matrix into a 3D tensor operator of size (4, 4, 49) and a 6D tensor operator of size (4, 4, 49; 16, 16, 4). .
[0210] Step 3: Perform quantized tensor train decomposition. First, rearrange the elements of the tensor operator into a 6D tensor of size (2, 2, 2, 2, 7, 7) , and then perform tensor decomposition to obtain:
[0211] ;
[0212] where, ; ; ; ; .
[0213] Then rearrange the tensor operator into a 12D tensor operator of size (2, 2, 2, 2, 7, 7; 4, 4, 4, 4, 2, 2) , and then perform tensor decomposition to obtain:
[0214] ;
[0215] where, ; ; ; ; ;
[0216] ; ; ; ; 。
[0217] Step 4: Thus, the product of vector a and matrix B can be obtained through vector-matrix multiplication using the corresponding kernel of the tensor and the tensor operator. The specific calculation is as follows:
[0218] 。
[0219] The resulting tensor is thus obtained , which consists of elements 。
[0220] For the vector-matrix multiplication completed in the above tensor form, the key point is that the computational complexity of the original vector-matrix multiplication depends quadratically on the dimension of the matrix, but in the tensor calculation, it depends linearly on the dimension of the tensor. Specifically, the computational complexity required for the original vector-matrix multiplication is , but the amount of computation required by the tensor calculation method is , where is a very small number. Therefore, the computational complexity can be significantly reduced compared to the original calculation method.
[0221] For the symbol description in the tensor operator in this example, please refer to Type 1.
[0222] Type 5: Scalar-matrix multiplication.
[0223] Table 7 is a schematic diagram of the parameter scale of scalar-matrix multiplication.
[0224] Table 7
[0225]
[0226] In the calculation process of the Transformer model, the above scalar-matrix product refers to performing 12 scalar-matrix product operations, and each calculation is a scalar multiplied by a matrix of size (197, 64). Then, the following steps are required to calculate using the quantized tensor train decomposition method:
[0227] Step 1: Perform a zero-padding operation on matrix A to expand it into a matrix of size (256, 64).
[0228] Step 2: Rearrange the elements of the zero-padded matrix into a 6-dimensional tensor operator , and the size of this tensor operator is (4, 4, 16; 4, 4, 4).
[0229] Step 3: Perform quantized tensor train decomposition. First, rearrange the elements of a 12-dimensional tensor operator of size (2, 2, 2, 2, 4, 4; 2, 2, 2, 2, 2, 2) , and then perform tensor decomposition to obtain:
[0230] ;
[0231] Among them, ; ; ; ; ;
[0232] ; ; .
[0233] Step 4: Thus, the product of the scalar a and the matrix A can be obtained through the tensor scalar multiplication of the scalar a and the tensor operator , which is essentially to multiply the scalar by any one of the kernels of the tensor operator. For example, multiply by the first kernel. The specific calculation is as follows: It is composed of elements
[0234] Each element is:
[0235] ;
[0236]
[0234] Among them, .
[0237] Thus, the calculation result of multiplying the scalar by the tensor operator is obtained. Through this operation method, the original scalar matrix multiplication operation requires computational complexity, but the computational complexity required for multiplying the scalar by the tensor operator is , where is a very small number. Therefore, the computational complexity can be reduced compared with the original calculation method.
[0238] For the symbol description in the tensor operator in this example, please refer to Type 1.
[0239] Type 6: Vector addition.
[0240] Table 8 is a schematic diagram of the parameter scale of vector addition.
[0241] Table 8
[0242]
[0243] To use the quantized tensor train decomposition method, the following steps are required:
[0244] Step 1: Perform a zero-padding operation on the vector a to expand it into a vector with 1024 elements.
[0245] Step 2: Rearrange its elements into a 3D tensor operator , the scale of this tensor is (16, 8, 8), so that the quantized tensor train decomposition can be used.
[0246] Step 3: First, perform quantization processing on this 3D tensor, that is, keep the number of elements of the tensor operator unchanged, only change the arrangement of its elements, and rearrange the tensor operator into a tensor with a scale of (2, 2, ……, 2) (a total of 4 + 3 + 3 = 10 multiplications of 2), that is to decompose this 10D tensor into:
[0247] ;
[0248] The left side of the equation is the element of the tensor operator with the corresponding index , and the right side of the equation is the element of the tensor with the corresponding index , and the two types of indices satisfy the following formula:
[0249] ; ; ;
[0250] ; ; ;
[0251] ; ; .
[0252] After that, perform tensor decomposition. Perform tensor train decomposition on the quantized tensor to obtain:
[0253] ;
[0254] Among them, is a small-scale 3D tensor operator with a scale of , because obtained after decomposition is a very small number. In addition , , are boundary conditions. Similarly, the vector b also performs the same operation as the vector a, and the tensor decomposition of the vector b corresponding to the 3D tensor B after padding with 0 can be obtained:
[0255] .
[0256] Step 4: Thus, the vector addition of vector a and vector b can be performed through the tensor addition of the tensor operator and the tensor operator To obtain an approximate result, the specific calculation is as follows:
[0257] ;
[0258] ;
[0259] Among them, , ; , ; , ;
[0260] ;
[0261] Thus, the calculation result of tensor addition is obtained. Through this method of operation, the original vector addition operation requires computational complexity, but tensor addition does not require calculation. It realizes the addition operation through element rearrangement, thus reducing the computational complexity.
[0262] For the symbol description in the tensor operator in this example, please refer to Type 1.
[0263] Through the above six types of calculation improvements, not only can the storage complexity of the parameters required for storing tensor calculations in the model calculation task be reduced, but also the computational complexity in the model calculation process can be reduced. In addition, in Figure 1 the calculation process of the Transformer model, the calculation data of some steps comes from the calculation results of the previous steps. Then, in actual calculation, the previous calculation results are stored in the video memory, and then the data is read from it for calculation in the subsequent calculation. However, this reading method can be further optimized because the data storage volume after tensor decomposition is very small, and the memory space it occupies is very small, so it can be stored in the cache, and the data is read from the cache in the subsequent calculation, so that the data can be obtained in a high-bandwidth reading method, thereby further accelerating the inference speed.
[0264] Therefore, by using the data processing method provided in the embodiments of the present invention, while occupying less storage space, the computational complexity of model calculation can be reduced, and the parameters required for calculation can be quickly read, which can reduce the demand for storage space while improving the model calculation performance.
[0265] Figure 3 FIG. is a schematic structural diagram of another data processing system provided in the embodiments of the present invention.
[0266] According to the data processing method provided in the embodiments of the present invention, a data processing system can be constructed as Figure 3As shown, in addition to including a model calculation module, it also includes a controller for controlling input and output and selecting calculation methods, and different processing modules are designed for different types of tensor calculations. For the six matrix calculations introduced in the embodiments of the present invention, a matrix addition processing module, a matrix multiplication processing module, a vector-matrix multiplication processing module, a vector addition processing module, and a scalar-matrix multiplication processing module can be designed respectively. In addition, it also includes a tensor decomposition module and a rearrangement element module required for processing a tensor operator into a third tensor operator.
[0267] The above embodiments introduce a data processing method applicable to models of any type. In addition to this, the embodiments of the present invention also provide an image processing method adapted to image processing tasks.
[0268] Figure 4 It is a flowchart of an image processing method provided by an embodiment of the present invention.
[0269] As Figure 4 shown, the image processing method provided by the embodiments of the present invention includes:
[0270] S401: Receive an input image;
[0271] S402: After vectorizing the input image using an image processing model, perform tensor calculations based on the obtained sequence data;
[0272] S403: During the process of performing tensor calculations, expand the scale of the tensor operator to be calculated so that the number of elements in each dimension is the product of z positive integers to obtain a first tensor operator; perform a rearrangement element process on the first tensor operator to increase the dimension of the first tensor operator while reducing the number of elements in each dimension of the first tensor operator to obtain a second tensor operator, and the dimension increase amount of each dimension of the second tensor operator corresponding to the first tensor operator is the same; perform a tensor train decomposition process on the second tensor operator to obtain a third tensor operator; perform tensor calculations using the third tensor operator;
[0273] S404: Output the image processing result.
[0274] Among them, z is a positive integer greater than 1.
[0275] In a specific implementation, the image processing model can be a Transformer model. Then, for the step of vectorizing the input image using the image processing model and performing tensor calculations based on the obtained sequence data in S402, reference can be made to the introduction in the above embodiments of the present invention.
[0276] For the specific implementation manner of S403, reference can be made to the description of the specific implementation manner of S103.
[0277] For S404, by adopting the processing method provided in the embodiments of the present invention during the process of the image processing model processing the input image when tensor calculation is involved, the image processing result of the input image can be output faster, and the storage space requirement for the parameters required for image processing is smaller.
[0278] Based on the above embodiments, in S403, processing the tensor operator into a third tensor operator may include: determining a conversion coefficient for processing the tensor operator into a third tensor operator according to the model calculation accuracy requirement parameters of the image processing model.
[0279] Among them, determining a conversion coefficient for processing the tensor operator into a third tensor operator according to the model calculation accuracy requirement parameters of the image processing model may include: determining the rank number of the third tensor operator according to the model calculation accuracy requirement parameters.
[0280] Among them, processing the tensor operator into a third tensor operator may include: determining a conversion coefficient for processing the tensor operator into a third tensor operator according to the model calculation accuracy requirement parameters of the image processing model and the size of the calculation parameter storage space allocated for tensor calculation.
[0281] In some optional implementation manners of the embodiments of the present invention, determining a conversion coefficient for processing the tensor operator into a third tensor operator according to the model calculation accuracy requirement parameters of the image processing model and the size of the calculation parameter storage space allocated for tensor calculation may include: determining the minimum rank number allowed for the third tensor operator according to the model calculation accuracy requirement parameters; generating the third tensor operator according to the minimum rank number; determining the number of parameters required to perform tensor calculation using the third tensor operator, and if it is not enough to occupy the calculation parameter storage space allocated for tensor calculation, increasing the rank number of the third tensor operator and regenerating the third tensor operator, and re-determining the number of parameters required to perform tensor calculation using the third tensor operator.
[0282] In some other alternative embodiments of the embodiments of the present invention, to determine the conversion coefficient for processing a tensor operator into a third tensor operator according to the model calculation accuracy requirement parameter of the image processing model and the size of the calculation parameter storage space allocated to tensor calculation, it may include: determining the minimum rank number allowed for the third tensor operator according to the model calculation accuracy requirement parameter; generating the third tensor operator according to the minimum rank number; expanding the scale of the tensor operator and performing element rearrangement processing to obtain a second tensor operator, where the second tensor operator rearranges the elements of each dimension of the corresponding tensor operator into k dimensions, and the number of elements in each dimension of the k-dimensional elements can be expressed in exponential form, and both the base and the exponent are positive integers; determining the number of parameters of a single third tensor operator according to the product of the bases at the corresponding positions of each dimension in the second tensor operator and the rank number of the third tensor operator, where if the base can be square-rooted to obtain a non-1 positive integer, then perform square-root calculation on the base and then perform product calculation; if the maximum value of the number of parameters of a group of third tensor operators is not sufficient to occupy the calculation parameter storage space allocated to a single group of tensor calculations, then change the scale expansion scale of the tensor operator or the element rearrangement method of the first tensor operator, regenerate the second tensor operator, and then return to determine the number of parameters of a single third tensor operator according to the product of the bases at the corresponding positions of each dimension in the second tensor operator and the rank number of the third tensor operator.
[0283] In some other alternative embodiments of the embodiments of the present invention, to determine the conversion coefficient for processing a tensor operator into a third tensor operator according to the model calculation accuracy requirement parameter of the image processing model and the size of the calculation parameter storage space allocated to tensor calculation, it may include: expanding the scale of the tensor operator and performing element rearrangement processing to obtain a second tensor operator, where the second tensor operator rearranges the elements of each dimension of the corresponding tensor operator into k dimensions, and the number of elements in each dimension of the k-dimensional elements can be expressed in exponential form, and both the base and the exponent are positive integers; determining the number of parameters of a single third tensor operator according to the product of the bases at the corresponding positions of each dimension in the second tensor operator and the rank number of the third tensor operator, where if the base can be square-rooted to obtain a non-1 positive integer, then perform square-root calculation on the base and then perform product calculation; determining the number of third tensor operators according to the sum of the exponents of a single dimension in the second tensor operator; determining the number of parameters of tensor calculation according to the number of parameters of a single third tensor operator and the number of third tensor operators; if the number of parameters of tensor calculation is not sufficient to occupy the calculation parameter storage space allocated to tensor calculation, then change the scale expansion scale of the tensor operator or the element rearrangement method of the first tensor operator, regenerate the second tensor operator, and then return to determine the number of parameters of a single third tensor operator according to the product of the bases at the corresponding positions of each dimension in the second tensor operator and the rank number of the third tensor operator and determine the number of third tensor operators according to the sum of the exponents of a single dimension in the second tensor operator.
[0284] In some other alternative embodiments of the embodiments of the present invention, processing the tensor operator into a third tensor operator may include: determining the rank number of the third tensor operator according to the model calculation accuracy requirement parameter; expanding the scale of the tensor operator and performing element rearrangement processing to obtain a second tensor operator, where the second tensor operator rearranges the elements of each dimension of the corresponding tensor operator into k dimensions, and the number of elements in each dimension of the k - dimensional elements can be expressed in exponential form, and both the base and the exponent are positive integers; determining the number of parameters of a single third tensor operator according to the product of the bases corresponding to each dimension in the second tensor operator and the rank number of the third tensor operator, where if the base can be square - rooted to obtain a non - 1 positive integer, then the base is square - rooted before performing the product calculation; determining the number of third tensor operators according to the sum of the exponents of a single dimension in the second tensor operator; determining the number of parameters of the tensor calculation according to the rank number of the third tensor operator, the number of parameters of a single third tensor operator, and the number of third tensor operators; taking the minimization of the number of parameters of the tensor calculation as the optimization goal, and solving to obtain the way to expand the tensor operator into the first tensor operator and the way to rearrange the second tensor operator into the third tensor operator.
[0285] In the data processing method provided by the embodiments of the present invention, the tensor calculation may be a multiplication calculation of two tensor operators; then, expanding the scale of the tensor operator in S403 may include: expanding both of the two tensor operators to the number of elements in each dimension being the product of z positive integers, and the number of elements in each dimension satisfies the requirement for the multiplication calculation of the two tensor operators.
[0286] Among them, if both of the two tensor operators for the multiplication calculation are two - dimensional tensor operators, then the number of columns of the multiplicand tensor operator in the two tensor operators is the same as the number of rows of the multiplier tensor operator in the two tensor operators.
[0287] In some alternative embodiments of the embodiments of the present invention, using the third tensor operator to perform tensor calculation may include: after storing the third tensor operator in the cache, reading the third tensor operator from the cache to perform tensor calculation.
[0288] It should be noted that in the embodiments of each data processing method and the embodiments of the image processing method of the present invention, some of the steps or features may be ignored or not executed. For the convenience of describing the divided hardware or software functional modules, it is not the only implementation form for implementing the data processing method and the image processing method provided by the embodiments of the present invention.
[0289] The above details the various embodiments corresponding to the data processing method and the image processing method. On this basis, the present invention also discloses a data processing device, an image processing device, a computing device, a non - volatile storage medium, and a computer program product corresponding to the above methods.
[0290] The data processing device provided by an embodiment of the present invention may include:
[0291] A data reading unit, configured to determine information of a storage device and information of a computing device according to a model calculation task, and read data to be processed corresponding to the model calculation task from the storage device;
[0292] A first calculation unit, configured to perform vectorization processing on the data to be processed by using a target model, and then perform tensor calculation according to the obtained sequence data; during the tensor calculation process, expand the scale of a tensor operator to be calculated to that the number of elements in each dimension is a product of z positive integers, to obtain a first tensor operator; perform element rearrangement processing on the first tensor operator, so as to increase the dimension of the first tensor operator while reducing the number of elements in each dimension of the first tensor operator, to obtain a second tensor operator, and the number of dimension increases of each dimension of the second tensor operator corresponding to the first tensor operator is the same; perform tensor train decomposition processing on the second tensor operator, to obtain a third tensor operator; write the third tensor operator into a calculation parameter storage space of the computing device, so as to perform the tensor calculation by using the third tensor operator;
[0293] A first output unit, configured to output a processing result corresponding to the data to be processed.
[0294] Where z is a positive integer greater than 1.
[0295] The image processing device provided by an embodiment of the present invention may include:
[0296] A receiving unit, configured to receive an input image;
[0297] A second calculation unit, configured to perform vectorization processing on the input image by using an image processing model, and then perform tensor calculation according to the obtained sequence data; during the tensor calculation process, expand the scale of a tensor operator to be calculated to that the number of elements in each dimension is a product of z positive integers, to obtain a first tensor operator; perform element rearrangement processing on the first tensor operator, so as to increase the dimension of the first tensor operator while reducing the number of elements in each dimension of the first tensor operator, to obtain a second tensor operator, and the number of dimension increases of each dimension of the second tensor operator corresponding to the first tensor operator is the same; perform tensor train decomposition processing on the second tensor operator, to obtain a third tensor operator; perform the tensor calculation by using the third tensor operator;
[0298] A second output unit, configured to output an image processing result.
[0299] Where z is a positive integer greater than 1.
[0300] It should be noted that in each embodiment of the data processing device and the image processing device provided by the embodiments of the present invention, the division of units is only a logical function division, and other division methods can be adopted. The connection methods between different units can adopt electrical, mechanical or other connection methods. The separated units can be located at the same physical location or distributed on multiple network nodes. Each unit can be implemented in the form of hardware or in the form of a software functional unit. That is, part or all of the units provided by the embodiments of the present invention can be selected according to actual needs and the corresponding connection methods or integration methods can be adopted to achieve the purpose of the solution of the embodiments of the present invention.
[0301] Since the embodiments of the device part correspond to the embodiments of the method part, please refer to the description of the embodiments of the method part for the embodiments of the device part, and will not be elaborated here.
[0302] Figure 5 It is a schematic structural diagram of a computing device provided by an embodiment of the present invention.
[0303] As Figure 5 shown, the computing device provided by the embodiment of the present invention includes: a memory 510 for storing a computer program 511; a processor 520 for executing the computer program 511, and when the computer program 511 is executed by the processor 520, it implements the steps of the computing method provided by any one of the above embodiments.
[0304] Among them, the processor 520 may include one or more processing cores, such as a 3-core processor, an 8-core processor, etc. The processor 520 can be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), and programmable logic array (PLA). The processor 520 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the central processing unit (CPU); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 520 may be integrated with a graphics processing unit (GPU), and the graphics processing unit is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 520 may further include an artificial intelligence (AI) processor, and the artificial intelligence processor is used to process computing operations related to machine learning.
[0305] The memory 510 may include one or more non-volatile storage media, which may be non-transitory. The memory 510 may also include high-speed random access memory, as well as non-volatile memory, such as one or more disk storage devices and flash storage devices. In this embodiment, the memory 510 is at least used to store the following computer program 511. After the computer program 511 is loaded and executed by the processor 520, it can implement the relevant steps in the calculation methods disclosed in any of the foregoing embodiments. In addition, the resources stored in the memory 510 may also include an operating system 512, data 513, etc., and the storage method may be transient storage or permanent storage. Among them, the operating system 512 may be Windows or other types of operating systems. The data 513 may include, but is not limited to, the data involved in the above methods.
[0306] In some embodiments, the computing device may further include a display screen 530, a power supply 540, a communication interface 550, an input / output interface 560, a sensor 570, and a communication bus 580.
[0307] Those skilled in the art can understand that Figure 5 the structure shown in
[0308] does not constitute a limitation on the computing device, and may include more or fewer components than those shown in the figure.
[0309] The computing device provided by the embodiment of the present invention includes a memory and a processor. When the processor executes the program stored in the memory, it can implement the steps of the calculation method provided in the above embodiment, and the effect is the same.
[0310] The non-volatile storage medium may include: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0311] For the introduction of the non-volatile storage medium provided by the embodiment of the present invention, please refer to the above method embodiment, and the effect thereof is the same as that of the data processing method and image processing method provided by the embodiment of the present invention. The present invention will not elaborate herein.
[0312] The embodiment of the present invention provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the steps of the data processing method or the image processing method provided in any of the foregoing embodiments.
[0313] For the introduction of the computer program product provided by the embodiments of the present invention, please refer to the above method embodiments, and the effects thereof are the same as those of the data processing method and the image processing method provided by the embodiments of the present invention, which will not be elaborated herein.
[0314] The above has introduced in detail an image processing method, a data processing method, a device, a medium and a product provided by the present invention. The embodiments in the specification are described in a progressive manner, and the key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the devices, computing devices, non-volatile storage media and computer program products disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part. It should be noted that those of ordinary skill in the art can make several improvements and modifications to the present invention without departing from the principle of the present invention, and these improvements and modifications also fall within the protection scope of the present invention.
[0315] It should also be noted that in this specification, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover a non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the existence of additional identical elements in the process, method, article or device comprising such element.
Claims
1. An image processing method, characterized in that: include: receiving an input image; After vectorizing the input image using an image processing model, tensor calculation is performed based on the obtained sequence data; In the process of performing tensor calculation, the tensor operator to be used for the tensor calculation is scaled up to the point where the number of elements in each dimension is the product of z positive integers, thereby obtaining a first tensor operator; the first tensor operator is subjected to element rearrangement processing to increase the dimension of the first tensor operator while reducing the number of elements in each dimension of the first tensor operator, thereby obtaining a second tensor operator, and the second tensor operator has the same number of increased dimensions as the first tensor operator in each dimension; the second tensor operator is subjected to tensor column decomposition processing to obtain a third tensor operator; and the tensor calculation is performed using the third tensor operator; Output image processing results; Wherein, z is a positive integer greater than 1; Processing the tensor operator into the third tensor operator includes: Determining the rank of the third tensor operator according to the model calculation accuracy requirement parameter; The tensor operator is expanded in scale and the elements are rearranged to obtain the second tensor operator, wherein the second tensor operator rearranges the elements of each dimension corresponding to the tensor operator into k dimensions, and the number of elements of each dimension in the k-dimensional elements can be expressed in exponential form, and the base and the exponent are both positive integers; Determine the number of parameters of a single third tensor operator according to the product of the bases of the corresponding positions of each dimension in the second tensor operator and the rank of the third tensor operator, wherein if the base can be square-rooted to obtain a non-1 positive integer, square-root the base before performing the product calculation; Determining the number of the third tensor operators according to the sum of the exponents of the single dimensions in the second tensor operators; Determining the number of parameters of the tensor calculation according to the rank of the third tensor operator, the number of parameters of a single third tensor operator, and the number of the third tensor operators; Taking minimizing the number of parameters of the tensor calculation as an optimization goal, a method of expanding the tensor operator into the first tensor operator and a method of rearranging the second tensor operator into the third tensor operator are solved.
2. The image processing method according to claim 1, characterized in that: Processing the tensor operator as the third tensor operator can be replaced by: Determining a minimum rank number allowed by the third tensor operator according to the model calculation accuracy requirement parameter; generating the third tensor operator according to the minimum rank number; The tensor operator is expanded in scale and the elements are rearranged to obtain the second tensor operator, wherein the second tensor operator rearranges the elements of each dimension corresponding to the tensor operator into k dimensions, and the number of elements of each dimension in the k-dimensional elements can be expressed in exponential form, and the base and the exponent are both positive integers; Determine the number of parameters of a single third tensor operator according to the product of the bases of the corresponding positions of each dimension in the second tensor operator and the rank of the third tensor operator, wherein if the base can be square-rooted to obtain a non-1 positive integer, square-root the base before performing the product calculation; If the maximum value of the number of parameters of a group of the third tensor operators is not sufficient to occupy the calculation parameter storage space allocated to the single group of calculations of the tensor calculations, then change the scale of the tensor operator expansion or the way of rearranging elements of the first tensor operator, regenerate the second tensor operator, and return to determine the number of parameters of a single third tensor operator based on the product of the bases of the corresponding positions of each dimension in the second tensor operator and the rank of the third tensor operator.
3. The image processing method according to claim 1, characterized in that: Processing the tensor operator as the third tensor operator can be replaced by: The tensor operator is expanded in scale and the elements are rearranged to obtain the second tensor operator, wherein the second tensor operator rearranges the elements of each dimension corresponding to the tensor operator into k dimensions, and the number of elements of each dimension in the k-dimensional elements can be expressed in exponential form, and the base and the exponent are both positive integers; Determine the number of parameters of a single third tensor operator according to the product of the bases of the corresponding positions of each dimension in the second tensor operator and the rank of the third tensor operator, wherein if the base can be square-rooted to obtain a non-1 positive integer, square-root the base before performing the product calculation; Determining the number of the third tensor operators according to the sum of the exponents of the single dimensions in the second tensor operators; Determining the number of parameters of the tensor calculation according to the number of parameters of a single third tensor operator and the number of the third tensor operators; If the number of parameters of the tensor calculation is insufficient to occupy the calculation parameter storage space allocated to the tensor calculation, then change the scale of the tensor operator expansion or the way of rearranging elements of the first tensor operator, regenerate the second tensor operator, and return to determine the number of parameters of a single third tensor operator based on the product of the bases of the corresponding positions of each dimension in the second tensor operator and the rank of the third tensor operator, and determine the number of the third tensor operators based on the sum of the exponents of the single dimensions in the second tensor operator.
4. The image processing method according to claim 1, characterized in that: The tensor calculation is a multiplication calculation of two tensor operators; The tensor operator is scaled up, including: The two tensor operators are expanded so that the number of elements in each dimension is the product of z positive integers and the number of elements in each dimension meets the requirements for multiplication calculation of the two tensor operators.
5. The image processing method according to claim 4, characterized in that: The two tensor operators performing multiplication calculation are both two-dimensional tensor operators, and the number of columns of the multiplicand tensor operators in the two tensor operators is the same as the number of rows of the multiplier tensor operators in the two tensor operators.
6. The image processing method according to claim 1, characterized in that: Performing the tensor calculation using the third tensor operator includes: After storing the third tensor operator in a cache, the third tensor operator is read from the cache to perform the tensor calculation.
7. A data processing method, characterized in that: include: After determining the information of the storage device and the information of the computing device according to the model computing task, reading the to-be-processed data corresponding to the model computing task from the storage device; After vectorizing the data to be processed using the target model, tensor calculation is performed based on the obtained sequence data; In the process of performing tensor calculation, the tensor operator to be used for the tensor calculation is scaled up to the point where the number of elements in each dimension is the product of z positive integers, so as to obtain a first tensor operator; the first tensor operator is subjected to element rearrangement processing to increase the dimension of the first tensor operator while reducing the number of elements in each dimension of the first tensor operator, so as to obtain a second tensor operator, and the second tensor operator has the same number of increased dimensions in each dimension as the first tensor operator; the second tensor operator is subjected to tensor column decomposition processing to obtain a third tensor operator; the third tensor operator is written into the calculation parameter storage space of the computing device, so as to use the third tensor operator to perform the tensor calculation; Outputting a processing result corresponding to the data to be processed; Wherein, z is a positive integer greater than 1; Processing the tensor operator into the third tensor operator includes: Determining the rank of the third tensor operator according to the model calculation accuracy requirement parameter; The tensor operator is expanded in scale and the elements are rearranged to obtain the second tensor operator, wherein the second tensor operator rearranges the elements of each dimension corresponding to the tensor operator into k dimensions, and the number of elements of each dimension in the k-dimensional elements can be expressed in exponential form, and the base and the exponent are both positive integers; Determine the number of parameters of a single third tensor operator according to the product of the bases of the corresponding positions of each dimension in the second tensor operator and the rank of the third tensor operator, wherein if the base can be square-rooted to obtain a non-1 positive integer, square-root the base before performing the product calculation; Determining the number of the third tensor operators according to the sum of the exponents of the single dimensions in the second tensor operators; Determining the number of parameters of the tensor calculation according to the rank of the third tensor operator, the number of parameters of a single third tensor operator, and the number of the third tensor operators; Taking minimizing the number of parameters of the tensor calculation as an optimization goal, a method of expanding the tensor operator into the first tensor operator and a method of rearranging the second tensor operator into the third tensor operator are solved.
8. A computing device, characterized in that include: Memory for storing computer programs; A processor is used to execute the computer program, and when the computer program is executed by the processor, the steps of the image processing method according to any one of claims 1 to 6 or the steps of the data processing method according to claim 7 are implemented.
9. A non-volatile storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the image processing method according to any one of claims 1 to 6 or the steps of the data processing method according to claim 7 are implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the image processing method according to any one of claims 1 to 6 or the steps of the data processing method according to claim 7 are implemented.
Citation Information
Patent Citations
Image compression processing method, system and device and medium
CN117319655A