Tensor structure information analysis method and system
Through the tensor structure information analysis method based on program semantic analysis, the problem of difficult to capture and utilize tensor structure information in custom neural network programs in the prior art is solved, and efficient computing resource utilization and model performance optimization are achieved.
Patent Information
- Application Number
- CN202510083524.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-01-20
AI Technical Summary
The prior art is difficult to effectively capture and utilize tensor structure information in custom neural network programs, resulting in waste of computing resources and delay in model inference training.
A tensor structure information analysis method based on program semantic analysis is proposed. Through index description list calculation, tensor structure information semantic analysis and index list optimization, neural network programs are deeply analyzed and optimized to reduce computational redundancy.
It realizes comprehensive perception and efficient processing of tensor structure information in neural network programs, reduces unnecessary computing operations, and improves computing resource utilization and overall operating performance.
Smart Images

Figure CN120066572A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of artificial intelligence and program analysis, and particularly to an analysis method and system for neural network tensor structure information. Background Art
[0002] With the remarkable results achieved by neural network models in various tasks, researchers have been constantly working on developing more complex and diverse models. To fully utilize the computational potential of these complex models, existing neural network frameworks (such as PyTorch) generally adopt many highly optimized operators. However, as the complexity and flexibility of the models increase, the phenomenon of wasted computational resources has become more and more significant. For example, many masking operations lead to unnecessary computational redundancy, reducing the efficiency of model inference and training. Therefore, how to effectively perceive and optimize the computational overhead brought by these semantic operations has become a key issue in neural network model optimization.
[0003] Existing neural network frameworks fail to deeply analyze and optimize various neural network application programs written by users, especially programs containing complex semantics. On the one hand, the use of masks increases the diversity of neural network semantics and the robustness of neural network performance, enabling the neural network to more flexibly adapt to the needs of different tasks. On the other hand, it objectively causes a large amount of computational redundancy, especially in practical applications. When masks are widely used, it may lead to the system processing a lot of invalid data during the calculation process, increasing the unnecessary computational burden. If developers do not use the specific coding paradigms provided by the framework, especially when writing custom operations or optimization strategies, it is usually difficult to avoid performing some unnecessary operations, which results in a significant decrease in computational efficiency.
[0004] In addition, these redundant calculations not only waste computational resources but also may cause a significant delay in the model inference and training processes, especially in resource-constrained embedded devices or edge computing environments, seriously affecting the actual deployment and application of neural networks.
[0005] Existing neural network frameworks usually have difficulty in deeply analyzing and optimizing when dealing with complex neural network programs written by users, especially programs containing complex semantics. Although these frameworks enhance the flexibility and adaptability of neural networks by supporting diverse syntax and operations, such as increasing semantic diversity and improving the robustness of neural networks through masking operations, these flexible operations also pose significant challenges to computational efficiency.
[0006] Although masking operations can enable neural networks to better handle different task requirements, they often lead to a large amount of computational redundancy in practical applications. The redundant technical problems existing in the prior art are mainly reflected in:
[0007] 1. Calculate invalid data: When masks are widely used, the system needs to process a large amount of invalid data during the calculation process. These data do not generate actual value during the calculation, but consume precious computing resources and increase the computational burden.
[0008] 2. Decrease in computational efficiency: If developers do not follow the specific coding paradigms provided by the framework, especially when writing custom operations or optimization strategies, it often leads to the execution of many unnecessary operations. This directly results in a significant decrease in computational efficiency and affects the inference and training processes of the model.
[0009] These computational redundancies not only waste computing resources but also may cause significant inference and training delays. This problem is particularly prominent in resource-constrained environments such as embedded devices or edge computing devices, seriously affecting the actual deployment and application effects of neural network models.
[0010] On this basis, several frameworks have proposed methods to optimize neural network models with sparse structures. However, the initial setting of the information structure of tensors, especially the sparse information structure, highly depends on manual experience. Frameworks such as SparTA need to manually set the sparse structure information of input tensors and weights before running. These prerequisite settings for optimization bring an additional burden to neural network programming.
[0011] Therefore, to solve the technical problem that the above optimization brings an additional burden to neural network programming, to address the challenge in the prior art of being difficult to effectively capture and utilize the tensor structure information in custom neural network programs, and to solve the technical problems of the difficulty in capturing tensor structure information and the difficulty in describing tensor structure information, it is urgently necessary to propose a method and system that can perceive tensor structure information through static analysis and can achieve tensor structure information analysis based on program semantic analysis; it is urgently necessary to have a system that can deeply analyze and optimize neural network programs while maintaining the diversity and flexibility of neural network semantic expressions, reduce computational redundancy, and improve overall efficiency. Summary of the Invention
[0012] To address the challenge in the prior art of being difficult to effectively capture and utilize the tensor structure information in custom neural network programs, the present invention proposes a tensor structure information analysis system based on program semantic analysis.
[0013] In a first aspect, embodiments of the present application provide an analysis method for tensor structure information. The method includes:
[0014] Index description list calculation step: For the input neural network application program, based on program semantic analysis, analyze the structure information of each tensor in the program, create a description tensor of the structure information of the tensor, convert the description tensor into a nested list, and construct a tree structure based on the nested list;
[0015] Steps for semantic analysis of tensor structure information: Traverse the tree structure, record all node paths, form a list of constant tensor index descriptions, and based on the list of constant tensor index descriptions, obtain a list of index descriptions of non-constant tensors through program semantic analysis;
[0016] Steps for optimizing the index list: Adopt a structure index information optimization mechanism to merge and optimize the obtained list of constant tensor and non-constant tensor index descriptions.
[0017] In a specific embodiment of the present invention, before the above-mentioned index description list calculation steps, the following steps are further included:
[0018] Steps for parsing program representation: Parse the source code of the program to obtain the specified initial values of all constant tensors, and create a structure information description tensor corresponding to the constant tensor, where the description tensor has the same shape as the corresponding original tensor, and the value at each position of the description tensor is used to label the specific value or precision information of the corresponding original tensor element;
[0019] Steps for initializing the structure information description of the tensor: Determine which description tensors corresponding to specific values are to be obtained according to the application program settings. For positions in the original tensor with a value of 0, the value at the corresponding position of the description tensor is set to 0; for positions with a preset specific value, the value at the corresponding position of the description tensor is set to the preset specific value.
[0020] In a specific embodiment of the present invention, the above-mentioned index description list calculation steps include:
[0021] Steps for constructing the tree structure: Convert the description tensor into a nested list and construct a tree structure, where each layer of the tree corresponds to a dimension of the tensor, and the nodes at any layer represent the index information of the corresponding dimension;
[0022] Steps for list contraction: Start performing contraction operations from the innermost dimension of the tree structure. The contraction operation is: If all elements in the nested list are preset specific values, replace the nested list with the preset specific value;
[0023] Steps for iterative processing: Iterate from the inner layer to the outer layer dimensions of the tree structure, perform contraction operations on each layer. If the elements of a certain list are not all preset specific values, do not perform the contraction operation;
[0024] Steps for index recording: By traversing the tree, record all node paths that satisfy the condition that all subtrees are preset specific values to form a list of index descriptions of the preset specific value.
[0025] In a specific embodiment of the present invention, the above-mentioned steps for semantic analysis of tensor structure information further include:
[0026] Analysis steps of assignment operation: When there is an assignment statement in the program and the constant at the i-th row and j-th column is set to 0, add an item [i, j] to the index description list with a value of 0, where i and j represent the number of rows and columns of the matrix respectively;
[0027] Recognition steps of specific structure tensor: For a two-dimensional tensor created by the program with a diagonal of a specific value scale and both the length and width of the two-dimensional tensor being l, add the arrays (0, 0), (1, 1), …, (l, l) to the index description list corresponding to scale.
[0028] In a specific embodiment of the present invention, the above-mentioned semantic analysis steps of tensor structure information further include:
[0029] Recognition and setting steps of regional unified value: The recognition and setting of the regional unified value are achieved by recognizing the regions with unified values in the tensor and setting a specific value for the regions with unified values;
[0030] Mining steps of operation characteristics of specific patterns: For the index positions with a specific interval pattern, establish a corresponding calculation template, define the step size and offset, and perform calculations on the positions separated by the step size and offset;
[0031] Steps of calculating the proportion of effective operations: The proportion of effective operations is the ratio of the non-zero values in each innermost dimension to the dimension length. Different operation modes are corresponding to different proportions of effective operations, and the operation modes are extracted through semantic analysis or set by the program.
[0032] In a specific embodiment of the present invention, the above-mentioned index list optimization steps further include:
[0033] Node merging steps: Starting from the deepest leaf nodes, traverse the tree upwards. If all the sub-node indices of a certain node are in the record and cover all positions of the corresponding dimension, then merge the sub-nodes, shrink the subtree into one node, and make the node a leaf node;
[0034] Iterative simplification steps: Repeatedly execute the node merging steps, merge layer by layer upwards until no further simplification is possible, and obtain the most simplified index description list.
[0035] In a second aspect, an analysis system for tensor structure information is provided in an embodiment of the present application. Using the analysis method of tensor structure information as described above, the system includes:
[0036] Index description list calculation module: For the input neural network application program, based on program semantic analysis, analyze the structure information of each tensor in the program, create a structure information description tensor corresponding to the tensor, convert the description tensor into a nested list, and construct a tree structure based on the nested list;
[0037] Tensor Structure Information Semantic Analysis Module: It is used to traverse the tree structure, record all node paths, form a list of constant tensor index descriptions, and based on the list of constant tensor index descriptions, obtain a list of index descriptions of non-constant tensors through program semantic analysis;
[0038] Index List Optimization Module: It is used to adopt a structure index information optimization mechanism to merge and optimize the obtained lists of constant tensor and non-constant tensor index descriptions.
[0039] In a specific embodiment of the present invention, before the above index description list calculation module, it further includes:
[0040] Program Representation Parsing Module: It is used to parse the source code of the program to obtain the specified initial values of all constant tensors, and create a structure information description tensor corresponding to the constant tensor. Among them, the description tensor has the same shape as the corresponding original tensor, and the value at each position of the description tensor is used to label the specific value or precision information of the corresponding original tensor element;
[0041] Initialization Tensor Structure Information Description Module: It is used to determine which description tensors corresponding to specific values are obtained according to the application program settings. For the positions in the original tensor where the value is 0, the value at the corresponding position of the description tensor is set to 0; for the positions where the value is a preset specific value, the value at the corresponding position of the description tensor is set to the preset specific value.
[0042] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the analysis method of the tensor structure information are implemented.
[0043] In a fourth aspect, an embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the analysis method of the tensor structure information are implemented.
[0044] Compared with the related prior art, it has the following outstanding beneficial effects:
[0045] 1) The method of the present invention designs a structure index information optimization mechanism. By merging and simplifying the index list, redundant records are reduced, and the accuracy and storage efficiency of index descriptions are improved;
[0046] 2) The structure index information optimization mechanism proposed in the method of the present invention can not only reduce the redundant storage of structure information, but also significantly improve the information reading and processing speed during calculation. In a complex neural network program, through the optimization processing of structural features, the system can more efficiently perceive and process tensor structure information, ensuring that the data processing and storage performance during the program execution reaches the optimal. Description of the Drawings
[0047] The accompanying drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:
[0048] Figure 1 It is a schematic diagram of the analysis method for the tensor structure information of the present invention;
[0049] Figure 2 It is a schematic diagram of the analysis method for the tensor structure information in the specific embodiment of the present invention;
[0050] Figure 3 It is a schematic diagram of the analysis system for the tensor structure information in the embodiment of the present invention;
[0051] Figure 4 It is a schematic diagram of the computer hardware of the present invention. Detailed implementation manners
[0052] It should be noted that the processor described in the present invention is the control center of the electronic device, which can be a single processor or a collective term for multiple processing elements. For example, it can be one or more central processing units (CPUs), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present invention. For example: one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs).
[0053] Optionally, the processor can execute various functions of the electronic device by running or executing software programs stored in the memory and calling data stored in the memory.
[0054] In a specific implementation, as an embodiment, the processor may include one or more CPUs. Each of these processors can be a single core processor or a multi-core processor. Here, the processor can refer to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions). The electronic device can include: servers, desktop computers, laptop computers, smart phones, tablet computers, embedded computers, etc., where the embedded computer includes vehicles and robots, etc.
[0055] The memory is used to store the software program for implementing the solution of the present invention and is controlled by a processor for execution. The specific implementation manner can refer to the above method embodiments and will not be elaborated herein.
[0056] It should be noted that the structure of the electronic device shown in the drawings of the present invention does not constitute a limitation thereto. The actual knowledge structure recognition device may include more or fewer components than shown in the drawings, or combine certain components, or have different component arrangements.
[0057] The above embodiments can be implemented in whole or in part by software, hardware (such as circuits), firmware, or any other arbitrary combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wired (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that contains one or more collections of available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.
[0058] It should also be understood that the term "and / or" in this article is merely a description of the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, which can be specifically understood with reference to the context.
[0059] In the present invention, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following items (pieces)" or its similar expressions refer to any combination of these items, including any combination of single items (pieces) or plural items (pieces). For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.
[0060] It should also be understood that in various embodiments of the present invention, the magnitudes of the sequence numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0061] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0062] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0063] In addition, in each embodiment of the present invention, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.
[0064] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks, etc., which can store program codes.
[0065] To make the above features and effects of the present invention more clearly and understandably described, specific embodiments are given below and will be described in detail in conjunction with the accompanying drawings of the specification. This specification discloses one or more embodiments incorporating the features of the present invention. The disclosed embodiments are only for illustrative purposes. The protection scope of the present invention is not limited to the disclosed embodiments, and the present invention is defined by the appended claims.
[0066] The following is a system embodiment corresponding to the above method embodiment, and this embodiment can be implemented in cooperation with the above embodiment. The relevant technical details mentioned in the above embodiment are still valid in this embodiment. To avoid repetition, they will not be elaborated here. Correspondingly, the relevant technical details mentioned in this embodiment can also be applied to the above embodiment.
[0067] Aiming at the problem of capturing tensor information of user-defined neural network programs, the present invention aims to propose a tensor structure information analysis system based on program semantic analysis. This system aims to solve the challenge in the prior art of effectively capturing and utilizing tensor structure information in user-defined neural network programs. The tensor structure information in neural network programs is usually hidden in complex program semantics, and this information is crucial for optimizing computational efficiency. To better optimize the execution efficiency of user-defined neural network programs, the system needs to deeply analyze program semantics, identify various structured information of tensors, and perform corresponding optimization processing in combination with the specific semantic attributes of operators.
[0068] The system provides as much useful information as possible for the optimization process by perceiving the structural information on the result tensors that can be analyzed under general semantic expressions by users and the specific structures caused by operators with special semantics. By embedding this structural information into the computational graph, the system can not only optimize the execution efficiency of individual operators but also achieve the overall optimization of the computational graph globally. The ability to comprehensively perceive and annotate structural information enables the system to significantly reduce unnecessary computational amounts and improve the execution efficiency of neural network programs in various hardware environments, especially in resource-constrained scenarios.
[0069] In this process, the main technical difficulties faced include:
[0070] 1. Difficulty in capturing tensor structure information: User-defined neural network programs usually have highly complex and dynamic structures, and structural information such as the sparsity of tensors is often hidden in complex program semantics. It is very difficult to directly extract this hidden information from program codes, and it is necessary to deeply analyze the semantics and execution logic of the program to accurately identify the operations of tensors and their impacts on the structure. This involves in-depth analysis of the abstract syntax tree of neural network programs and various intermediate representation forms to capture the structural features of tensors.
[0071] 2. Difficulty in Describing Tensor Structure Information: How to accurately and systematically describe the tensor structure information captured from program semantics is also a major challenge. A general and efficient representation method needs to be designed to convert the captured sparse information into a form that can be used to represent the sparse structure of tensors. This description method must be able to cover various possible structural features so that the subsequent optimization process can effectively utilize this information. At the same time, redundancy should be avoided to improve the efficiency of information processing.
[0072] The following is a detailed description in combination with specific embodiments:
[0073] Embodiment 1
[0074] As Figure 1 shown, Figure 1 is a schematic flowchart of the analysis method for tensor structure information of the present invention. The embodiments of the present application provide an analysis method for tensor structure information, and the method includes:
[0075] Index Description List Calculation Step 101: For the input neural network application program, based on program semantic analysis, analyze the structure information of each tensor in the program, create a description tensor of the structure information of the tensor, convert the description tensor into a nested list, and build a tree structure based on the nested list;
[0076] Tensor Structure Information Semantic Analysis Step 102: Traverse the tree structure, record all node paths, form a constant tensor index description list, and based on the constant tensor index description list, obtain the index description list of non-constant tensors through program semantic analysis;
[0077] Index List Optimization Step 103: Adopt a structure index information optimization mechanism to merge and optimize the obtained constant tensor and non-constant tensor index description lists.
[0078] 1) Capturing Neural Network Tensor Structure Information Based on Program Semantics
[0079] When current neural network systems process tensors on a large scale, there are often a large number of structural features, such as sparsity. In specific operation modes, zero elements or low-impact weights contribute little or even are completely unnecessary to the actual calculation, but traditional calculation methods fail to fully utilize this information, resulting in a waste of computing resources. Existing neural network compilers and optimization systems mainly rely on shallow structural analysis of operators or computational graphs, or highly depend on manual annotation and optimization, and cannot automatically delve into the semantic level of the program for a comprehensive identification of sparse structures. Therefore, how to intelligently obtain and annotate the structure information of tensors in neural network programs has become one of the key issues in optimization.
[0080] To solve this problem, the present invention proposes a tensor structure information capture mechanism based on program semantic analysis. By parsing the semantic information of neural network programs, the system can deeply understand tensor operations and their results, thereby accurately capturing structural features. By analyzing the structural information of each tensor in the program, the system can identify semantics such as all-zero elements and masking operations, and generate corresponding structural information. These structural information can not only provide an optimization basis for subsequent calculations, but also provide complete basic data for the system to perceive inefficient calculation paths.
[0081] Specifically, the present invention automatically analyzes program code to generate an associated matrix and an index description list for describing tensor structural features, and details and annotates specific information in the tensor. The system can identify operators with masking semantics, zero-element regions of tensors, and structural features of temporary tensors generated by code logic in user-defined neural network programs. In addition, it also supports user-defined operation features. By recording this information, it can be further ensured that the program can accurately know the structure of the tensor during execution.
[0082] 2) Index-based description of neural network tensor structural features
[0083] After obtaining tensor structure information, how to efficiently represent and store this information becomes another important technical issue for optimizing system performance. Existing methods usually only record individual structural features of tensors, resulting in high storage complexity of structural information, lack of generalization, and inability to achieve efficient description and calculation. In complex neural networks, as the dimension and tensor scale increase, the processing efficiency of structural information will significantly decrease, and how to simplify and optimize this information becomes a challenge.
[0084] To solve this problem, the present invention designs a structural index information optimization mechanism. By merging and simplifying the index list, redundant records are reduced, and the accuracy and storage efficiency of index descriptions are improved. Specifically, when the system captures the sparse structure of a tensor, it merges the same or similar elements in the structure. For example, if all elements in a certain dimension have the same precision, the system will merge these elements into a unified description, thereby reducing storage space and the indexing difficulty of sparse regions.
[0085] The structural index information optimization mechanism in the present invention can not only reduce the redundant storage of structural information, but also significantly improve the information reading and processing speed during calculation. In complex neural network programs, through the optimization processing of structural features, the system can more efficiently perceive and process tensor structure information, ensuring that the data processing and storage performance during program execution reaches the optimal level.
[0086] In summary, by solving the two major technical problems of obtaining and describing tensor structure information, the present invention realizes the comprehensive perception and efficient processing of tensor structure information in neural network programs, providing strong support for subsequent optimization.
[0087] Embodiment 2
[0088] As Figure 2 shown, an embodiment of the present application provides an analysis method for tensor structure information, and the method includes:
[0089] Program representation parsing step 201: Parse the source code of the program to obtain the specified initial values of all constant tensors, and create a structure information description tensor corresponding to the constant tensor. Among them, the description tensor has the same shape as the corresponding original tensor, and the value at each position of the description tensor is used to label the specific value or precision information of the corresponding original tensor element;
[0090] Initialization of tensor structure information description step 202: Determine which description tensors corresponding to specific values to obtain according to the settings of the application program. For the positions in the original tensor where the value is 0, the value at the corresponding position of the description tensor is set to 0; for the positions where the value is the preset specific value, the value at the corresponding position of the description tensor is set to the preset specific value.
[0091] Index description list calculation step 203: For the input neural network application program, based on program semantic analysis, analyze the structure information of each tensor in the program, create a description tensor of the structure information of the tensor, convert the description tensor into a nested list, and build a tree structure based on the nested list;
[0092] Tensor structure information semantic analysis step 204: Traverse the tree structure, record all node paths, form a constant tensor index description list, and based on the constant tensor index description list, obtain the index description list of non-constant tensors through program semantic analysis;
[0093] Index list optimization step 205: Adopt a structure index information optimization mechanism to merge and optimize the obtained constant tensor and non-constant tensor index description lists.
[0094] In a specific embodiment of the present invention, the above index description list calculation step 203 includes:
[0095] Tree structure construction step: Convert the description tensor into a nested list and build a tree structure. Among them, each layer of the tree corresponds to a dimension of the tensor, and the nodes of any layer represent the index information of the corresponding any dimension;
[0096] List contraction step: Start performing a contraction operation from the innermost dimension of the tree structure. The contraction operation is: if all elements in the nested list are the preset specific value, then replace the nested list with the preset specific value;
[0097] Iterative processing step: Starting from the inner layer of the tree structure, iterate outwards to the outer layer dimensions in sequence, perform a contraction operation on each layer. If the elements of a certain list are not all preset specific values, the contraction operation is not performed.
[0098] Index recording step: By traversing the tree, record all node paths where the subtrees are all preset specific values, forming a list of index descriptions of the preset specific values.
[0099] In a specific embodiment of the present invention, for the index description of tensor structure information, through in-depth semantic analysis of tensor operations in a neural network program, a systematic method is proposed to accurately capture the structure in the tensor. The system defaults to recognizing 0 and 1 elements in the tensor, and the index description of capturing the fixed value num can be set through the API setCatch Const(num, flag). In addition, the system captures the structured sparsity introduced by mask operations and the structure information of temporary tensors generated by user-defined program logic. By generating an associated matrix and a list of index descriptions, the system details the structure information of the tensor and uses this information for subsequent calculation processing.
[0100] Without loss of generality, assume that the shapes of tensors S and C are (d 1 , d 2 , …, d n ), where d i (i = 1, 2, …, n) represents the length of the i-th dimension. First, create a nested list describing the specific value a corresponding to the original tensor. First, create a list with a shape of (d 1 , d 2 , …, d n ), whose innermost layer is a list with a length of d n , d n being the number of specific values of the tensor being a. For the corresponding original tensor being the required specific value a, it is set to a, otherwise it is set to 0. The second-to-last layer contains d n-1 innermost layers with a length of d n …… The outermost layer list is d 1Sub - list. Assume a list is l. If all elements inside l are a specific value a, then replace l at its original position with a, or it can be said that l is contracted to a. Thus, the property of its parent list may become a new flat list all represented by a. Perform the above - mentioned contraction action of attempting replacement iteratively for each deepest - level list. If the elements inside a list are not all a, no contraction operation is performed. This nested list is essentially equivalent to a tree structure, which may be called T. The node value i represents the index i of the information structure represented by its subtree in the current dimension. The subtree corresponding to the node at height h represents the structural information of the h - th dimension. Assume the dimension of the original tensor S is n. Then a certain node at the h - th layer represents all the structural information in the h - th dimension, and the maximum depth of its subtree is n - h. By traversing this tree in breadth - first order, record the paths corresponding to all nodes that meet the following conditions to form an index. For example, i→j represents the path corresponding to the node with value j at depth 2 under the node with value i at depth 1, which is the index of the current value a, and all leaf nodes of its corresponding subtree are a. By obtaining a list of all such indexes, an index description of the specific value a can be obtained. By default, calculate the index description of the specific value 0 for each tensor. The generation of the index description for 0 can be cancelled by setCatch Const(0, false), and the generation of the index description of the tensor for a can be enabled by setting setCatch Const(a, true).
[0101] Furthermore, after the process of generating a sparse index list through semantics, for an index list, starting from the list item with the deepest nesting level, consider the nested list as a tree structure. The h - th layer corresponds to a node in the h - th dimension of the tensor. If the recorded index value is i, then its subtree represents the structural information of the i - th bit in the h - th dimension of the tensor. For a leaf node, assume the values on the path to the current node (including the leaf node itself) are (i 1 ,i 2 ,…,i h ), then it means that the sub - tensors corresponding to S[i 1 [i 2 …[i n of the source tensor S are all 0. Consider the following process. Start traversing from the parent node (the j - th layer) of the deepest - level leaf node. Whenever all the child nodes of this subtree can cover 0…d j , then contract the subtree into a node, that is, when all the child nodes of a parent node can exactly ensure that all the values of all sub - tensors are 0, then delete all the child nodes of this node and make this node a leaf node. Iterate the above process to obtain the most simplified index description.
[0102] Technical effects:
[0103] 1. Improve the accuracy of structural feature description: By generating index descriptions corresponding to the tensor structure for each required value, such as 0, 1, etc., the system can accurately label the structural features contained in the tensor for each specific required value, ensuring that important structural information is not missed during the calculation process.
[0104] 2. Enhance computational efficiency: By generating a list of index descriptions, the system can quickly locate specific structural information in the tensor, reducing redundant calculation operations. Utilizing features such as structured sparsity and the compacted index descriptions, the system can efficiently skip unnecessary elements in subsequent computational processing, thus significantly shortening the calculation time and optimizing the overall running performance of the neural network.
[0105] In a specific embodiment of the present invention, the above-mentioned tensor structure information semantic analysis step 204 further includes:
[0106] Assignment operation analysis step: When there is an assignment statement in the program and the constant at the i-th row and j-th column is set to 0, then add an item [i, j] to the index description list with a value of 0, where i and j respectively represent the number of rows and columns of the matrix;
[0107] Specific structure tensor recognition step: For a two-dimensional tensor with a diagonal of a specific value scale created by the program, where the length and width of the two-dimensional tensor are both l, add the array (0, 0), (1, 1), …, (l, l) to the index description list corresponding to scale.
[0108] In a specific embodiment of the present invention, the above-mentioned tensor structure information semantic analysis step 204 further includes:
[0109] Recognition and setting step of region unified value: The recognition and setting of the region unified value, by recognizing the region with a unified value in the tensor, setting the unified value region to a specific value;
[0110] Operation feature mining step of specific pattern: For the index positions with a specific interval pattern, establish a corresponding calculation template, define the step size and offset, and calculate the positions separated by the step size and offset;
[0111] Calculation of effective operation ratio step: The effective operation ratio is to statistically calculate the ratio of non-zero values in each innermost dimension to the dimension length. According to the effective operation ratio, different operation modes are corresponding, and the operation modes are extracted through semantic analysis or set by the program.
[0112] In the specific embodiments of the present invention, structural information and operation characteristics are captured from semantics. Based on obtaining the above sparse index list for constant tensors that are known initially for all programs, the sparse index lists of all non-constant tensors can be analyzed through program semantics. For the semantics of setting the (i,j)-th dimension of a certain tensor A to 0 (usually occurring in the pruning process, directly pruning a certain channel, here it can be pruning all channels of the j-th in the i-th batch), that is, the semantics of A[i][j]=0, then an item [i,j] is added to the index list corresponding to the 0 element. For another example, if a two-dimensional tensor with a length and width of l and a diagonal of a specific value scale is created in the program, then (0,0),(1,1),…,(l,l) are added to the index description corresponding to scale (create a description if there is none). This information is scalable and highly compressed, laying a solid foundation for runtime analysis and optimization.
[0113] By identifying regions with a unified value in the tensor (usually of great significance for matrix partitioning), these regions can be specially processed during the calculation. For a region that is all a specific value a, this value can be directly used in the calculation, avoiding calculating each element one by one. Use the specific API setRegionValue((i,j,…),(x1,y1),(x2,y2),a) to set that the region with the upper left corner at (x1,y1) and the lower right corner at (x2,y2) in the sub-tensor with the index (i,j,…) is all a. In particular, for a sparse operation mode, setRegionSparse((i,j,…),(x1,y1),(x2,y2),true) can be used to set that the region with the upper left corner at (x1,y1) and the lower right corner at (x2,y2) in the sub-tensor with the index (i,j,…) is all a. The flag true indicates that all other parts of this sub-tensor are 0 except for this region, and false indicates that the values of the sub-tensor outside this region are not set.
[0114] For index positions with a specific interval pattern, a corresponding calculation template can be established to perform calculations only on necessary positions. This can be achieved by defining the stride and offset. Starting from the starting address represented by the current index (i, j, …), the function setPatternStride((i, j, …), offset, t, m, a) is called only when calculating m consecutive data every t positions. When m is 1, it degenerates to the classical unstructured sparsity. While setting, the ratio p of the number of non-zero values in each innermost dimension to the length of that dimension can be counted, also known as the effective operation ratio. If the effective operation count for all one-dimensional dimensions of a two-dimensional matrix is set or detected (using getPatternStride((i, j, …))) to be 0.5, it indicates that only half of the original tensor size of operations are performed on each dimension, corresponding to a 2:4 operation pattern. Special operation patterns can be extracted through semantic analysis or set by the program.
[0115] Technical effects:
[0116] 1. Improve the utilization rate of computing resources: Optimized calculations for specific regions and patterns reduce unnecessary calculation operations and improve the utilization efficiency of computing resources.
[0117] 2. Enhance the computing efficiency: The optimized index description can accelerate the reading and processing of structural information, ensure that the structural information in the calculation process can be efficiently utilized, and avoid unnecessary waste of computing resources.
[0118] In the specific embodiments of the present invention, the above index list optimization step 205 further includes:
[0119] Node merging step: Starting from the deepest leaf node, traverse the tree upward. If all the child node indexes of a certain node are in the record and cover all positions of the corresponding dimension, then merge the child nodes and shrink the subtree into one node, making the node a leaf node;
[0120] Iterative simplification step: Repeat the node merging step, merge layer by layer upward until no further simplification is possible, and obtain the most simplified index description list.
[0121] The following uses Python source code as a specific implementation example of the present invention to illustrate the method of the present invention:
[0122] Step 1: Program representation parsing
[0123] Taking Python source code as input, parse its abstract syntax tree (AST) or intermediate representation (IR) to generate a data flow graph corresponding to the source code. By traversing the AST or IR, obtain all tensor definitions, assignments, and operations in the program, laying a foundation for subsequent capture of structural information.
[0124] Step 2: Initialize the structural information description of tensors
[0125] For all constant tensors with initial values parsed in Step 1, create corresponding structural information description tensors D. The description tensor has the same shape as the original tensor S, and the value at each position is used to label the specific value or precision information of the corresponding source tensor element. Determine which description tensors corresponding to specific values to obtain according to the API setting. By default, obtain the descriptions corresponding to 0 and 1. At the positions in the source tensor where the value is 0, set the value at the corresponding position in the description tensor to 0; at the positions where the value is a specific non-zero constant a, set the value at the corresponding position in the description tensor to a.
[0126] Step 3: Index description list calculation
[0127] Convert the description tensor into a nested list and construct a tree structure T. Each layer of the tree corresponds to a dimension of the tensor, and the nodes in the h-th layer represent the index information of the h-th dimension. The specific process is as follows
[0128] List contraction: Starting from the innermost dimension, if all elements in a list are a specific value a, then replace the list with a.
[0129] Iterative processing: Iterate outwards to the outer dimensions in turn, and try to perform the above contraction operation on each layer. If the elements in a list are not all a, then do not perform contraction.
[0130] Index recording: By performing a breadth-first traversal of tree T, record all node paths that satisfy that all subtrees are the specific value a, forming an index description list. For example, the path [i,j] indicates that at the position with index i in the first dimension and index j in the second dimension, the corresponding subtensor is filled with the specific value a.
[0131] Step 4: Semantic analysis of tensor structural information
[0132] On the basis of capturing the structural information of constant tensors, through the semantic analysis of the program, further capture the structural information of non-constant tensors, specifically including:
[0133] Assignment operation analysis: When there is an assignment statement in the program, such as A[i][j]=0, add an item [i,j] to the index description list where the value is 0. This usually occurs during the pruning process when directly pruning a certain channel.
[0134] Recognition of specific structure tensors: For a two-dimensional tensor created by the program with a diagonal of a specific value scale, add (0,0),(1,1),…,(l,l) to the corresponding index description, where l is the size of the tensor. This information is scalable and highly compressed, laying a solid foundation for runtime analysis and optimization.
[0135] Step 5: Identification and Setting of Region Unified Values
[0136] By identifying regions with unified values in the tensor, special processing can be performed on these regions. The setRegionValue is used to set the specific value for the unified region, and setRegionSparse is used to set the sparse specific region.
[0137] Step 6: Mining of Operation Characteristics of Specific Patterns
[0138] For index positions with a specific interval pattern, a corresponding calculation template is established, and calculations are only performed on necessary positions. By defining the stride and offset, the setPatternStride API is used to set them in the program.
[0139] Step 7: Merging and Simplification of Index Descriptions
[0140] Merge and optimize the list of index descriptions obtained in Steps 3, 4, 5, and 6:[[]]
[0141] Node merging: Starting from the deepest leaf nodes, traverse the tree T upward. If the index of all child nodes of a certain node is in the record and covers all positions of this dimension, then merge these child nodes to make this node a leaf node.
[0142] Iterative simplification: Repeat the above process, merge layer by layer upward until no further simplification is possible, and obtain the most simplified list of index descriptions.
[0143] Embodiment III
[0144] As Figure 3 shown, the embodiment of the present application provides an analysis system for tensor structure information, which adopts the analysis method of tensor structure information as described above. The system includes:[[]]
[0145] Program Representation Parsing Module 301: Used to parse the source code of the program, obtain the specified initial values of all constant tensors, and create a structure information description tensor corresponding to the constant tensor. Among them, the description tensor has the same shape as the corresponding original tensor, and the value at each position of the description tensor is used to label the specific value or precision information of the corresponding original tensor element;
[0146] Initialization Tensor Structure Information Description Module 302: Used to determine which description tensors corresponding to specific values to obtain according to the settings of the application program. For positions in the original tensor with a value of 0, the value at the corresponding position of the description tensor is set to 0; for positions with a preset specific value, the value at the corresponding position of the description tensor is set to the preset specific value.
[0147] Index description list calculation module 303: For the input neural network application program, based on program semantic analysis, analyze the structural information of each tensor in the program, create a structural information description tensor corresponding to the tensor, convert the description tensor into a nested list, and construct a tree structure based on the nested list;
[0148] Tensor structure information semantic analysis module 304: For traversing the tree structure, recording all node paths, forming a constant tensor index description list, and obtaining an index description list of non-constant tensors through program semantic analysis based on the constant tensor index description list;
[0149] Index list optimization module 305: For adopting a structural index information optimization mechanism to merge and optimize the obtained constant tensor and non-constant tensor index description lists.
[0150] Embodiment 4
[0151] An embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the method for analyzing tensor structure information described above are implemented.
[0152] Embodiment 5
[0153] An embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the method for analyzing tensor structure information as described above are implemented.
[0154] In addition, combined with Figure 1 The method for analyzing tensor structure information in the embodiments of the present application described can be implemented by an electronic device, such as a computer device. Figure 4 It is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present application.
[0155] In some of these embodiments, the computer device may further include a communication interface 83 and a bus 80. Among them, as Figure 4 shown, the processor 81, the memory 82, and the communication interface 83 are connected through the bus 80 and complete communication with each other.
[0156] Specifically, the above-mentioned processor 81 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured as one or more integrated circuits implementing the embodiments of the present application.
[0157] The memory 82 may be used to store or cache various data files required for processing and / or communication, as well as possible computer program instructions executed by the processor 81.
[0158] The processor 81 reads and executes the computer program instructions stored in the memory 82 to implement any one of the tensor structure information analysis methods in the above embodiments.
[0159] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0160] The above embodiments only represent several implementation manners of the present application, and the description is relatively specific and detailed. However, it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. A method for analyzing tensor structure information, characterized in that: The method comprises: Index description list calculation step: for the input neural network application, based on program semantic analysis, analyzing the structural information of each tensor in the program, creating a description tensor of the structural information of the tensor, converting the description tensor into a nested list, and building a tree structure based on the nested list; The tensor structure information semantic analysis step: traverse the tree structure, record all node paths, form a constant tensor index description list, and obtain an index description list of non-constant tensors through program semantic analysis based on the constant tensor index description list; Index list optimization step: using a structural index information optimization mechanism, the constant tensor and non-constant tensor index description lists obtained are merged and optimized.
2. The method for analyzing tensor structure information according to claim 1, characterized in that: Before the index description list calculation step, the method further includes: Program representation parsing step: parsing the source code of the program to obtain the specified initial values of all constant tensors, creating a structural information description tensor corresponding to the constant tensor, wherein the description tensor has the same shape as the corresponding original tensor, and the value of each position of the description tensor is used to mark the specific value or precision information of the corresponding original tensor element; Initialize the structural information description step of the tensor: determine which specific values to obtain the description tensor corresponding to according to the application settings, and set the value of the corresponding position of the description tensor to 0 for the position with a value of 0 in the original tensor; and set the value of the corresponding position of the description tensor to the preset specific value for the position with a value of the preset specific value.
3. The method for analyzing tensor structure information according to claim 1 or 2, characterized in that: The index description list calculation step comprises: Tree structure construction step: convert the description tensor into a nested list and construct a tree structure, wherein each layer of the tree corresponds to a dimension of the tensor, and the node of any layer represents the index information of the corresponding any dimension; List compaction step: performing a compaction operation starting from the innermost dimension of the tree structure, wherein the compaction operation is: if all elements in the nested list are preset specific values, then replacing the nested list with the preset specific value; Iterative processing step: iterating from the inner layer of the tree structure to the outer layer dimension in sequence, performing the compaction operation on each layer, and not performing the compaction operation if not all elements of a list are the preset specific values; Index recording step: by traversing the tree, recording all node paths that satisfy the subtrees having the preset specific values, forming an index description list of the preset specific values.
4. The method for analyzing tensor structure information according to claim 1 or 2, characterized in that: The tensor structure information semantic analysis step also includes: Assignment operation analysis steps: When there is an assignment statement in the program, and the constant in row i and column j is set to 0, add an item [i,j] to the index description list with a value of 0, where i and j represent the number of rows and columns of the matrix respectively; Specific structure tensor identification steps: For a two-dimensional tensor created by the program with a diagonal of a specific value scale, the length and width of the two-dimensional tensor are both l, then add an array (0,0), (1,1), ..., (l,l) to the index description list corresponding to the scale.
5. The method for analyzing tensor structure information according to claim 1 or 2, characterized in that: The tensor structure information semantic analysis step also includes: Identification and setting of regional uniform values: identification and setting of regional uniform values, by identifying regions with uniform values in the tensor, and setting regional uniform specific values of the uniform values; The operation feature mining step of a specific pattern: for the index position with a specific interval pattern, a corresponding calculation template is established, a step length and an offset are defined, and the position separated by the step length and the offset is calculated; Step of calculating effective operation ratio: the effective operation ratio is to count the ratio of non-zero values in each innermost dimension to the length of the dimension, and different operation modes correspond to the effective operation ratio, and the operation mode is extracted through semantic analysis or set by the program.
6. The method for analyzing tensor structure information according to claim 1 or 2, characterized in that: The index list optimization step further includes: Node merging step: starting from the deepest leaf node, traverse the tree upwards. If all child node indexes of a node are in the record and cover all positions of the corresponding dimension, merge the child nodes and shrink the subtree into one node, making the node a leaf node. Iterative simplification step: Repeat the node merging step, merging upward layer by layer until it cannot be simplified any further, and obtain the most simplified index description list.
7. A tensor structure information analysis system, using the tensor structure information analysis method according to any one of claims 1 to 6, characterized in that: The system comprises: Index description list calculation module: for analyzing the structural information of each tensor in the input neural network application based on program semantic analysis, creating a structural information description tensor corresponding to the tensor, converting the description tensor into a nested list, and building a tree structure based on the nested list; A tensor structure information semantic analysis module: used to traverse the tree structure, record all node paths, form a constant tensor index description list, and obtain an index description list of non-constant tensors through program semantic analysis based on the constant tensor index description list; Index list optimization module: used to merge and optimize the constant tensor and non-constant tensor index description lists obtained by adopting a structure index information optimization mechanism.
8. The method for analyzing tensor structure information according to claim 7, characterized in that: Before the index description list calculation module, it also includes: Program representation parsing module: used to parse the source code of the program, obtain the specified initial values of all constant tensors, and create a structural information description tensor corresponding to the constant tensor, wherein the description tensor has the same shape as the corresponding original tensor, and the value of each position of the description tensor is used to mark the specific value or precision information of the corresponding original tensor element; The structural information description module of the initialized tensor is used to determine which specific values to obtain the description tensor corresponding to according to the application settings. For positions in the original tensor where the value is 0, the value of the corresponding position in the description tensor is set to 0; for positions where the value is a preset specific value, the value of the corresponding position in the description tensor is set to the preset specific value.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the tensor structure information analysis method described in any one of claims 1 to 6 are implemented.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the tensor structure information analysis method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Method for optimizing calculation of embedded module in language model
CN115034198A
Method and device for jointly training model
CN115345298A
Neural network feature extraction and classification method based on multi-kernel width graph
CN115346055A
Compilation optimization method and device, computer equipment and storage medium
CN115756722A
Sparse tensor operation acceleration method and system, computer equipment and storage medium
CN117149778A