Data processing method and related device
By filtering the subgraph structure with high frequency in the database and building a subgraph-level Benchmark library, the problem of low construction efficiency of subgraph-level Benchmark library in the existing technology is solved, and more efficient chip performance optimization is achieved.
Patent Information
- Application Number
- CN202311765309.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-19
- Publication Date
- 2025-06-20
AI Technical Summary
The prior art lacks effective methods for building subgraph-level Benchmark libraries, resulting in inefficient performance optimization in chip design.
By obtaining the complete graph structure of the database, filter out subgraph structures with frequency greater than the preset threshold, and build a subgraph-level Benchmark library based on these subgraph structures to improve construction efficiency.
It realizes efficient construction of sub-graph Benchmark library, improves the efficiency of chip performance optimization, and avoids the problem of missing important sub-graph structures.
Smart Images

Figure CN120179836A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence, and in particular, to a data processing method and related devices. Background Art
[0002] In the chip design stage, it is often necessary to build a series of benchmark libraries to monitor the relevant performance of chip design. By establishing a suitable benchmark load library, the optimizations at the application level can be extracted to the chip level and incorporated into the chip microarchitecture, which helps to improve the performance of the chip. The commonly used benchmark libraries include two levels: the macro-benchmark library and the micro-benchmark library. Among them, the macro-benchmark library usually refers to the model-level benchmark library, and the micro-benchmark library usually refers to the operator-level benchmark library or the instruction-level benchmark library. The model-level benchmark library is too large and the care cost is relatively high. The operator-level benchmark library lacks the connection relationships between operators, and the load characteristics are inaccurately characterized. Therefore, it is also necessary to build a subgraph-level benchmark library, which includes operator sequences and the connection relationships between operators, so as to build a complete benchmark library system.
[0003] However, for the subgraph-level benchmark library, there is currently no effective construction method. Summary of the Invention
[0004] This application provides a data processing method and related devices for improving the construction efficiency of the subgraph-level benchmark library.
[0005] In the first aspect of this application, a data processing method is provided, which can be applied to a data processing device. The method includes: obtaining the complete graph structure of a database, where the database includes multiple models, the complete graph structure includes multiple graph structures, each model in the multiple models corresponds to one graph structure in the multiple graph structures, and each graph structure in the multiple graph structures further includes a subgraph structure, and the subgraph structure is used to indicate multiple operators and the connection relationships between the multiple operators; screening out a first subgraph structure from the subgraph structures included in the complete graph structure, where the frequency of occurrence of the first subgraph structure in the complete graph structure is greater than a preset threshold, and the first subgraph structure is used to determine a target subgraph structure; and constructing a subgraph-level benchmark library based on the target subgraph structure.
[0006] The database includes numerous models, and each model corresponds to a graph structure respectively. Therefore, the complete graph structure of the database includes multiple graph structures. The graph structure further includes a subgraph structure, and the subgraph structure is a subset of the graph structure. Moreover, the subgraph structure can be a non-proper subset, that is, the subgraph structure can be the same as the graph structure. The subgraph structure is used to indicate multiple operators and the connection relationships between the multiple operators.
[0007] After obtaining the complete graph structure of the database, a first sub-graph structure with an occurrence frequency greater than a preset threshold is screened out from the complete graph structure, and a target sub-graph structure for constructing a sub-graph level Benchmark library is determined based on the first sub-graph structure. The fact that the sub-graph structure appears more frequently indicates that the sub-graph structure will be frequently used, that is, it can be considered that the sub-graph structure is an important sub-graph structure and can be used to construct a sub-graph level Benchmark library.
[0008] In the first aspect of the present application, the first sub-graph structure for constructing a sub-graph level Benchmark library is quickly screened out based on the occurrence frequency of the sub-graph structure, greatly improving the construction efficiency of the sub-graph level Benchmark library.
[0009] In a possible implementation manner of the first aspect, the method further includes: screening out a first type of models from the database, the first type of models includes multiple models, and the multiple models included in the first type of models are of the same type; determining a first type of graph structure corresponding to the first type of models in the complete graph structure; screening out a second sub-graph structure from the sub-graph structures included in the first type of graph structure, the occurrence frequency of the second sub-graph structure in the first type of graph structure is greater than a preset threshold, and the second sub-graph structure is used to determine the target sub-graph structure.
[0010] When the types and quantities of models in the database are large, the number of sub-graph structures will be very large. Moreover, as the number of sub-graph structures increases, the support ratio of the sub-graph structures will also decrease. Therefore, some important sub-graph structures may be missed during the screening process. To address this issue, the complete graph structure can be screened multiple times to ensure the integrity of the sub-graph level Benchmark library.
[0011] There are usually more identical sub-graph structures among models of the same type. In other words, there are more important sub-graph structures for constructing a sub-graph level Benchmark library in models of the same type. Therefore, secondary mining for models of the same type can avoid missing some important sub-graph structures. It can be understood that models of the same type can also be referred to as models in the same scenario, that is, models applied to the same application scenario.
[0012] Specifically, first classify the models in the database and screen out the first-class models from them. The first-class models include multiple models, and all the multiple models are of the same type. Classifying the models can be done by manually identifying models in the same scenario for classification, or by the data processing device extracting the features of the models and classifying them according to the features. After determining the first-class models, then find the first-class graph structures corresponding to the first-class models from the complete graph structure. The first-class graph structures include multiple graph structures, and each model in the first-class models corresponds to one graph structure in the first-class graph structures. After determining the first-class graph structures, mine the first-class graph structures again, screen out the second sub-graph structures whose occurrence frequency is greater than the preset threshold from the first-class graph structures, and determine the target sub-graph structures for constructing the sub-graph level Benchmark library based on the second sub-graph structures. By re-mining the models of the same type with more frequent sub-graph structures, it is possible to avoid missing important sub-graph structures as much as possible, thereby improving the integrity of the sub-graph level Benchmark library.
[0013] In a possible implementation manner of the first aspect, the sub-graph structure is counted once in each graph structure where it appears.
[0014] In this possible implementation manner, when counting the occurrence frequency of the sub-graph structure, it is only counted once in each graph structure that contains the sub-graph structure. That is, when screening the first sub-graph structure from the complete graph structure, it is only counted once in each graph structure that contains the first sub-graph structure. When screening the second sub-graph structure from the first-class graph structures, it is only counted once in each graph structure that contains the second sub-graph structure.
[0015] Since both the complete graph structure and the first-class graph structures include multiple graph structures, the occurrence frequency of the same sub-graph structure may be relatively high. Therefore, only counting a sub-graph structure once in the graph structure can achieve fast screening and further improve the construction efficiency of the sub-graph level Benchmark library.
[0016] Of course, when counting the occurrence frequency of the sub-graph structure, it can also be counted as many times as the sub-graph structure appears, so as to perform more accurate screening. In this application, no specific limitation is made on the statistical rule.
[0017] In a possible implementation manner of the first aspect, the method further includes: screening out the first model from the database, where the first model is a model that meets the importance condition; determining the first graph structure corresponding to the first model in the complete graph structure; screening out the third sub-graph structure from the sub-graph structures included in the first graph structure, where the occurrence frequency of the third sub-graph structure in the first graph structure is greater than the preset threshold, and the third sub-graph structure is used to determine the target sub-graph structure.
[0018] In the model of a database, there are usually some important models, and the sub-graph structures in the graph structures corresponding to the important models are also relatively important. Therefore, the graph structures corresponding to the key models are mined again to avoid missing important sub-graph structures.
[0019] Specifically, first screen out the important models (i.e., the first models) from the models included in the database. The screening method can be that the data processing device judges the importance of the models, and regards the models that meet the importance conditions as the first models. Or, the user screens the models, and when a model that meets the importance conditions is recognized, it is selected as the first model. Then, find the first graph structure corresponding to the first model from the complete graph structure, screen out the third sub-graph structures with a support greater than the preset threshold from the first graph structure, and finally determine the target sub-graph structure based on the third sub-graph structures. By re-mining the graph structures of the important models, it is possible to avoid missing important sub-graph structures as much as possible, ensuring the integrity of the sub-graph level Benchmark library.
[0020] In a possible implementation manner of the first aspect, the method further includes: performing a merging process on the first sub-graph structure, the second sub-graph structure, and the third sub-graph structure to obtain the first type of sub-graph structure; deleting the duplicate sub-graph structures in the first type of sub-graph structure to obtain the target sub-graph structure.
[0021] After multiple levels of screening, there will be some duplicate sub-graph structures. To reduce redundancy, the same parts in the first sub-graph structure, the second sub-graph structure, and the third sub-graph structure are deleted. Specifically, perform a merging process on the first sub-graph structure, the second sub-graph structure, and the third sub-graph structure to obtain the first type of sub-graph structure. Then delete the duplicate sub-graph structures in the first type of sub-graph structure, thereby obtaining the target sub-graph structure for constructing the sub-graph level Benchmark library.
[0022] In this possible implementation manner, deleting the duplicate sub-graph structures reduces the redundancy of the sub-graph structures and can improve the construction efficiency of the sub-graph level Benchmark library.
[0023] In a possible implementation manner of the first aspect, the above step: obtaining the complete graph structure of the database includes: converting multiple models into a complete graph structure through a graph generation tool.
[0024] Some machine learning frameworks applied in the database can directly convert the model into a graph structure. For example, TensorFlow or MindSpore. In this case, the complete graph structure of the database can be directly obtained. However, some machine learning frameworks cannot convert the model into a graph structure. For example, the PyTorch framework. In this case, a graph generation tool needs to be used. Specifically, first build a graph generation tool for the machine learning framework, and then use the built graph generation tool to convert the model in the database into a graph structure, so as to obtain the complete graph structure of the database.
[0025] In this possible implementation method, the way to obtain the complete graph structure is defined, which improves the feasibility of the solution.
[0026] In a possible implementation manner of the first aspect, the method further includes: converting the target sub-graph structure into a target model.
[0027] The sub-graph level Benchmark library is used to quickly evaluate the performance of a computing system (such as a chip). However, the target sub-graph structure cannot be directly executed on the computing system. Therefore, the target sub-graph structure is converted into a target model that can be executed on the computing system and is represented by a machine learning framework. Machine learning frameworks such as PyTorch, MindSpore, and TensorFlow, etc. Specifically, first build a sub-graph to model tool for the machine learning framework. Through the sub-graph to model tool, the nodes in the target sub-graph structure are converted into statements or application programming interfaces (APIs) in the machine learning framework, and then the statements or APIs are associated according to the connection relationships in the target sub-graph structure, so as to form the target model. Subsequently, the performance of the computing system can be evaluated by executing the target model.
[0028] In this possible implementation method, the target sub-graph structure is pre-converted into a model that can be directly run on the computing system, which can improve the efficiency of using the sub-graph level Benchmark library to evaluate the computing system.
[0029] In a possible implementation manner of the first aspect, the method further includes: extracting the features of the target model; classifying the target sub-graph structure based on the features.
[0030] After converting the target sub-graph structure into a target model, run the target model on the computing system, and extract the relevant features of the target model (i.e., the target sub-graph structure) according to the running process. The extracted features can be features at the framework level, such as graph structure features, computing features, and access features, etc. Specifically, they are operator types, computing amounts, memory access amounts, and in-degree and out-degree of nodes, etc. The features can also be features during the execution process of the load, such as the time-consuming of running the target model and the access of different pipelines, etc.
[0031] After extracting the features of the target model, the target subgraph structures are divided into different categories according to the features. The features of the target subgraph structures in the same category are the same, and the requirements for the computing system are also the same. Therefore, when evaluating the performance of the computing system through the subgraph-level Benchmark library later, only a few target subgraph structures need to be selected from each category to run, instead of executing all the target subgraph structures, thus greatly improving the efficiency of evaluating the computing system.
[0032] The second aspect of the present application provides a data processing device, including an acquisition unit, a screening unit, and a construction unit. The acquisition unit is configured to acquire the complete graph structure of the database. The database includes multiple models, the complete graph structure includes multiple graph structures, each model in the multiple models corresponds to one graph structure in the multiple graph structures, and each graph structure in the multiple graph structures includes a subgraph structure, and the subgraph structure is used to indicate multiple operators and the connection relationships between the multiple operators; the screening unit is configured to screen out the first subgraph structure from the subgraph structures included in the complete graph structure, and the frequency of occurrence of the first subgraph structure in the complete graph structure is greater than a preset threshold, and the first subgraph structure is used to determine the target subgraph structure; the construction unit is configured to construct a subgraph-level Benchmark library based on the target subgraph structure.
[0033] In a possible implementation manner of the second aspect, the screening unit is further configured to screen out the first type of models from the database. The first type of models includes multiple models, and the multiple models included in the first type of models are of the same type; the device further includes a determination unit configured to determine the first type of graph structures corresponding to the first type of models in the complete graph structure; the screening unit is further configured to screen out the second subgraph structure from the subgraph structures included in the first type of graph structures, and the frequency of occurrence of the second subgraph structure in the first type of graph structures is greater than a preset threshold, and the second subgraph structure is used to determine the target subgraph structure.
[0034] In a possible implementation manner of the second aspect, the subgraph structure is counted once in each graph structure where it appears.
[0035] In a possible implementation manner of the second aspect, the screening unit is further configured to screen out the first model from the database, and the first model is a model that meets the importance condition; the determination unit is further configured to determine the first graph structure corresponding to the first model in the complete graph structure; the screening unit is further configured to screen out the third subgraph structure from the subgraph structures included in the first graph structure, and the frequency of occurrence of the third subgraph structure in the first graph structure is greater than a preset threshold, and the third subgraph structure is used to determine the target subgraph structure.
[0036] In a possible implementation of the second aspect, the apparatus further includes: a merging unit, configured to perform a merging process on the first sub-graph structure, the second sub-graph structure, and the third sub-graph structure to obtain a first type of sub-graph structure; a deleting unit, configured to delete the duplicate sub-graph structures in the first type of sub-graph structure to obtain a target sub-graph structure.
[0037] In a possible implementation of the second aspect, the obtaining unit is specifically configured to convert multiple models into a complete graph structure through a graph generation tool.
[0038] In a possible implementation of the second aspect, the apparatus further includes a conversion unit, configured to convert the target sub-graph structure into a target model.
[0039] In a possible implementation of the second aspect, the apparatus further includes: an extracting unit, configured to extract features of the target model; a classifying unit, configured to classify the target sub-graph structure based on the features.
[0040] The data processing apparatus provided in the second aspect of the present application is used to execute the method described in the first aspect or any possible implementation manner of the first aspect.
[0041] A data processing apparatus provided in the third aspect of the present application includes a processor and a memory. The memory is used to store instructions, and the processor is used to obtain the instructions stored in the memory to execute the method described in the first aspect or any possible implementation manner of the first aspect.
[0042] A computer-readable storage medium provided in the fourth aspect of the present application includes instructions, and when the instructions run on a computer, the computer is caused to execute the method described in the first aspect or any possible implementation manner of the first aspect.
[0043] A computer program product containing instructions provided in the fifth aspect of the present application, when the computer program product runs on a computer, causes the computer to execute the method described in the first aspect or any possible implementation manner of the first aspect.
[0044] A chip system provided in the sixth aspect of the present application includes at least one processor and a communication interface. The communication interface and the at least one processor are interconnected by a line, and the at least one processor is used to run a computer program or instructions to execute the method described in the first aspect or any possible implementation manner of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 It is a schematic diagram of an embodiment of the data processing method provided in the embodiment of the present application;
[0046] Figure 2 It is a schematic diagram of the sub - graph processing process in the embodiment of the present application;
[0047] Figure 3 It is a schematic diagram of the target sub - graph structure in the embodiment of the present application;
[0048] Figure 4 It is another schematic diagram of the target sub - graph structure in the embodiment of the present application;
[0049] Figure 5 It is another schematic diagram of the target sub - graph structure in the embodiment of the present application;
[0050] Figure 6 It is another schematic diagram of the target sub - graph structure in the embodiment of the present application;
[0051] Figure 7 It is another schematic diagram of the data processing method provided by the embodiment of the present application;
[0052] Figure 8 It is a schematic diagram of the structure of the data processing device provided by the embodiment of the present application;
[0053] Figure 9 It is another schematic diagram of the structure of the data processing device provided by the embodiment of the present application. Detailed implementation manners
[0054] The embodiment of the present application provides a data processing method for improving the construction efficiency of the sub - graph - level Benchmark library. The embodiment of the present application also provides corresponding devices, computer - readable storage media, computer program products, etc. The following will be described separately.
[0055] Next, in combination with the accompanying drawings, the embodiments of the present application will be described. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. Those of ordinary skill in the art know that with the development of technology and the emergence of new scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.
[0056] In the description and claims of this application, and in the above-mentioned drawings, terms such as "system" and "network", "weight" and "weight value" can be used interchangeably. Unless otherwise specified, ordinal numbers such as "first" and "second" are used to distinguish multiple objects and are not used to limit the order, time sequence, priority or importance of multiple objects. It should be understood that such terms can be interchanged under appropriate circumstances so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units need not be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0057] For ease of understanding, the following first introduces the relevant terms and concepts mainly involved in the embodiments of this application.
[0058] 1. Benchmark
[0059] Benchmark is an evaluation method that has been applied for a long time in the entire computer field. A typical application is computer performance testing, which mainly uses stress testing to explore the performance of the entire computer system.
[0060] 2. Sub-graph level Benchmark library
[0061] The sub-graph structure identified from the original load contains the operator sequence and the connection relationship between operators. The performance of the computing system is evaluated by executing the sub-graph structure in the sub-graph level Benchmark library on the computing system (such as a chip).
[0062] In the chip design stage, it is often necessary to build a series of benchmark (Benchmark) libraries to monitor the relevant performance of the chip design. By establishing a suitable Benchmark load library, the optimizations at the application level can be extracted to the chip level and incorporated into the chip micro-architecture, which helps to improve the performance of the chip. The commonly used Benchmark libraries include two levels: the macro-Benchmark library and the micro-Benchmark library. Among them, the macro-Benchmark library usually refers to the model-level benchmark library, and the micro-Benchmark library usually refers to the operator-level benchmark library or the instruction-level benchmark library. The model-level benchmark library is too large and the care cost is relatively high. The operator-level benchmark library lacks the connection relationship between operators, and the load characteristics are not accurately characterized. Therefore, it is also necessary to build a sub-graph level benchmark library, which includes the operator sequence and the connection relationship between operators, so as to build a complete Benchmark library system.
[0063] Currently, the commonly used screening method for the model-level Benchmark library is feature analysis plus expert identification. By using feature extraction and clustering algorithms to screen relevant loads, finally, experts manually screen out the key loads that their own companies and the industry are concerned about. The operator-level Benchmark library is usually screened from the model level. Therefore, the construction methods of the model-level and operator-level Benchmark libraries are relatively mature. However, the construction method of the subgraph-level Benchmark library is not mature in the industry. For example, the method of pure manual screening is adopted, and experts select some substructures from key artificial intelligence (AI) models as the subgraph-level Benchmark library. However, there are numerous types and quantities of AI loads. If pure manual selection is carried out, the workload is too large and the efficiency is low. Moreover, the integrity of load features cannot be guaranteed.
[0064] In view of this, the embodiments of the present application provide a data processing method, which can be applied to a data processing device. By identifying the subgraph structures that appear more frequently in the database and using these subgraph structures as the subgraph-level Benchmark library, the construction efficiency of the subgraph-level Benchmark library can be improved, and the integrity of the subgraph-level Benchmark library can also be improved.
[0065] Please refer to Figure 1 , Figure 1 which is a schematic diagram of an embodiment of the data processing method provided by the embodiments of the present application. As Figure 1 shown, this embodiment includes steps 101 to 105.
[0066] 101. Obtain the complete graph structure of the database.
[0067] First, obtain the complete graph structure of the database. The database includes numerous models, and each model corresponds to a graph structure respectively. That is to say, the complete graph structure of the database includes multiple graph structures. The graph structure also includes subgraph structures, and the subgraph structure is a subset of the graph structure. Moreover, the subgraph structure can be a non-proper subset, that is, the subgraph structure can be the same as the graph structure. The subgraph structure is used to indicate multiple operators and the connection relationships between multiple operators.
[0068] Some machine learning frameworks applied in databases can directly convert models into graph structures. For example, TensorFlow or MindSpore. In this case, the complete graph structure of the database can be directly obtained. However, some machine learning frameworks cannot convert models into graph structures. For example, the PyTorch framework. In this case, a graph generation tool needs to be used. Specifically, first build a graph generation tool for the machine learning framework, and then use the built graph generation tool to convert the models in the database into graph structures, so as to obtain the complete graph structure of the database.
[0069] 102. Screen out the target sub-graph structure from the complete graph structure. The target sub-graph structure is used to construct the sub-graph level Benchmark library.
[0070] After obtaining the complete graph structure of the database, screen out the first sub-graph structure whose occurrence frequency is greater than the preset threshold from the complete graph structure, and determine the target sub-graph structure for constructing the sub-graph level Benchmark library based on the first sub-graph structure. If a sub-graph structure appears more frequently, it means that this sub-graph structure is often used, that is, it can be considered that this sub-graph structure is an important sub-graph structure and can be used to construct the sub-graph level Benchmark library. The preset threshold can be set according to actual needs. For example, the threshold can be set to 5 times, or the threshold can be set to 10 times. Specifically, it is not limited here.
[0071] The frequency of occurrence of a sub-graph structure in the complete graph structure can also be called the support degree of the sub-graph structure. A sub-graph structure whose occurrence frequency is greater than the preset threshold can be called a frequent sub-graph structure. When the types and quantities of models in the database are large, the number of sub-graph structures will be very large. Moreover, as the number of sub-graph structures increases, the proportion of the support degree of the sub-graph structure will also decrease. Therefore, some important sub-graph structures may be missed during the screening process. To address this problem, the complete graph structure can be screened multiple times to ensure the integrity of the sub-graph level Benchmark library.
[0072] Optionally, there are usually more identical sub-graph structures among models of the same type. In other words, there are more important sub-graph structures for constructing the sub-graph level Benchmark library in models of the same type. Therefore, secondary mining for models of the same type can avoid missing some important sub-graph structures. It can be understood that models of the same type can also be called models of the same scenario, that is, models applied to the same application scenario.
[0073] Specifically, first classify the models in the database and screen out the first-class models from them. The first-class models include multiple models, and all the multiple models are of the same type. Classifying the models can be done by manually identifying models in the same scenario for classification, or by the data processing device extracting the features of the models and classifying them according to the features. After determining the first-class models, then find the first-class graph structures corresponding to the first-class models from the complete graph structure. The first-class graph structures include multiple graph structures, and each model in the first-class models corresponds to one graph structure in the first-class graph structures.
[0074] After determining the first-class graph structures, mine the first-class graph structures again, screen out the second sub-graph structures with the occurrence frequency greater than the preset threshold from the first-class graph structures, and determine the target sub-graph structures for constructing the sub-graph level Benchmark library based on the second sub-graph structures. By re-mining the models of the same type with more frequent sub-graph structures, it is possible to avoid missing important sub-graph structures as much as possible, thereby improving the integrity of the sub-graph level Benchmark library.
[0075] In a possible solution, when counting the occurrence frequency of the sub-graph structures, it is only counted once in each graph structure containing the sub-graph structure. That is, when screening the first sub-graph structures from the complete graph structure, it is only counted once in each graph structure containing the first sub-graph structure. When screening the second sub-graph structures from the first-class graph structures, it is only counted once in each graph structure containing the second sub-graph structure. Exemplarily, the complete graph structure includes four graph structures A, B, C, and D. The first sub-graph structure is included in graph structures A and D, and the first sub-graph structure appears 3 times in A and 2 times in D. Since it is only counted once in each graph structure, the occurrence frequency of the first sub-graph structure in the complete graph structure is 2 times.
[0076] Since both the complete graph structure and the first-class graph structures include multiple graph structures, the occurrence frequency of the same sub-graph structure may be relatively high. Therefore, only counting once for a sub-graph structure in the graph structure can achieve fast screening and further improve the construction efficiency of the sub-graph level Benchmark library.
[0077] Of course, when counting the occurrence frequency of the sub-graph structures, it can also be counted as many times as the sub-graph structure appears, so as to more accurately screen out the frequent sub-graph structures. The specific counting rule is not specifically limited in this embodiment. Then in the above example, the occurrence frequency of the first sub-graph structure in the complete graph structure is the 3 times it appears in graph structure A plus the 2 times it appears in graph structure D, totaling 5 times.
[0078] The statistical rule that the sub - graph structure is counted only once in each graph structure is called Scheme One, and the statistical rule that the sub - graph structure is counted as many times as it appears is called Scheme Two. It should be noted that for the statistics of the support (i.e., frequency) of the sub - graph structure, it can be that Scheme One is adopted when screening the first sub - graph structure and Scheme Two is adopted when screening the second sub - graph structure. It can also be that Scheme Two is adopted when screening the first sub - graph structure and Scheme One is adopted when screening the first sub - graph structure. It can also be that Scheme One is adopted when screening both the first sub - graph structure and the second sub - graph structure, or Scheme Two is adopted, and the specific situation is not limited here.
[0079] Taking the statistical rule of Scheme One as an example, the calculation process is described below.
[0080] Specifically, use GS = {G i |i = 0…n} to represent the complete graph structure, where i and n are integers, and G i is the graph structure among them. The sub - graph structures g and G are sub - graph structures in G i . When g and G are the same (i.e., g is a sub - graph isomorphism of G), record ζ(g,G)=1. When g and G are different (i.e., g is not a sub - graph isomorphism of G), record ζ(g,G)=0. That is, when G i includes g, ζ(g,G i ) = 1. When g does not appear in G i , ζ(g,G i ) = 0. Then ∑ Gi∈GS ζ(g,G i ) is the frequency of the sub - graph structure g appearing in GS, and use σ(g,GS)=∑ Gi∈GS ζ(g,G i ) to represent the frequency of the sub - graph structure g appearing.
[0081] When σ(g,GS) is greater than the preset threshold, then g is considered a frequent sub - graph structure, that is, g is the first sub - graph structure or the second sub - graph structure.
[0082] Optionally, there are usually some important models in the database model, and the sub - graph structures in the graph structures corresponding to the important models are also relatively important. Therefore, the graph structures corresponding to the key models are mined again to avoid missing important sub - graph structures.
[0083] Specifically, first, important models (i.e., the first models) are screened out from the models included in the database. The screening method can be that the data processing device evaluates the importance of the models, and the models that meet the importance conditions are regarded as the first models. Alternatively, the user screens the models, and when a model that meets the importance conditions is identified, it is selected as the first model. Then, the first graph structure corresponding to the first model is found from the complete graph structure, and then single-graph frequent subgraph mining is performed on the first graph structure. The third subgraph structure with a support greater than the preset threshold is screened out from the first graph structure. Finally, the target subgraph structure is determined based on the third subgraph structure. By re-mining the graph structure of the important models, it is possible to avoid missing important subgraph structures as much as possible and ensure the integrity of the subgraph-level Benchmark library.
[0084] It can be understood that the preset thresholds in the above multiple levels of screening processes can be the same or different. Exemplarily, the thresholds for screening the first subgraph structure, the second subgraph structure, and the third subgraph structure are the same, all being 5 times. Or, since the scale of the complete graph structure is large, the threshold for screening the first subgraph structure is set higher, to 8 times. The number of subgraph structures within the same type of models and a single key model is small, so the thresholds for screening the second subgraph structure and the third subgraph structure are set lower. The threshold for screening the second subgraph structure is set to 5 times, and the threshold for screening the third subgraph structure is set to 3 times.
[0085] Optionally, after the above multiple levels of screening, there will be some duplicate subgraph structures. To reduce redundancy, the same parts in the first subgraph structure, the second subgraph structure, and the third subgraph structure are deleted. Specifically, the first subgraph structure, the second subgraph structure, and the third subgraph structure are merged to obtain the first type of subgraph structure. Then, the duplicate subgraph structures in the first type of subgraph structure are deleted, thereby obtaining the target subgraph structure for constructing the subgraph-level Benchmark library.
[0086] Exemplarily, please refer to Figure 2 , in the mining process of the full quantum graph (i.e., the complete graph structure), the subgraph mining process of the same type of models, and the subgraph mining process of the key model, the subgraph structure shown in Figure 2 is screened out, that is, the first subgraph structure, the second subgraph structure, and the third subgraph structure all include this subgraph structure. Therefore, the duplicate subgraph structure is deleted to reduce the redundancy of the subgraph structure and improve the construction efficiency of the subgraph-level Benchmark library.
[0087] It can be understood that when determining the target sub-graph structure, redundant parts can be deleted or not. That is to say, after merging the first sub-graph structure, the second sub-graph structure and the third sub-graph structure to obtain the first type of sub-graph structure, the repeated parts in the first type of sub-graph structure can be deleted to obtain the target sub-graph structure. It can also be directly taking the first type of sub-graph structure as the target sub-graph structure, and specifically there is no limitation here.
[0088] After determining the target sub-graph structure, use the target sub-graph structure to construct a sub-graph level Benchmark library.
[0089] The selected target sub-graph structure can be referred to Figures 3 to 6 for understanding. As Figure 3 shown, the target sub-graph structure is a multi-input sub-graph structure, that is, a node (operator) has multiple inputs. Figure 4 It is a closed sub-graph structure, that is, a node has two inputs, and the input data of one of the input nodes is obtained based on the output data of the other input node. Figure 5 It is a multi-output sub-graph structure, that is, the output data of a node is the input data of multiple nodes. Figure 6 It is a complex sub-graph structure, and the connection relationship between nodes is relatively complex.
[0090] 103. Convert the target sub-graph structure into a target model.
[0091] The sub-graph level Benchmark library is used to quickly evaluate the performance of a computing system (such as a chip). However, the target sub-graph structure cannot be directly executed on the computing system. Therefore, the target sub-graph structure is converted into a target model that can be executed on the computing system represented by a machine learning framework. Machine learning frameworks such as pytorch, MinSpore, and TensorFlow, etc. Specifically, first build a sub-graph to model tool for the machine learning framework. Through the sub-graph to model tool, the nodes in the target sub-graph structure are converted into statements or application programming interfaces (APIs) in the machine learning framework, and then the statements or APIs are associated according to the connection relationship in the target sub-graph structure to form the target model. Subsequently, the performance of the computing system can be evaluated by executing the target model.
[0092] 104. Extract the features of the target sub-graph structure.
[0093] After converting the target subgraph structure into a target model, run the target model on a computing system and extract relevant features of the target model (i.e., the target subgraph structure) according to the running process. The extracted features can be features at the framework level, such as graph structure features, computing features, and access features, etc., specifically operator types, computing amounts, memory access amounts, and in-degrees and out-degrees of nodes, etc. The features can also be features during the execution of the workload, such as the time taken to run the target model and accesses in different pipelines, etc.
[0094] 105. Perform clustering processing on the target subgraph structure according to the extracted features.
[0095] After extracting the features of the target subgraph structure, divide the target subgraph structure into different categories according to the features. The target subgraph structures in the same category have the same features and the same requirements for the computing system. Therefore, when evaluating the performance of the computing system through the subgraph-level Benchmark library later, only a few target subgraph structures need to be selected from each category to run, instead of executing all the target subgraph structures, thus greatly improving the efficiency of evaluating the computing system.
[0096] It can be understood that steps 103 to 105 are optional steps, and the construction of the subgraph-level Benchmark library can be achieved through steps 101 and 102. However, by executing steps 103 to 105, the evaluation efficiency of the subgraph-level Benchmark library can be improved.
[0097] In this embodiment, the target subgraph structures for constructing the subgraph-level Benchmark library are quickly screened based on the occurrence frequency of the subgraph structures, greatly improving the construction efficiency of the subgraph-level Benchmark library. Moreover, through screening at three levels of the full quantum subgraph structure (i.e., the subgraph structure included in the complete graph structure), the same-scenario subgraph structure, and the important subgraph structure, important subgraph structures can be avoided from being omitted as much as possible, improving the integrity of the subgraph-level Benchmark library.
[0098] Combined with the above content, the data processing method provided in the embodiments of the present application will be exemplarily described below. Please refer to Figure 7 , which is another schematic diagram of the data processing method provided in the embodiments of the present application.
[0099] As Figure 7As shown, first convert the workload (i.e., the model) into a complete graph structure, and then perform hierarchical mining on the complete graph structure. The complete graph structure is divided into three levels: the full quantum graph structure, the sub-graph structure of the same type of model (i.e., the first type of model), and the sub-graph structure of the important model (i.e., the first model). Both the full quantum graph structure and the sub-graph structure of the same type correspond to multiple models. Therefore, perform multi-graph frequent sub-graph mining (graph-based mining) on the full quantum graph structure and the sub-graph structure of the same type, and filter out the frequent sub-graph structures. During the statistical process of multi-graph frequent sub-graph mining, a sub-graph structure that appears in a graph structure (model) is only counted once. For example, although the sub-graph structure A appears 3 times in the graph structure B, during the statistical process, the sub-graph structure A in the graph structure B is only counted once. After performing multi-graph frequent sub-graph mining on the full quantum graph structure, filter out the first sub-graph structure. After performing multi-graph frequent sub-graph mining on the sub-graph structure of the same type, filter out the second sub-graph structure. The sub-graph structure of the important model corresponds to a single model. Therefore, perform single-graph frequent sub-graph mining (embedding-based mining) on the sub-graph structure of the important model, and filter out the frequent sub-graph structure (i.e., the third sub-graph structure). During the statistical process of single-graph frequent sub-graph mining, the number of times a sub-graph structure appears in a graph structure is counted as many times as it appears. Exemplarily, if the sub-graph structure D appears 5 times in the graph structure C, then the occurrence frequency of the sub-graph structure D is 5 times.
[0100] After sub-graph mining at three levels, duplicate sub-graph structures will be filtered out. To reduce redundancy and improve the construction efficiency of the sub-graph-level Benchmark library, perform fusion processing on the filtered first sub-graph structure, second sub-graph structure, and third sub-graph structure. That is, first merge the first sub-graph structure, second sub-graph structure, and third sub-graph structure into the first type of sub-graph structure, and then delete the duplicate sub-graph structures in the first type of sub-graph structure to obtain the target sub-graph structure, which is used to construct the sub-graph-level Benchmark library.
[0101] The sub-graph level Benchmark library is used to evaluate a computing system. However, the target sub-graph structure cannot be directly run on the computing system. Therefore, the target sub-graph structure is converted into a target model. Specifically, a tool for converting sub-graphs into models is constructed to convert the target sub-graph structure into a target model represented by a machine learning framework that can be directly executed on the computing system. After converting into the target model, the features of the target sub-graph structure are extracted by running the target model on the computing system. The features include operator type, computational volume, memory access volume, and in-degree and out-degree of nodes, etc. After extracting the features of the target sub-graph structure, the target sub-graph structure is divided into different types according to the features. The target sub-graph structures in the same type have the same features and the same requirements for the computing system. Therefore, when evaluating the computing system, only a few models corresponding to the target sub-graph structures in each type need to be selected for evaluation, which can greatly improve the evaluation efficiency. After dividing the target sub-graph structure into different types, the construction of the sub-graph level Benchmark library is completed.
[0102] The above describes the embodiments of the present application from the perspective of the method. Next, the related devices in the embodiments of the present application are introduced from the perspective of the specific device implementation.
[0103] Please refer to Figure 8 , a schematic diagram of a data processing device 800 is provided in an embodiment of the present application. Among them, the data processing device 800 includes an acquisition unit 801, a screening unit 802, and a construction unit 803.
[0104] The acquisition unit 801 is configured to acquire the complete graph structure of the database. The database includes multiple models. The complete graph structure includes multiple graph structures. Each model in the multiple models corresponds to one graph structure in the multiple graph structures. Each graph structure in the multiple graph structures includes a sub-graph structure, and the sub-graph structure is used to indicate multiple operators and the connection relationships between the multiple operators.
[0105] The screening unit 802 is configured to screen out a first sub-graph structure from the sub-graph structures included in the complete graph structure. The frequency of occurrence of the first sub-graph structure in the complete graph structure is greater than a preset threshold, and the first sub-graph structure is used to determine the target sub-graph structure.
[0106] The construction unit 803 is configured to construct a sub-graph level Benchmark library based on the target sub-graph structure.
[0107] Optionally, the screening unit 802 is further configured to screen out a first type of models from the database. The first type of models includes multiple models, and the multiple models included in the first type of models are of the same type. The data processing device 800 further includes a determination unit 804, configured to determine a first type of graph structure corresponding to the first type of models in the complete graph structure. The screening unit 802 is further configured to screen out a second sub-graph structure from the sub-graph structures included in the first type of graph structure. The frequency of occurrence of the second sub-graph structure in the first type of graph structure is greater than a preset threshold, and the second sub-graph structure is used to determine the target sub-graph structure.
[0108] Optionally, the sub-graph structure is counted once in each graph structure in which it appears.
[0109] Optionally, the screening unit 802 is further configured to screen out a first model from the database. The first model is a model that meets the importance condition. The determination unit 804 is further configured to determine a first graph structure corresponding to the first model in the complete graph structure. The screening unit 802 is further configured to screen out a third sub-graph structure from the sub-graph structures included in the first graph structure. The frequency of occurrence of the third sub-graph structure in the first graph structure is greater than a preset threshold, and the third sub-graph structure is used to determine the target sub-graph structure.
[0110] Optionally, the data processing device 800 further includes a merging unit 805, configured to perform a merging process on the first sub-graph structure, the second sub-graph structure, and the third sub-graph structure to obtain a first type of sub-graph structure. A deletion unit 806 is configured to delete the duplicate sub-graph structures in the first type of sub-graph structure to obtain the target sub-graph structure.
[0111] Optionally, the obtaining unit 801 is specifically configured to convert multiple models into a complete graph structure through a graph generation tool.
[0112] Optionally, the data processing device 800 further includes a conversion unit 807, configured to convert the target sub-graph structure into a target model.
[0113] Optionally, the data processing device 800 further includes an extraction unit 808, configured to extract features of the target model. A classification unit 809 is configured to classify the target sub-graph structure based on the features.
[0114] Each unit in the data processing device 800 performs the operations of the data processing device in the foregoing Figure 1 and Figure 7 illustrated embodiments, and details are not described herein again.
[0115] Please refer to the following Figure 9, which is a possible structural schematic diagram of the data processing device 900 provided by the embodiments of the present application, includes a processor 901, a communication interface 902, a memory 903, and a bus 904. The processor 901, the communication interface 902, and the memory 903 are interconnected through the bus 904. In the embodiments of the present application, the processor 901 is used to control and manage the operations of the data processing device. For example, the processor 901 is used to execute Figure 1 the steps executed by the data processing device in the method embodiments shown. The communication interface 902 is used to support the data processing device to communicate. The memory 903 is used to store the program code and data of the data processing device.
[0116] Among them, the processor 901 may be a central processing unit, a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field programmable gate array, or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logic blocks, modules, and circuits described in combination with the disclosure of the present application. The processor may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a digital signal processor and a microprocessor, and so on. The bus 904 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 9 only a thick line is shown in, but it does not mean that there is only one bus or one type of bus.
[0117] The embodiments of the present application also provide a computer-readable storage medium. The computer-readable storage medium includes instructions. When the instructions run on a computer, the computer is caused to execute the foregoing Figure 1 and Figure 7 methods in the embodiments shown.
[0118] The embodiments of the present application also provide a computer program product containing instructions. When the computer program product runs on a computer, the computer is caused to execute the foregoing Figure 1 and Figure 7 methods in the embodiments shown.
[0119] The embodiments of the present application also provide a chip system. The chip system includes at least one processor and a communication interface. The communication interface and the at least one processor are interconnected by a line. The at least one processor is used to run a computer program or instructions to execute the foregoing Figure 1 and Figure 7 methods in the embodiments shown.
[0120] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.
[0121] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.
[0122] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0123] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0124] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0125] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, read-only memory), random access memories (RAM, random access memory), magnetic disks, or optical discs that can store program codes.
Claims
1. A data processing method, characterized in that, Including: Obtain the complete graph structure of the database, where the database includes multiple models, the complete graph structure includes multiple graph structures, each model in the multiple models corresponds to one graph structure in the multiple graph structures, and each graph structure in the multiple graph structures includes a sub-graph structure, and the sub-graph structure is used to indicate multiple operators and the connection relationships between the multiple operators; Screen out a first sub-graph structure from the sub-graph structures included in the complete graph structure, where the frequency of occurrence of the first sub-graph structure in the complete graph structure is greater than a preset threshold, and the first sub-graph structure is used to determine the target sub-graph structure; Construct a sub-graph level Benchmark library based on the target sub-graph structure.
2. The method according to claim 1, characterized in that, The method further includes: Screen out a first type of models from the database, where the first type of models includes multiple models, and the multiple models included in the first type of models are of the same type; Determine a first type of graph structures corresponding to the first type of models in the complete graph structure; Screen out a second sub-graph structure from the sub-graph structures included in the first type of graph structures, where the frequency of occurrence of the second sub-graph structure in the first type of graph structures is greater than a preset threshold, and the second sub-graph structure is used to determine the target sub-graph structure.
3. The method according to claim 1 or 2, characterized in that, The sub-graph structure is counted once in each graph structure where it appears.
4. The method according to any one of claims 1 to 3, characterized in that, The method further includes: Screen out a first model from the database, where the first model is a model that meets the importance condition; Determine a first graph structure corresponding to the first model in the complete graph structure; Screen out a third sub-graph structure from the sub-graph structures included in the first graph structure, where the frequency of occurrence of the third sub-graph structure in the first graph structure is greater than a preset threshold, and the third sub-graph structure is used to determine the target sub-graph structure.
5. The method according to claim 4, characterized in that, The method further includes: Perform a merging process on the first sub-graph structure, the second sub-graph structure, and the third sub-graph structure to obtain a first type of sub-graph structure; Delete the duplicate sub-graph structures in the first type of sub-graph structure to obtain the target sub-graph structure.
6. The method according to any one of claims 1 to 5, characterized in that, The obtaining of the complete graph structure of the database includes: Convert the multiple models into the complete graph structure through a graph generation tool.
7. The method according to any one of claims 1 to 6, characterized in that, The method further includes: Convert the target sub-graph structure into a target model.
8. The method according to claim 7, characterized in that, The method further includes: Extract the features of the target model; Classify the target sub-graph structure based on the features.
9. A data processing device, characterized in that, Including: An obtaining unit, configured to obtain the complete graph structure of the database, where the database includes multiple models, the complete graph structure includes multiple graph structures, each model in the multiple models corresponds to one graph structure in the multiple graph structures, and each graph structure in the multiple graph structures includes a sub-graph structure, and the sub-graph structure is used to indicate multiple operators and the connection relationships between the multiple operators; A screening unit, configured to screen out a first sub-graph structure from the sub-graph structures included in the complete graph structure, where the frequency of occurrence of the first sub-graph structure in the complete graph structure is greater than a preset threshold, and the first sub-graph structure is used to determine the target sub-graph structure; A constructing unit, configured to construct a sub-graph level Benchmark library based on the target sub-graph structure.
10. The device according to claim 9, characterized in that, The screening unit is further configured to: screen out a first type of models from the database, the first type of models includes a plurality of models, and the plurality of models included in the first type of models are of the same type; The apparatus further includes a determining unit, configured to determine a first type of graph structure corresponding to the first type of models in the complete graph structure; The screening unit is further configured to: screen out a second sub-graph structure from the sub-graph structures included in the first type of graph structure, the frequency of occurrence of the second sub-graph structure in the first type of graph structure is greater than a preset threshold, and the second sub-graph structure is used to determine the target sub-graph structure.
11. The device according to claim 9 or 10, characterized in that, The sub-graph structure is counted once in each graph structure where it appears.
12. The device according to any one of claims 9 to 11, characterized in that, The screening unit is further configured to: screen out a first model from the database, the first model is a model that meets the importance condition; The determining unit is further configured to: determine a first graph structure corresponding to the first model in the complete graph structure; The screening unit is further configured to: screen out a third sub-graph structure from the sub-graph structures included in the first graph structure, the frequency of occurrence of the third sub-graph structure in the first graph structure is greater than a preset threshold, and the third sub-graph structure is used to determine the target sub-graph structure.
13. The device according to claim 12, characterized in that, The apparatus further includes: a merging unit, configured to perform a merging process on the first sub-graph structure, the second sub-graph structure, and the third sub-graph structure to obtain a first type of sub-graph structure; a deleting unit, configured to delete the duplicate sub-graph structures in the first type of sub-graph structure to obtain the target sub-graph structure.
14. The device according to any one of claims 9 to 13, characterized in that, The obtaining unit is specifically configured to: convert the plurality of models into the complete graph structure through a graph generation tool.
15. The device according to any one of claims 9 to 14, characterized in that, The apparatus further includes: a conversion unit, configured to convert the target sub-graph structure into a target model.
16. The device according to claim 15, characterized in that, The apparatus further includes: an extraction unit, configured to extract features of the target model; a classification unit, configured to classify the target sub-graph structure based on the features.
17. A data processing device, characterized in that, including: a processor and a memory; the memory is used to store instructions; the processor is used to execute the instructions stored in the memory to implement the method according to any one of claims 1 to 8.
18. A computer-readable storage medium, on which a computer program is stored, characterized in that, The computer program, when executed by one or more processors, implements the method according to any one of claims 1 to 8.
19. A computer program product comprising instructions, characterized in that, When the computer program product runs on a computer, the computer is caused to execute the method according to any one of claims 1 to 8.