Fusion operator testing method, device, computer equipment, storage medium and computer program product

Through automated graph structure comparison and acceleration ratio calculation, the problem of inefficiency of traditional manual testing methods is solved, and efficient fusion operator testing and performance evaluation are achieved.

CN118968246BActive Publication Date: 2025-05-13BEIJING QINGCHENG JIZHI TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411090802.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-09
Publication Date
2025-05-13
Estimated Expiration
2044-08-09

AI Technical Summary

Technical Problem

Traditional manual testing methods lead to low testing efficiency of fusion operators, time-consuming and labor-intensive.

Method used

By obtaining the target graph structure and historical graph structure of the deep learning model to be constructed, the fusion operator and historical operator are determined, the output time of the two under the same input parameters is compared, the acceleration ratio is calculated, and the performance improvement of the fusion operator is evaluated.

Benefits of technology

Automatic testing is realized, the testing efficiency of the fusion operator is improved, manual intervention is reduced, and the performance improvement of the new operator can be accurately evaluated.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118968246B_ABST
    Figure CN118968246B_ABST
Patent Text Reader

Abstract

The present application relates to a fusion operator testing method, device, computer equipment, storage medium and computer program product. The method includes: obtaining a target graph structure of a deep learning model to be constructed, and obtaining a historical graph structure of a historical deep learning model corresponding to the deep learning model to be constructed; determining a fusion operator in the target graph structure, and a historical operator corresponding to the fusion operator in the historical graph structure; inputting a first input parameter associated with the deep learning model to be constructed into the fusion operator to obtain a first output result and a first output time corresponding to the first output result, and inputting the first input parameter into the historical operator to obtain a second output result and a second output time corresponding to the second output result; determining the acceleration ratio of the fusion operator based on the first output time and the second output time; and determining the first target test result of the fusion operator based on the acceleration ratio. The present method can improve the test efficiency of the fusion operator.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a fusion operator testing method, apparatus, computer equipment, computer-readable storage medium, and computer program product. Background Art

[0002] In the field of model optimization, testing the fusion operators in the optimized model is crucial to improving the computing performance of the model.

[0003] In traditional technology, manual testing is generally used to test fusion operators; however, this method is cumbersome and easily consumes a lot of time and manpower, resulting in low testing efficiency of fusion operators. Summary of the invention

[0004] Based on this, it is necessary to provide a fusion operator testing method, apparatus, computer equipment, computer-readable storage medium and computer program product that can improve the testing efficiency of the fusion operator in response to the above technical problems.

[0005] In a first aspect, the present application provides a fusion operator testing method, comprising:

[0006] Obtaining a target graph structure of a deep learning model to be constructed, and obtaining a historical graph structure of a historical deep learning model corresponding to the deep learning model to be constructed;

[0007] According to the target graph structure and the historical graph structure, determining a fusion operator in the target graph structure and a historical operator in the historical graph structure corresponding to the fusion operator;

[0008] Inputting a first input parameter associated with the deep learning model to be constructed into the fusion operator to obtain a first output result and a first output time corresponding to the first output result, and inputting the first input parameter into the history operator to obtain a second output result and a second output time corresponding to the second output result;

[0009] Determine a speedup ratio of the fusion operator according to the first output time and the second output time;

[0010] According to the acceleration ratio of the fusion operator, a first target test result of the fusion operator is determined.

[0011] In one embodiment, determining the speedup ratio of the fusion operator according to the first output time and the second output time includes:

[0012] Obtaining a historical processing time of the historical deep learning model; the historical processing time is obtained by inputting the first input parameter into the historical deep learning model;

[0013] A speedup ratio of the fusion operator is determined according to the number of the fusion operators, the first output time, the second output time, and the historical processing time.

[0014] In one embodiment, before determining the first target test result of the fusion operator according to the speedup ratio of the fusion operator, the method further includes:

[0015] Determining a processing accuracy of the fusion operator according to the first output result and the second output result;

[0016] Determining a first target test result of the fusion operator according to the acceleration ratio of the fusion operator includes:

[0017] According to the acceleration ratio and processing accuracy of the fusion operator, a first target test result of the fusion operator is determined.

[0018] In one embodiment, determining, according to the target graph structure and the history graph structure, a fusion operator in the target graph structure and a history operator in the history graph structure corresponding to the fusion operator includes:

[0019] Parsing the target graph structure to obtain a parsed target graph structure, and parsing the history graph structure to obtain a parsed history graph structure;

[0020] Comparing the first operator in the parsed target graph structure with the second operator in the parsed history graph structure, obtaining a target operator in the first operator that is not repeated with the second operator as a fusion operator in the target graph structure;

[0021] Determine a historical operator in the historical graph structure that corresponds to the fusion operator.

[0022] In one embodiment, after determining the first target test result of the fusion operator according to the speedup ratio of the fusion operator, the method further includes:

[0023] When the target test result satisfies the preset test result, constructing the deep learning model to be constructed according to the target graph structure;

[0024] Inputting a second input parameter associated with the deep learning model to be constructed into the deep learning model to be constructed, obtaining a third output result and a third output time corresponding to the third output result, and inputting the second input parameter into the historical deep learning model, obtaining a fourth output result and a fourth output time corresponding to the fourth output result;

[0025] A second target test result of the deep learning model to be constructed is determined according to the third output time and the fourth output time.

[0026] In one embodiment, constructing the deep learning model to be constructed according to the target graph structure includes:

[0027] Obtaining weight information of the historical deep learning model;

[0028] Construct the deep learning model to be constructed according to the target graph structure and the weight information.

[0029] In a second aspect, the present application also provides a fusion operator testing device, including:

[0030] A structure acquisition module, used to acquire a target graph structure of a deep learning model to be constructed, and to acquire a historical graph structure of a historical deep learning model corresponding to the deep learning model to be constructed;

[0031] An operator determination module, used to determine, according to the target graph structure and the historical graph structure, a fusion operator in the target graph structure and a historical operator in the historical graph structure corresponding to the fusion operator;

[0032] An operator processing module, configured to input a first input parameter associated with the deep learning model to be constructed into the fusion operator to obtain a first output result and a first output time corresponding to the first output result, and to input the first input parameter into the history operator to obtain a second output result and a second output time corresponding to the second output result;

[0033] A speedup ratio determination module, configured to determine a speedup ratio of the fusion operator according to the first output time and the second output time;

[0034] A result determination module is used to determine a first target test result of the fusion operator according to the acceleration ratio of the fusion operator.

[0035] In a third aspect, the present application further provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the following steps are implemented:

[0036] Obtaining a target graph structure of a deep learning model to be constructed, and obtaining a historical graph structure of a historical deep learning model corresponding to the deep learning model to be constructed;

[0037] According to the target graph structure and the historical graph structure, determining a fusion operator in the target graph structure and a historical operator in the historical graph structure corresponding to the fusion operator;

[0038] Inputting a first input parameter associated with the deep learning model to be constructed into the fusion operator to obtain a first output result and a first output time corresponding to the first output result, and inputting the first input parameter into the history operator to obtain a second output result and a second output time corresponding to the second output result;

[0039] Determine a speedup ratio of the fusion operator according to the first output time and the second output time;

[0040] According to the acceleration ratio of the fusion operator, a first target test result of the fusion operator is determined.

[0041] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the following steps are implemented:

[0042] Obtaining a target graph structure of a deep learning model to be constructed, and obtaining a historical graph structure of a historical deep learning model corresponding to the deep learning model to be constructed;

[0043] According to the target graph structure and the historical graph structure, determining a fusion operator in the target graph structure and a historical operator in the historical graph structure corresponding to the fusion operator;

[0044] Inputting a first input parameter associated with the deep learning model to be constructed into the fusion operator to obtain a first output result and a first output time corresponding to the first output result, and inputting the first input parameter into the history operator to obtain a second output result and a second output time corresponding to the second output result;

[0045] Determine a speedup ratio of the fusion operator according to the first output time and the second output time;

[0046] According to the acceleration ratio of the fusion operator, a first target test result of the fusion operator is determined.

[0047] In a fifth aspect, the present application further provides a computer program product, including a computer program, which implements the following steps when executed by a processor:

[0048] Obtaining a target graph structure of a deep learning model to be constructed, and obtaining a historical graph structure of a historical deep learning model corresponding to the deep learning model to be constructed;

[0049] According to the target graph structure and the historical graph structure, determining a fusion operator in the target graph structure and a historical operator in the historical graph structure corresponding to the fusion operator;

[0050] Inputting a first input parameter associated with the deep learning model to be constructed into the fusion operator to obtain a first output result and a first output time corresponding to the first output result, and inputting the first input parameter into the history operator to obtain a second output result and a second output time corresponding to the second output result;

[0051] Determine a speedup ratio of the fusion operator according to the first output time and the second output time;

[0052] According to the acceleration ratio of the fusion operator, a first target test result of the fusion operator is determined.

[0053] The above-mentioned fusion operator testing method, apparatus, computer equipment, storage medium and computer program product first obtain the target graph structure of the deep learning model to be constructed, and obtain the historical graph structure of the historical deep learning model corresponding to the deep learning model to be constructed, and then determine the fusion operator in the target graph structure and the historical operator corresponding to the fusion operator in the historical graph structure based on the target graph structure and the historical graph structure. Then, input the first input parameter associated with the deep learning model to be constructed into the fusion operator to obtain a first output result and a first output time corresponding to the first output result, and input the first input parameter into the historical operator to obtain a second output result and a second output time corresponding to the second output result. Then, determine the acceleration ratio of the fusion operator based on the first output time and the second output time. Finally, determine the first target test result of the fusion operator based on the acceleration ratio of the fusion operator. In this way, when testing the fusion operator, by comparing the output time of the fusion operator and the historical operator, the acceleration ratio of the fusion operator can be accurately determined, so that the performance improvement of the new fusion operator when processing the same input parameters can be accurately evaluated; moreover, the entire process does not require human intervention, avoiding the defect of low test efficiency of the fusion operator caused by manual testing, which is cumbersome and easy to spend a lot of time and manpower, thereby improving the test efficiency of the fusion operator. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the drawings required for use in the embodiments of the present application or related technical descriptions will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0055] Figure 1 A schematic diagram of a flow chart of a fusion operator testing method in one embodiment;

[0056] Figure 2 A schematic flow chart of the steps of determining the speedup ratio of a fusion operator in one embodiment;

[0057] Figure 3 A schematic diagram of a flow chart of a fusion operator testing method in another embodiment;

[0058] Figure 4 A flowchart of a method for optimizing the fusion operator test performance and correctness of a general deep learning model in one embodiment;

[0059] Figure 5 is a structural block diagram of a fusion operator testing device in an embodiment;

[0060] Figure 6 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0061] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0062] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0063] In an exemplary embodiment, Figure 1As shown, a fusion operator testing method is provided. This embodiment uses the method applied to a server as an example for illustration; it can be understood that the method can also be applied to a terminal, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. Among them, the terminal can be but is not limited to various personal computers, laptops, smart phones and tablets; the server can be implemented with an independent server or a server cluster composed of multiple servers. In this embodiment, the method includes the following steps:

[0064] Step S101, obtaining a target graph structure of a deep learning model to be constructed, and obtaining a historical graph structure of a historical deep learning model corresponding to the deep learning model to be constructed.

[0065] The deep learning model to be constructed may refer to an image processing model used to identify the image type of the image to be processed, and the deep learning model to be constructed may also refer to a text processing model used to identify the sentiment type of the text to be processed. It should be noted that the deep learning model to be constructed includes two model formats: ONNX (Open Neural Network Exchange) and PyTorch (a deep learning framework). In actual scenarios, the deep learning model to be constructed refers to a deep learning model optimized by a fusion operator.

[0066] The fusion operator refers to an operator combination obtained by fusing at least two operators in a deep learning model. For example, in an image processing model, the fusion operator refers to an operator combination corresponding to the combination of a convolution operator and a pooling operator; in a text processing model, the fusion operator refers to an operator combination corresponding to the combination of a feature extraction operator and a concatenation operator.

[0067] Among them, the image processing model includes the transformer model.

[0068] Among them, the text processing model includes the GPT-2 (Generative Pretrained Transformer 2) model.

[0069] Among them, the target graph structure refers to the graph structure of the deep learning model to be constructed.

[0070] The graph structure includes node information and edge information. It should be noted that the format of the graph structure is text.

[0071] The historical deep learning model refers to the original model corresponding to the deep learning model to be constructed. In actual scenarios, the deep learning model to be constructed refers to the deep learning model that has not been optimized by the fusion operator. It should be noted that the historical deep learning model is also called the original model.

[0072] Among them, the historical graph structure refers to the graph structure of the historical deep learning model.

[0073] Exemplarily, in response to a test instruction for a fusion operator, the server determines a deep learning model corresponding to the fusion operator as the deep learning model to be constructed; then, the server obtains a graph structure in the form of an adjacency matrix of the deep learning model to be constructed; then, the server determines the nodes in the graph structure (e.g., node 1, node 2, node 3, node 4) and element information of the nodes in the graph structure (e.g., the element in the 1st row and 2nd column is 1) based on the graph structure in the form of the adjacency matrix, and determines the connection information between the nodes in the graph structure based on the element information of the nodes in the graph structure (e.g., the element in the 1st row and 2nd column is 1, indicating that node 1 is connected to node 2); then, the server inputs the nodes in the graph structure and the connection information between the nodes in the graph structure into a GNN (Graph Neural Networks) model to obtain a graph structure in text form as the target graph structure of the deep learning model to be constructed; then, the server determines the original model corresponding to the deep learning model to be constructed as the historical deep learning model; then, the server obtains the graph structure of the historical deep learning model as the historical graph structure.

[0074] Step S102: According to the target graph structure and the historical graph structure, a fusion operator in the target graph structure and a historical operator corresponding to the fusion operator in the historical graph structure are determined.

[0075] Among them, the historical operator refers to the operator corresponding to the fusion operator in the historical graph structure.

[0076] Exemplarily, the server parses the target graph structure to obtain a parsed target graph structure, and parses the historical graph structure to obtain a parsed historical graph structure; then, the server determines the fusion operator in the target graph structure and the historical operator corresponding to the fusion operator in the historical graph structure based on the parsed target graph structure and the parsed historical graph structure.

[0077] Step S103: input the first input parameter associated with the deep learning model to be constructed into the fusion operator to obtain the first output result and the first output time corresponding to the first output result, and input the first input parameter into the history operator to obtain the second output result and the second output time corresponding to the second output result.

[0078] The first input parameter refers to an input parameter used to test the fusion operator. For example, in an image processing model, the first input parameter refers to an image to be processed; in a text processing model, the first input parameter refers to a text to be processed.

[0079] The first output result refers to the processing result obtained by inputting the first input parameter into the fusion operator. For example, in an image processing model, the first output result refers to the image type of the image to be processed; in a text processing model, the first output result refers to the sentiment type of the text to be processed.

[0080] The first output time refers to the time required for the fusion operator to process the first input parameter.

[0081] The second output result refers to the processing result obtained by inputting the first input parameter into the history operator. For example, in an image processing model, the second output result refers to the image type of the image to be processed; in a text processing model, the second output result refers to the sentiment type of the text to be processed.

[0082] The second output time refers to the time required for the history operator to process the first input parameter.

[0083] Exemplarily, the server automatically generates multiple groups of random input parameters as the first input parameters associated with the deep learning model to be constructed according to the model information of the deep learning model to be constructed; for example, the server automatically generates multiple groups of random input parameters as the first input parameters associated with the deep learning model to be constructed according to the dimension (such as a two-dimensional matrix or a three-dimensional tensor), length (the number of parameters) and value type (such as an integer or a floating-point number) of the deep learning model to be constructed; then, the server inputs the first input parameter associated with the deep learning model to be constructed into the fusion operator, processes the first input parameter through the fusion operator, and obtains the output result corresponding to the first input parameter as the first output result; then, the server obtains the input time of the first input parameter and the output time of the first output result, and according to the input time of the first input parameter and the output time of the first output result The server determines the output time of an output result, and the first output time corresponding to the first output result; for example, the server performs a difference processing on the input time of the first input parameter and the output time of the first output result to obtain the first output time corresponding to the first output result; then, the server inputs the first input parameter into the history operator, processes the first input parameter through the history operator, and obtains the output result corresponding to the first input parameter as the second output result; then, the server obtains the input time of the first input parameter and the output time of the second output result, and determines the second output time corresponding to the second output result according to the input time of the first input parameter and the output time of the second output result; for example, the server performs a difference processing on the input time of the first input parameter and the output time of the second output result to obtain the second output time corresponding to the second output result.

[0084] Step S104: determining a speedup ratio of the fusion operator according to the first output time and the second output time.

[0085] The speedup ratio is used to characterize the performance improvement of the fusion operator relative to the historical operator, such as 50%.

[0086] Exemplarily, the server obtains a target correspondence between the first output time, the second output time, and the acceleration ratio; then, the server queries the target correspondence based on the first output time and the second output time, and obtains the acceleration ratio corresponding to the first output time and the second output time as the acceleration ratio of the fusion operator.

[0087] Step S105, determining the first target test result of the fusion operator according to the acceleration ratio of the fusion operator.

[0088] Among them, the first target test result refers to the comprehensive test result of the fusion operator.

[0089] Exemplarily, the server obtains the operator level corresponding to the acceleration ratio of the fusion operator according to the acceleration ratio of the fusion operator and the correspondence between the acceleration ratio and the operator level as the operator level of the fusion operator; for example, the operator level corresponding to the acceleration ratio of 50% is good, and the operator level corresponding to the acceleration ratio of 10% is needs improvement, etc.; then, the server uses the operator level of the fusion operator as the first target test result of the fusion operator.

[0090] In the above-mentioned fusion operator testing method, the target graph structure of the deep learning model to be constructed and the historical graph structure of the historical deep learning model corresponding to the deep learning model to be constructed are first obtained, and then the fusion operator in the target graph structure and the historical operator corresponding to the fusion operator in the historical graph structure are determined according to the target graph structure and the historical graph structure. Next, the first input parameter associated with the deep learning model to be constructed is input into the fusion operator to obtain the first output result and the first output time corresponding to the first output result, and the first input parameter is input into the historical operator to obtain the second output result and the second output time corresponding to the second output result. Then, the acceleration ratio of the fusion operator is determined according to the first output time and the second output time. Finally, according to the acceleration ratio of the fusion operator, the first target test result of the fusion operator is determined. In this way, when testing the fusion operator, by comparing the output time of the fusion operator and the historical operator, the acceleration ratio of the fusion operator can be accurately determined, so that the performance improvement of the new fusion operator when processing the same input parameters can be accurately evaluated; moreover, the entire process does not require human intervention, avoiding the defect of low test efficiency of the fusion operator caused by manual testing, which is cumbersome and easy to spend a lot of time and manpower, thereby improving the test efficiency of the fusion operator.

[0091] In an exemplary embodiment, Figure 2As shown, the above step S104, determining the speedup ratio of the fusion operator according to the first output time and the second output time, specifically includes the following steps:

[0092] Step S201, obtain the historical processing time of the historical deep learning model; the historical processing time is obtained by inputting the first input parameter into the historical deep learning model.

[0093] Step S202, determining the speedup ratio of the fusion operator according to the number of fusion operators, the first output time, the second output time and the historical processing time.

[0094] Here, the historical processing time refers to the time required for the historical deep learning model to process the first input parameter.

[0095] Here, the number refers to the number of fusion operators in the target graph structure, such as 3.

[0096] Exemplarily, the server inputs the first input parameter into the historical deep learning model, processes the first input parameter through the historical deep learning model, and obtains the output result of the historical deep learning model and the output time corresponding to the output result; then, the server uses the output time as the historical processing time of the historical deep learning model; then, the server obtains the number of fusion operators in the target graph structure as the number of fusion operators; then, the server inputs the number of fusion operators, the first output time, the second output time and the historical processing time into the acceleration ratio prediction model, and obtains the acceleration ratio corresponding to the number of fusion operators, the first output time, the second output time and the historical processing time as the acceleration ratio of the fusion operator.

[0097] For example, the server can determine the speedup ratio of the fusion operator through the following formula:

[0098] , formula (1)

[0099] Among them, the number of fused operators refers to the number of fused operators, the operator time before fusion refers to the second output time, the operator time after fusion refers to the first output time, and the overall time of the original model refers to the historical processing time.

[0100] In this embodiment, by comprehensively considering a variety of different factors, the acceleration ratio of the fusion operator is obtained, so that the performance improvement of the fusion operator can be evaluated more accurately and comprehensively, which is conducive to determining the actual effect of the fusion operator on optimizing the deep learning model to be constructed, and providing a reliable basis for further improvement of the deep learning model to be constructed.

[0101] In an exemplary embodiment, the above step S105, before determining the first target test result of the fusion operator according to the acceleration ratio of the fusion operator, specifically includes the following content: determining the processing accuracy of the fusion operator according to the first output result and the second output result.

[0102] Then, the above step S105, determining the first target test result of the fusion operator according to the acceleration ratio of the fusion operator, specifically includes the following contents: determining the first target test result of the fusion operator according to the acceleration ratio and processing accuracy of the fusion operator.

[0103] The processing accuracy rate is used to characterize the accuracy of the processing of the first input parameter by the fusion operator.

[0104] Exemplarily, the server inputs multiple first input parameters into the fusion operator and the history operator respectively, and obtains the first output results and the second output results corresponding to the multiple first input parameters; then, the server compares the first output result and the second output result corresponding to each first input parameter to obtain multiple groups of comparison results, and determines the processing accuracy of the fusion operator based on the multiple groups of comparison results; for example, the server inputs 10 first input parameters into the fusion operator and the history operator respectively, and obtains the first output results and the second output results corresponding to the 10 first input parameters; then, the server compares the first output result and the second output result corresponding to each first input parameter to obtain 10 groups of comparison results, for example, 9 groups of comparison results are equal, and 1 group of comparison results are unequal, and based on these 10 groups of comparison results, it is determined that the processing accuracy of the fusion operator is 90%; finally, the server takes the acceleration ratio and processing accuracy of the fusion operator as the first target test result of the fusion operator.

[0105] In this embodiment, by combining the two key indicators of the acceleration ratio and processing accuracy of the fusion operator, the performance of the fusion operator can be evaluated more comprehensively and comprehensively, which is conducive to determining a fusion operator that is both efficient and accurate, thereby improving the overall performance of the deep learning model to be constructed.

[0106] In an exemplary embodiment, the above step S102, based on the target graph structure and the historical graph structure, determines the fusion operator in the target graph structure and the historical operator corresponding to the fusion operator in the historical graph structure, specifically including the following contents: parsing the target graph structure to obtain the parsed target graph structure, and parsing the historical graph structure to obtain the parsed historical graph structure; comparing the first operator in the parsed target graph structure with the second operator in the parsed historical graph structure to obtain the target operator in the first operator that is not repeated with the second operator as the fusion operator in the target graph structure; determining the historical operator corresponding to the fusion operator in the historical graph structure.

[0107] The parsed target graph structure refers to a plurality of operators obtained after parsing the target graph structure.

[0108] The parsed historical graph structure refers to a plurality of operators obtained after parsing the historical graph structure.

[0109] The first operator refers to the operator in the target graph structure after parsing.

[0110] The second operator refers to the operator in the parsed history graph structure.

[0111] The target operator refers to an operator in the first operator that is not repeated in the second operator.

[0112] Exemplarily, the server responds to the parsing instruction for the target graph structure and parses the target graph structure to obtain a parsed target graph structure; then, the server responds to the parsing instruction for the historical graph structure and parses the historical graph structure to obtain a parsed historical graph structure; then, the server compares the first operator in the parsed target graph structure with the second operator in the parsed historical graph structure to obtain a comparison result; then, the server extracts the operators in the first operator that are not repeated with the second operator from the comparison result as target operators; then, the server uses these target operators as fusion operators in the target graph structure; finally, the server determines the historical operators in the historical graph structure corresponding to the fusion operators based on the comparison result.

[0113] For example, the first operator includes operator 1, operator 2, and operator 3, and the second operator includes operator 1, operator 4, operator 5, and operator 3. After the server performs comparison processing, the comparison result is that operator 1 in the first operator corresponds to operator 1 in the second operator, operator 2 in the first operator corresponds to operators 4 and 5 in the second operator, and operator 3 in the first operator corresponds to operator 3 in the second operator; then, the server extracts the operator in the first operator that is not repeated in the second operator, that is, operator 2, and uses operator 2 as the fusion operator in the target graph structure; then, based on the comparison result, the server determines the historical operators in the historical graph structure that correspond to the fusion operator, that is, operator 4 and operator 5.

[0114] In this embodiment, by analyzing and comparing the target graph structure and the historical graph structure, non-repetitive target operators can be accurately found as fusion operators, which is beneficial to improving the accuracy of determining the fusion operator; moreover, this method does not require human intervention, which is beneficial to reducing the workload of manual screening and judgment, thereby improving the efficiency of determining the fusion operator.

[0115] In an exemplary embodiment, the above step S105, after determining the first target test result of the fusion operator according to the acceleration ratio of the fusion operator, specifically includes the following contents: when the target test result meets the preset test result, constructing the deep learning model to be constructed according to the target graph structure; inputting the second input parameter associated with the deep learning model to be constructed into the deep learning model to be constructed, obtaining the third output result and the third output time corresponding to the third output result, and inputting the second input parameter into the historical deep learning model, obtaining the fourth output result and the fourth output time corresponding to the fourth output result; determining the second target test result of the deep learning model to be constructed according to the third output time and the fourth output time.

[0116] The preset test result refers to a pre-set test result, which is used to judge the target test result, such as 30%. It should be noted that the preset test result may be determined according to the circumstances.

[0117] The second input parameter refers to an input parameter used to test the deep learning model to be constructed. For example, in an image processing model, the second input parameter refers to the image to be processed; in a text processing model, the second input parameter refers to the text to be processed.

[0118] The third output result refers to the processing result obtained by inputting the second input parameter into the deep learning model to be constructed. For example, in an image processing model, the third output result refers to the image type of the image to be processed; in a text processing model, the third output result refers to the sentiment type of the text to be processed.

[0119] The third output time refers to the time required for the deep learning model to be constructed to process the second input parameter.

[0120] The fourth output result refers to the processing result obtained by inputting the second input parameter into the historical deep learning model. For example, in an image processing model, the third output result refers to the image type of the image to be processed; in a text processing model, the third output result refers to the sentiment type of the text to be processed.

[0121] Among them, the fourth output time refers to the time required for the historical deep learning model to process the second input parameter.

[0122] Among them, the second target test result refers to the comprehensive test result of the deep learning model to be constructed.

[0123] Exemplarily, when the target test result meets the preset test result, the server constructs the deep learning model to be constructed according to the target graph structure; then, the server inputs the second input parameter associated with the deep learning model to be constructed into the deep learning model to be constructed, processes the second input parameter through the deep learning model to be constructed, and obtains the output result corresponding to the second input parameter as the third output result; then, the server obtains the input time of the second input parameter and the output time of the third output result, and determines the third output time corresponding to the third output result according to the input time of the second input parameter and the output time of the third output result; for example, the server performs a difference processing on the input time of the second input parameter and the output time of the third output result to obtain the third output time corresponding to the third output result; then, the server inputs the second input parameter and the output time of the third output result The server inputs an input parameter into the historical deep learning model, processes the second input parameter through the historical deep learning model, and obtains an output result corresponding to the second input parameter as a fourth output result; then, the server obtains the input time of the second input parameter and the output time of the fourth output result, and determines a fourth output time corresponding to the fourth output result according to the input time of the second input parameter and the output time of the fourth output result; for example, the server performs a difference processing on the input time of the second input parameter and the output time of the fourth output result to obtain a fourth output time corresponding to the fourth output result; then, the server performs a difference processing on the third output time and the fourth output time to obtain the time difference as the optimized duration of the deep learning model to be constructed; then, the server uses the optimized duration as the second target test result of the deep learning model to be constructed.

[0124] In this embodiment, the deep learning model to be constructed is constructed only when the target test results meet the preset conditions, thereby avoiding unnecessary model construction attempts and saving computing resources and time; moreover, by inputting the same input parameters into the deep learning model to be constructed and the historical deep learning model respectively, and comparing the output time, the performance improvement of the deep learning model to be constructed can be more accurately evaluated.

[0125] In an exemplary embodiment, a deep learning model to be constructed is constructed according to a target graph structure, which specifically includes the following contents: obtaining weight information of a historical deep learning model; and constructing a deep learning model to be constructed according to the target graph structure and the weight information.

[0126] Among them, weight information is used to characterize the strength parameters of the connections between neurons in the historical deep learning model.

[0127] Exemplarily, the server obtains weight information of historical deep learning models from a database; then, the server determines the layers of the model, such as convolutional layers, pooling layers, fully connected layers, etc., based on the target graph structure; then, the server loads the weight information into the layers corresponding to the weight information to obtain an initial deep learning model; then, the server uses a validation set to test the initial deep learning model to obtain test results of the initial deep learning model, such as the accuracy, recall rate, mean square error, loss value, and running time of the initial deep learning model; finally, when the test results meet the preset conditions (for example, the accuracy of the initial deep learning model is 92%, which is greater than the preset accuracy of 90%), the server uses the initial deep learning model that meets the preset conditions as the deep learning model to be constructed.

[0128] In this embodiment, by utilizing the weight information of the historical deep learning model, the existing resources can be migrated to the deep learning model to be constructed, avoiding the need to build the deep learning model from scratch, which is beneficial to saving a lot of time and computing resources.

[0129] In an exemplary embodiment, Figure 3 As shown, another fusion operator testing method is provided, and the method is applied to a server as an example, including the following steps:

[0130] Step S301, obtaining a target graph structure of a deep learning model to be constructed, and obtaining a historical graph structure of a historical deep learning model corresponding to the deep learning model to be constructed.

[0131] Step S302: parse the target graph structure to obtain a parsed target graph structure, and parse the history graph structure to obtain a parsed history graph structure.

[0132] Step S303 : Compare the first operator in the parsed target graph structure with the second operator in the parsed history graph structure to obtain a target operator in the first operator that is not repeated with the second operator as a fusion operator in the target graph structure.

[0133] Step S304: determine the historical operator corresponding to the fusion operator in the historical graph structure.

[0134] Step S305: input the first input parameter associated with the deep learning model to be constructed into the fusion operator to obtain the first output result and the first output time corresponding to the first output result, and input the first input parameter into the history operator to obtain the second output result and the second output time corresponding to the second output result.

[0135] Step S306, obtaining the historical processing time of the historical deep learning model; the historical processing time is obtained by inputting the first input parameter into the historical deep learning model.

[0136] Step S307, determining the speedup ratio of the fusion operator according to the number of fusion operators, the first output time, the second output time and the historical processing time.

[0137] Step S308: Determine the processing accuracy of the fusion operator according to the first output result and the second output result.

[0138] Step S309, determining the first target test result of the fusion operator according to the acceleration ratio and processing accuracy of the fusion operator.

[0139] In the above-mentioned fusion operator testing method, when testing the fusion operator, by comparing the output time of the fusion operator and the historical operator, the acceleration ratio of the fusion operator can be accurately determined, so that the performance improvement of the new fusion operator when processing the same input parameters can be accurately evaluated; moreover, the entire process does not require human intervention, avoiding the defect of low testing efficiency of the fusion operator caused by manual testing, which is cumbersome and easy to spend a lot of time and manpower, thereby improving the testing efficiency of the fusion operator.

[0140] In an exemplary embodiment, in order to more clearly illustrate the fusion operator testing method provided in the embodiment of the present application, the fusion operator testing method is specifically described below with a specific embodiment. Figure 4 As shown, the present application also provides a method for optimizing the performance and correctness of the fusion operator test of a general deep learning model. When testing the fusion operator, first obtain the target graph structure of the deep learning model to be constructed, and obtain the historical graph structure of the historical deep learning model corresponding to the deep learning model to be constructed. Then, based on the target graph structure and the historical graph structure, determine the fusion operator in the target graph structure, and the historical operator corresponding to the fusion operator in the historical graph structure. Then, input the first input parameter associated with the deep learning model to be constructed into the fusion operator to obtain the first output result and the first output time corresponding to the first output result, and input the first input parameter into the historical operator to obtain the second output result and the second output time corresponding to the second output result. Then, based on the first output time and the second output time, determine the acceleration ratio of the fusion operator. Finally, based on the acceleration ratio of the fusion operator, determine the first target test result of the fusion operator. Specifically including the following contents:

[0141] This application implements the fusion operator optimization test performance and correctness of the general deep learning model through the following steps:

[0142] 1. Write the original model file and the fused graph structure into the configuration area in text form, and write the custom fusion operator into the custom operator area.

[0143] 2. The system reads the input model file and the fused graph structure text information for graph analysis.

[0144] 3. When an operator in a custom operator area is scanned in the graph structure, multiple sets of random input parameters are automatically generated based on the input parameters, the custom operator is called, and the results and performance are compared with the subgraph before fusion, and the results are printed.

[0145] 4. After completing the performance and correctness verification of all custom operators, the system tests the model as a whole.

[0146] The configuration area includes: the original model file, including ONNX, PyTorch and other formats; the fused graph structure, which is a text description of the structure of the model after fusion; and the custom fusion operator, which is a text description of the custom fusion operator.

[0147] Among them, graph parsing includes: the system reads and parses the input model file and the fused graph structure; identifies the operators in the custom operator area, and automatically generates multiple sets of random input parameters.

[0148] The performance and correctness comparison includes: calling the custom operator and comparing its results and performance with the subgraph before fusion; printing and recording the comparison results.

[0149] The overall model test includes: calling the original model, obtaining the model output and calculation time; building a new model based on the fused graph-text structure, and writing the original model weights; running the new model, obtaining the output and calculation time, and comparing them with the original model; analyzing the performance improvement of each fusion operator and its share of time consumption in the overall operation of the model, and calculating the acceleration ratio.

[0150] The speedup ratio formula is as follows:

[0151] , formula (1)

[0152] The data flow description includes: data generation, that is, the system reads the original model file and the fused graph structure; data processing, including parsing the graph structure, identifying custom operators, generating random input parameters, and calling custom operators for testing; data output, including output comparison results and performance analysis reports.

[0153] Among them, the application scenarios of this application: This technical solution can be widely used in scenarios such as deep learning model optimization, operator fusion performance improvement, etc., and is particularly suitable for deep learning models that need to process multiple data formats (such as ONNX, PyTorch).

[0154] In the above embodiment, when testing the fusion operator, by comparing the output time of the fusion operator and the historical operator, the acceleration ratio of the fusion operator can be accurately determined, so that the performance improvement of the new fusion operator when processing the same input parameters can be accurately evaluated; moreover, the entire process does not require human intervention, avoiding the defect of low test efficiency of the fusion operator caused by manual testing, which is cumbersome and easy to spend a lot of time and manpower, thereby improving the test efficiency of the fusion operator.

[0155] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.

[0156] Based on the same inventive concept, the embodiment of the present application also provides a fusion operator testing device for implementing the fusion operator testing method involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in one or more fusion operator testing device embodiments provided below can refer to the limitations on the fusion operator testing method above, and will not be repeated here.

[0157] In an exemplary embodiment, Figure 5 As shown, a fusion operator testing device is provided, including: a structure acquisition module 501, an operator determination module 502, an operator processing module 503, a speedup ratio determination module 504 and a result determination module 505, wherein:

[0158] The structure acquisition module 501 is used to obtain the target graph structure of the deep learning model to be constructed, and to obtain the historical graph structure of the historical deep learning model corresponding to the deep learning model to be constructed.

[0159] The operator determination module 502 is used to determine the fusion operator in the target graph structure and the historical operator corresponding to the fusion operator in the historical graph structure according to the target graph structure and the historical graph structure.

[0160] The operator processing module 503 is used to input the first input parameter associated with the deep learning model to be constructed into the fusion operator to obtain the first output result and the first output time corresponding to the first output result, and to input the first input parameter into the history operator to obtain the second output result and the second output time corresponding to the second output result.

[0161] The speedup ratio determination module 504 is used to determine the speedup ratio of the fusion operator according to the first output time and the second output time.

[0162] The result determination module 505 is used to determine the first target test result of the fusion operator according to the acceleration ratio of the fusion operator.

[0163] In an exemplary embodiment, the acceleration ratio determination module 504 is also used to obtain the historical processing time of the historical deep learning model; the historical processing time is obtained by inputting the first input parameter into the historical deep learning model; the acceleration ratio of the fusion operator is determined according to the number of fusion operators, the first output time, the second output time and the historical processing time.

[0164] In an exemplary embodiment, the fusion operator testing device further includes a correctness determination module for determining the processing correctness of the fusion operator based on the first output result and the second output result.

[0165] In an exemplary embodiment, the result determination module 505 is further used to determine the first target test result of the fusion operator according to the acceleration ratio and processing accuracy of the fusion operator.

[0166] In an exemplary embodiment, the operator determination module 502 is also used to parse the target graph structure to obtain a parsed target graph structure, and to parse the historical graph structure to obtain a parsed historical graph structure; compare the first operator in the parsed target graph structure with the second operator in the parsed historical graph structure to obtain a target operator in the first operator that is not repeated with the second operator as a fusion operator in the target graph structure; and determine the historical operator in the historical graph structure that corresponds to the fusion operator.

[0167] In an exemplary embodiment, the fusion operator testing device also includes a model testing module, which is used to construct a deep learning model to be constructed according to a target graph structure when the target test result meets the preset test result; input a second input parameter associated with the deep learning model to be constructed into the deep learning model to be constructed to obtain a third output result and a third output time corresponding to the third output result, and input the second input parameter into the historical deep learning model to obtain a fourth output result and a fourth output time corresponding to the fourth output result; determine the second target test result of the deep learning model to be constructed based on the third output time and the fourth output time.

[0168] In an exemplary embodiment, the model testing module is also used to obtain weight information of historical deep learning models; and construct a deep learning model to be constructed based on the target graph structure and weight information.

[0169] Each module in the above-mentioned fusion operator testing device can be implemented in whole or in part by software, hardware and a combination thereof. Each of the above-mentioned modules can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute the operations corresponding to each of the above modules.

[0170] In an exemplary embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as shown in FIG. Figure 6 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. Among them, the processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data such as a historical graph structure and a first input parameter. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a fusion operator testing method is implemented.

[0171] Those skilled in the art will understand that Figure 6 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0172] In an exemplary embodiment, a computer device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above-mentioned method embodiments when executing the computer program.

[0173] In an exemplary embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.

[0174] In an exemplary embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0175] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but are not limited to this.

[0176] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0177] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.

Claims

1. A fusion operator testing method, characterized in that: The method comprises: Obtaining a target graph structure of a deep learning model to be constructed, and obtaining a historical graph structure of a historical deep learning model corresponding to the deep learning model to be constructed; According to the target graph structure and the historical graph structure, determining a fusion operator in the target graph structure and a historical operator in the historical graph structure corresponding to the fusion operator; Inputting a first input parameter associated with the deep learning model to be constructed into the fusion operator to obtain a first output result and a first output time corresponding to the first output result, and inputting the first input parameter into the history operator to obtain a second output result and a second output time corresponding to the second output result; Determine a speedup ratio of the fusion operator according to the first output time and the second output time; the speedup ratio of the fusion operator is used to represent the speedup ratio corresponding to the number of the fusion operators, the first output time, the second output time, and the historical processing time, and the historical processing time is obtained by inputting the first input parameter into the historical deep learning model; According to the acceleration ratio of the fusion operator, a first target test result of the fusion operator is determined.

2. The method according to claim 1, characterized in that: Determining the speedup ratio of the fusion operator according to the first output time and the second output time includes: Obtaining the historical processing time of the historical deep learning model; A speedup ratio of the fusion operator is determined according to the number of the fusion operators, the first output time, the second output time, and the historical processing time.

3. The method according to claim 1, characterized in that Before determining the first target test result of the fusion operator according to the acceleration ratio of the fusion operator, the method further includes: Determining a processing accuracy of the fusion operator according to the first output result and the second output result; Determining a first target test result of the fusion operator according to the acceleration ratio of the fusion operator includes: According to the acceleration ratio and processing accuracy of the fusion operator, a first target test result of the fusion operator is determined.

4. The method according to claim 1, characterized in that: The step of determining, according to the target graph structure and the historical graph structure, a fusion operator in the target graph structure and a historical operator in the historical graph structure corresponding to the fusion operator comprises: Parsing the target graph structure to obtain a parsed target graph structure, and parsing the history graph structure to obtain a parsed history graph structure; Comparing the first operator in the parsed target graph structure with the second operator in the parsed history graph structure, obtaining a target operator in the first operator that is not repeated with the second operator as a fusion operator in the target graph structure; Determine a historical operator in the historical graph structure that corresponds to the fusion operator.

5. The method according to any one of claims 1 to 4, characterized in that After determining the first target test result of the fusion operator according to the acceleration ratio of the fusion operator, the method further includes: When the target test result satisfies the preset test result, constructing the deep learning model to be constructed according to the target graph structure; Inputting a second input parameter associated with the deep learning model to be constructed into the deep learning model to be constructed, obtaining a third output result and a third output time corresponding to the third output result, and inputting the second input parameter into the historical deep learning model, obtaining a fourth output result and a fourth output time corresponding to the fourth output result; A second target test result of the deep learning model to be constructed is determined according to the third output time and the fourth output time.

6. The method according to claim 5, characterized in that The step of constructing the deep learning model to be constructed according to the target graph structure includes: Obtaining weight information of the historical deep learning model; Construct the deep learning model to be constructed according to the target graph structure and the weight information.

7. A fusion operator testing device, characterized in that: The device comprises: A structure acquisition module, used to acquire a target graph structure of a deep learning model to be constructed, and to acquire a historical graph structure of a historical deep learning model corresponding to the deep learning model to be constructed; An operator determination module, used to determine, according to the target graph structure and the historical graph structure, a fusion operator in the target graph structure and a historical operator in the historical graph structure corresponding to the fusion operator; An operator processing module, configured to input a first input parameter associated with the deep learning model to be constructed into the fusion operator to obtain a first output result and a first output time corresponding to the first output result, and to input the first input parameter into the history operator to obtain a second output result and a second output time corresponding to the second output result; a speedup ratio determination module, configured to determine a speedup ratio of the fusion operator according to the first output time and the second output time; the speedup ratio of the fusion operator is used to represent a speedup ratio corresponding to the number of the fusion operators, the first output time, the second output time, and a historical processing time, wherein the historical processing time is obtained by inputting the first input parameter into the historical deep learning model; A result determination module is used to determine a first target test result of the fusion operator according to the acceleration ratio of the fusion operator.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Deep learning model reasoning method, machine translation method and device

    CN116011468A