Hardware operation deviation positioning method and device, equipment and storage medium
By loading the same deep learning model on multiple hardware platforms, automatically collecting and comparing operation data, and generating hardware operation deviation positioning reports, the problem of inefficiency in the existing technology is solved and efficient and accurate hardware operation deviation positioning is achieved.
Patent Information
- Application Number
- CN202510897919.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-08-01
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, the positioning of hardware operation deviation is inefficient by manually inserting print statements or manually comparing the output results of the model layer by layer, especially when a large number of printing points or comparison tasks are required, resulting in low positioning of hardware operation deviation.
By loading the same preset deep learning model on multiple hardware platforms, collecting the operation data of multi-layer operator nodes, and automatically comparing the operation data of the reference hardware platform and the hardware platform to be tested, generating hardware operation deviation positioning reports to improve positioning efficiency.
It realizes efficient positioning of hardware operation deviations, reduces human operation errors, improves positioning accuracy and accuracy, and enhances cross-platform compatibility and deployment compatibility and iterative reliability of deep learning models.
Smart Images

Figure CN120407363A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and particularly to a method, device, equipment, and storage medium for locating hardware operation deviations. Background Art
[0002] Hardware operation deviation location refers to the technical process of identifying and diagnosing numerical result differences generated when a deep learning model runs on different hardware platforms through a systematic method. The core challenge stems from the heterogeneity of hardware. That is, the precision processing of floating-point operations, parallel computing strategies, and the underlying implementation differences of vendor-optimized libraries by computing units with different architectures will cause deviations in model inference that may affect business logic. Therefore, with the increasing requirements of business logic, efficient hardware operation deviation location is particularly important.
[0003] In related technologies, the current hardware operation deviation location method detects and locates the operation deviations of the same deep learning model on different hardware platforms by manually inserting print statements or manually comparing the model output results layer by layer. However, in related technologies, the process of manually inserting print statements or manually comparing the model output results layer by layer is time-consuming. When a large number of print points need to be inserted in the model or there are a large number of comparison tasks, it will cause the problem of low efficiency in hardware operation deviation location. Summary of the Invention
[0004] This application provides a method, device, equipment, and storage medium for locating hardware operation deviations to at least solve the problem in related technologies that the process is time-consuming by manually inserting print statements or manually comparing the model output results layer by layer, and when a large number of print points need to be inserted in the model or there are a large number of comparison tasks, it will cause the problem of low efficiency in hardware operation deviation location.
[0005] This application provides a method for locating hardware operation deviations, including:
[0006] Determine multiple hardware platforms, where the multiple hardware platforms include a reference hardware platform and multiple hardware platforms to be tested;
[0007] Load the same preset deep learning model on each hardware platform, where the preset deep learning model includes multiple layers of operator nodes;
[0008] For each hardware platform, input the preset test sample data into the preset deep learning model loaded on each hardware platform respectively, so that the preset deep learning model in each hardware platform starts to perform model inference processing;
[0009] On each hardware platform, during the process from the preset deep learning model starting to perform model inference processing to completing model inference processing, collect the operation data of the multiple layers of operator nodes of the preset deep learning model;
[0010] For each hardware platform, summarize the operation data of each operator node collected into model operation data;
[0011] Obtain the model operation data corresponding to the reference hardware platform and the model operation data corresponding to each hardware platform to be tested;
[0012] Compare the model operation data corresponding to each hardware platform to be tested with the model operation data corresponding to the reference hardware platform to obtain the comparison error data corresponding to each hardware platform to be tested;
[0013] Generate a hardware operation deviation positioning report for each hardware platform to be tested for a preset deep learning model according to the comparison error data corresponding to each hardware platform to be tested;
[0014] Output the hardware operation deviation positioning report.
[0015] This application also provides a hardware operation deviation positioning device, including:
[0016] A first determination module, configured to determine a plurality of hardware platforms, where the plurality of hardware platforms include a reference hardware platform and a plurality of hardware platforms to be tested;
[0017] A loading module, configured to load the same preset deep learning model on each hardware platform, where the preset deep learning model includes multiple layers of operator nodes;
[0018] An input module, configured to, for each hardware platform, input preset test sample data into the preset deep learning model loaded on each hardware platform respectively, so that the preset deep learning model in each hardware platform starts to perform model inference processing;
[0019] A collection module, configured to, on each hardware platform, collect the operation data of multiple layers of operator nodes of the preset deep learning model during the process from the start of the model inference processing of the preset deep learning model to the completion of the model inference processing;
[0020] A summarization module, configured to, for each hardware platform, summarize the operation data of each operator node collected into model operation data;
[0021] A first acquisition module, configured to obtain the model operation data corresponding to the reference hardware platform and the model operation data corresponding to each hardware platform to be tested;
[0022] A first comparison module, configured to compare the model operation data corresponding to each hardware platform to be tested with the model operation data corresponding to the reference hardware platform to obtain the comparison error data corresponding to each hardware platform to be tested;
[0023] A first generation module, configured to generate a hardware operation deviation localization report for each hardware platform to be tested with respect to a preset deep learning model according to the comparison error data corresponding to each hardware platform to be tested;
[0024] A first output module, configured to output the hardware operation deviation localization report.
[0025] This application also provides an electronic device, including: a memory, configured to store a computer program; a processor, configured to implement the steps of any of the above-mentioned hardware operation deviation localization methods when executing the computer program.
[0026] This application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned hardware operation deviation localization methods are implemented.
[0027] Through the hardware operation deviation localization method, device, equipment and storage medium provided by the embodiments of this application, a same preset deep learning model is loaded on each pre-determined hardware platform; for each hardware platform, preset test sample data is respectively input into the preset deep learning model loaded on each hardware platform, so that the preset deep learning model in each hardware platform performs inference; on each hardware platform, during the process from the start of inference to the completion of inference of the preset deep learning model, operation data of multiple layers of operator nodes of the preset deep learning model is collected; for each hardware platform, the collected operation data of each operator node is aggregated into model operation data; the model operation data corresponding to a reference hardware platform and the model operation data corresponding to each hardware platform to be tested are obtained; the model operation data corresponding to each hardware platform to be tested is compared with the model operation data corresponding to the reference hardware platform to obtain comparison error data corresponding to each hardware platform to be tested; a hardware operation deviation localization report for each hardware platform to be tested with respect to the preset deep learning model is generated according to the comparison error data corresponding to each hardware platform to be tested; the hardware operation deviation localization report is output, and by automatically comparing the operation data of the corresponding operator nodes of the preset deep learning model on the reference hardware platform and each hardware platform to be tested, the efficiency of hardware operation deviation localization is improved. Description of the Drawings
[0028] In order to more clearly illustrate the embodiments of this application, the following will briefly introduce the drawings required for the embodiments. Obviously, the drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to these drawings.
[0029] Figure 1 It is a schematic diagram of the application scenario of the hardware operation deviation localization method provided by the embodiments of this application;
[0030] Figure 2 Flow schematic of the hardware operation deviation positioning method provided by the embodiment of the present application Figure 1 ;
[0031] Figure 3 Flow schematic of the hardware operation deviation positioning method provided by the embodiment of the present application Figure 2 ;
[0032] Figure 4 Structural schematic diagram of the hardware operation deviation positioning device provided by the embodiment of the present application;
[0033] Figure 5 Hardware structural schematic diagram of the electronic device provided by the embodiment of the present application. Specific embodiments
[0034] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.
[0035] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variant thereof are intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects, rather than to describe a specific order or sequence.
[0036] Hardware operation deviation positioning refers to the technical process of identifying and diagnosing the numerical result differences generated when a deep learning model runs on different hardware platforms through a systematic method. Its core challenge stems from the heterogeneity of hardware, that is, the differences in the precision processing of floating-point operations, parallel computing strategies, and the underlying implementation of vendor optimization libraries by computing units with different architectures will cause deviations in model inference that may affect business logic. Therefore, with the increasing requirements of business logic, efficient hardware operation deviation positioning is particularly important. In related technologies, the current hardware operation deviation positioning method detects and locates the operation deviations of the same deep learning model on different hardware platforms by manually inserting print statements or manually comparing the model output results layer by layer. However, in related technologies, the method of manually inserting print statements or manually comparing the model output results layer by layer is time-consuming. When a large number of print points need to be inserted in the model or there are a large number of comparison tasks, it will cause the problem of low efficiency in hardware operation deviation positioning.
[0037] To solve the above technical problems, the embodiments of the present application propose the following technical concepts: The inventors considered a reference hardware platform and multiple hardware platforms to be tested, loaded the same preset deep learning model including multiple layers of operator nodes on each hardware platform, inferred preset test sample data based on each preset deep learning model to obtain the operation data of the multiple layers of operator nodes; for each hardware platform, determined the operation data of each operator node as the model operation data; compared the model operation data corresponding to each hardware platform to be tested with the model operation data corresponding to the reference hardware platform to obtain the comparison error data corresponding to each hardware platform to be tested; determined a hardware operation deviation positioning report according to each comparison error data. By automatically comparing the operation data of each operator node of the preset deep learning model on the reference hardware platform and each hardware platform to be tested, the efficiency of hardware operation deviation positioning is improved.
[0038] To enable those skilled in the art of the present technology to better understand the solution of the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0039] Combined with the specific application environment architecture or specific hardware architecture on which the execution of the hardware operation deviation positioning method depends, the specific application environment architecture or specific hardware architecture will be described herein. Refer to Figure 1 , Figure 1 FIG. 12 is a schematic diagram of an application scenario of the hardware operation deviation positioning method provided by the embodiments of the present application.
[0040] As Figure 1 shown, the application scenario of the hardware operation deviation positioning method includes: an electronic device 101 and multiple hardware platforms 102.
[0041] The electronic device 101 can be a single server or a cluster composed of multiple servers, or other devices.
[0042] The multiple hardware platforms 102 include a reference hardware platform 1021 and multiple hardware platforms 1022 to be tested.
[0043] An electronic device 101 determines multiple hardware platforms 102; loads the same preset deep learning model on each hardware platform 102; for each hardware platform 102, inputs preset test sample data into the preset deep learning model loaded on each hardware platform 102 respectively, so that the preset deep learning model in each hardware platform 102 starts to perform model inference processing; on each hardware platform, during the process from the preset deep learning model starting to perform model inference processing to completing the model inference processing, collects the operation data of multiple operator nodes of the preset deep learning model; for each hardware platform 102, summarizes the collected operation data of each operator node into model operation data; obtains the model operation data corresponding to the reference hardware platform 1021 and the model operation data corresponding to each hardware platform 1022 to be tested; compares the model operation data corresponding to each hardware platform 1022 to be tested with the model operation data corresponding to the reference hardware platform 1021 to obtain the comparison error data corresponding to each hardware platform 1022 to be tested; generates a hardware operation deviation localization report for each hardware platform 1022 to be tested for the preset deep learning model according to the comparison error data corresponding to each hardware platform 1022 to be tested; outputs the hardware operation deviation localization report.
[0044] Figure 2 Schematic flow of the hardware operation deviation localization method provided by the embodiment of the present application Figure 1 , such as Figure 2 shown, an embodiment of the present application provides a hardware operation deviation localization method, and the method is described in detail as follows:
[0045] S201: Determine multiple hardware platforms, where the multiple hardware platforms include a reference hardware platform and multiple hardware platforms to be tested.
[0046] Exemplarily, the multiple hardware platforms may be GPUs, MLUs, image recognizers, and other hardware platforms.
[0047] Among them, GPU is a graphics processor and MLU is a machine learner.
[0048] Exemplarily, since the GPU has relatively complete support for the deep learning model, the GPU is determined as the reference hardware platform; the MLU and MX are determined as the corresponding hardware platforms to be tested.
[0049] S202: Load the same preset deep learning model on each hardware platform, where the preset deep learning model includes multiple operator nodes.
[0050] In this embodiment, the model structures and model parameters of each preset deep learning model are completely the same, and each preset deep learning model is set to the evaluation mode.
[0051] In addition, according to the multi-layer operator nodes in the preset deep learning model, the corresponding input tensors and output tensors are pre-registered for the operator nodes of each layer of forward propagation.
[0052] S203: For each hardware platform, input the preset test sample data into the preset deep learning model loaded in each hardware platform respectively, so that the preset deep learning model in each hardware platform starts to perform model inference processing.
[0053] S204: On each hardware platform, during the process from the preset deep learning model starting to perform model inference processing to completing the model inference processing, collect the running data of the multi-layer operator nodes of the preset deep learning model.
[0054] In this embodiment, the running data of the operator nodes includes one or more of the following: input tensor, output tensor, operator node name, hierarchical path, and corresponding hardware platform information.
[0055] Among them, the operator node name, hierarchical path, and hardware platform information can also be determined as meta-information.
[0056] In addition, before step S204, it further includes: dynamically insert a hook function for the preset deep learning model through the register_forward_hook function or the torch.fx function.
[0057] Among them, the register_forward_hook function is the corresponding function of the hook function for intercepting the forward propagation process of the neural network; the torch.fx function is the corresponding function of the code conversion tool chain.
[0058] Among them, the hook function is a programming mechanism used to collect the running data of the operator nodes of each layer.
[0059] Specifically, step S204 is specifically: on each hardware platform, during the process from the preset deep learning model starting to perform model inference processing to completing the model inference processing, collect the running data of the multi-layer operator nodes of the preset deep learning model through the hook function.
[0060] In addition, after step S204, it further includes: store the running data of the multi-layer operator nodes of the preset deep learning model of each hardware platform collected to the local cache for subsequent invocation.
[0061] S205: For each hardware platform, summarize the running data of each operator node collected into model running data.
[0062] S206: Obtain the model running data corresponding to the reference hardware platform and the model running data corresponding to each hardware platform to be tested.
[0063] S207: Compare the model operation data corresponding to each hardware platform to be tested with the model operation data corresponding to the reference hardware platform to obtain the comparison error data corresponding to each hardware platform to be tested.
[0064] Specifically, step S207 specifically includes:
[0065] S2071: Obtain the operation data of each operator node corresponding to the model operation data of any hardware platform to be tested.
[0066] S2072: Obtain the reference operation data of each operator node corresponding to the model operation data of the reference hardware platform.
[0067] S2073: Compare the operation data of each operator node with the corresponding reference operation data layer by layer to obtain the comparison error value corresponding to each operator node.
[0068] Specifically, compare the output tensors in the operation data of each operator node with the output tensors in the corresponding reference operation data layer by layer, and use the calculation methods of absolute error, relative error, maximum difference, and / or numerical distribution statistics to obtain the comparison error value corresponding to each operator node.
[0069] [[ID=|18]]S2074: Integrate each comparison error value into the comparison error data of the hardware platform to be tested.
[0070] S2075: Determine the comparison error data of the remaining hardware platforms to be tested.
[0071] Specifically, according to the processing procedures of steps S2071 to S2074, determine the comparison error data of the remaining hardware platforms to be tested.
[0072] S208: Generate a hardware operation deviation localization report for each hardware platform to be tested for the preset deep learning model according to the comparison error data corresponding to each hardware platform to be tested.
[0073] In this embodiment, the comparison error data includes the comparison error values of each operator node; correspondingly, step S208 specifically includes:
[0074] S2081: Determine whether the comparison error value of each operator node of any hardware platform to be tested is greater than the preset error threshold.
[0075] In this embodiment, the preset error threshold is a configurable error tolerance threshold, and the corresponding value can be 1e-4 or other values.
[0076] S208: If it is determined that the comparison error value of any operator node is greater than the preset error threshold, generate a serious operation deviation localization report for the operator node.
[0077] S2083: If it is determined that the comparison error value of any operator node is less than or equal to the preset error threshold, a minor operation deviation location report of the operator node is generated.
[0078] S2084: According to the serious operation deviation location reports or minor operation deviation location reports corresponding to each operator node, an operation deviation location report of the to-be-tested hardware platform for the preset deep learning model is generated.
[0079] S2085: According to the operation deviation location reports corresponding to each to-be-tested hardware platform, a hardware operation deviation location report of each to-be-tested hardware platform for the preset deep learning model is generated.
[0080] Specifically, the operation deviation location reports corresponding to each to-be-tested hardware platform are integrated to generate a hardware operation deviation location report of each to-be-tested hardware platform for the preset deep learning model.
[0081] Exemplarily, the hardware operation deviation location report includes: information of each to-be-tested hardware platform, the names of the operator nodes with serious operation deviations and the names of the operator nodes with minor operation deviations corresponding to each to-be-tested hardware platform, the shapes of the input tensors and output tensors of each operator node, the maximum error value of the output tensors of each operator node, and the complete hierarchical path.
[0082] S209: Output the hardware operation deviation location report.
[0083] In addition, after step S209, steps a to c are further included:
[0084] Step a: In response to the location operation of the preset deep learning model in the corresponding to-be-tested hardware platform according to the hardware operation deviation location report, one or more operator nodes are determined.
[0085] In this embodiment, the one or more operator nodes are the operator nodes with serious operation deviations.
[0086] Step b: Extract the operation data and pre-stored log information of each operator node.
[0087] Step c: In response to the repair operation of the preset deep learning model according to the operation data and pre-stored log information of each operator node, a repaired deep learning model is generated.
[0088] Specifically, in response to the operation data and pre-stored log information of each operator node, the problem is reproduced in the local environment, and the preset deep learning model is repaired according to the reproduced problem, and a repaired deep learning model is generated.
[0089] In summary, the hardware operation deviation positioning method provided in this embodiment loads the same preset deep learning model on each predetermined hardware platform; for each hardware platform, the preset test sample data is respectively input into the preset deep learning model loaded on each hardware platform, so that the preset deep learning model in each hardware platform performs inference; on each hardware platform, during the process from the start of inference to the completion of inference of the preset deep learning model, the operation data of multiple operator nodes of the preset deep learning model is collected; for each hardware platform, the operation data of each collected operator node is summarized into model operation data; the model operation data corresponding to the reference hardware platform and the model operation data corresponding to each hardware platform to be tested are obtained; the model operation data corresponding to each hardware platform to be tested is compared with the model operation data corresponding to the reference hardware platform to obtain the comparison error data corresponding to each hardware platform to be tested; according to the comparison error data corresponding to each hardware platform to be tested, a hardware operation deviation positioning report for each hardware platform to be tested for the preset deep learning model is generated; the hardware operation deviation positioning report is output, and by automatically comparing the operation data of the corresponding operator nodes of the preset deep learning model on the reference hardware platform and each hardware platform to be tested, the efficiency of hardware operation deviation positioning is improved.
[0090] In addition, the hardware operation deviation positioning method provided in this embodiment, by automatically comparing the operation data of the corresponding operator nodes of the preset deep learning model on the reference hardware platform and each hardware platform to be tested, does not require manual participation, avoids human operation errors, and improves the accuracy of hardware operation deviation positioning.
[0091] In addition, the hardware operation deviation positioning method provided in this embodiment can perform deviation detection at the operator node level, thus realizing the accuracy of hardware operation deviation positioning.
[0092] In addition, the hardware operation deviation positioning method provided in this embodiment enhances cross-platform compatibility and adaptability by performing hardware operation deviation positioning processing on multiple different hardware platforms.
[0093] In addition, the hardware operation deviation positioning method provided in this embodiment repairs the corresponding preset deep learning model through each operator node with operation deviation to obtain a repaired deep learning model, which improves the compatibility of the deployment of the deep learning model and the reliability of the iteration of the deep learning model.
[0094] Figure 3 It is a schematic flow of the hardware operation deviation positioning method provided in the embodiment of the present application Figure 2 。In the embodiment of the present application, based on the Figure 2 embodiment provided, a detailed description is given of a specific implementation method for another form of hardware operation deviation positioning after step S203. As Figure 3As shown in the figure, the hardware operation deviation positioning method includes:
[0095] S301: Determine multiple hardware platforms, where the multiple hardware platforms include a reference hardware platform and multiple hardware platforms to be tested.
[0096] In this embodiment, the descriptions of the multiple hardware platforms, the reference hardware platform, and the multiple hardware platforms to be tested have been elaborated in step S201, and will not be repeated here.
[0097] S302: Load the same preset deep learning model on each hardware platform, where the preset deep learning model includes multiple operator nodes.
[0098] In this embodiment, the descriptions of the preset deep learning model and the multiple operator nodes have been elaborated in step S202, and will not be repeated here.
[0099] S303: For each hardware platform, input the preset test sample data into the preset deep learning model loaded on each hardware platform respectively, so that the preset deep learning models in each hardware platform start to perform model inference processing.
[0100] S304: Obtain the reference inference result of the preset deep learning model in the reference hardware platform, and the inference results of the preset deep learning models in each hardware platform to be tested.
[0101] In addition, the user can specify the module name or level in the preset deep learning model to be monitored through a configuration file or command-line parameters to obtain the corresponding intermediate inference results.
[0102] S305: Compare each inference result with the reference inference result respectively to obtain the comparison error data corresponding to each hardware platform to be tested.
[0103] Specifically, compare each inference result with the reference inference result respectively, and obtain the comparison error data corresponding to each hardware platform to be tested by calculating the absolute error, relative error, maximum difference, and / or numerical distribution statistics.
[0104] S306: Generate a hardware operation deviation positioning report for each hardware platform to be tested for the preset deep learning model according to the comparison error data corresponding to each hardware platform to be tested.
[0105] Specifically, step S306 specifically includes:
[0106] S3061: Determine whether the comparison error data corresponding to each hardware platform to be tested is greater than a preset error threshold.
[0107] In this embodiment, the discussion on the preset error threshold has been described in detail in step S2081, and will not be elaborated here.
[0108] S3062: If it is determined that the comparison error data corresponding to any hardware platform to be tested is greater than the preset error threshold, a serious operation deviation positioning report for the hardware platform to be tested is generated.
[0109] S3063: If it is determined that the comparison error data corresponding to any hardware platform to be tested is less than or equal to the preset error threshold, a minor operation deviation positioning report for the hardware platform to be tested is generated.
[0110] S3064: According to the serious operation deviation positioning report or minor operation deviation positioning report corresponding to each hardware platform to be tested, a hardware operation deviation positioning report for each hardware platform to be tested for the preset deep learning model is generated.
[0111] Exemplarily, the hardware operation deviation positioning report includes: the hardware platforms to be tested with serious operation deviations, the hardware platforms to be tested with minor operation deviations, the maximum error values corresponding to each hardware platform to be tested, the information of each hardware platform to be tested, and the complete module call path.
[0112] S307: Output the hardware operation deviation positioning report.
[0113] In summary, the hardware operation deviation positioning method provided in this embodiment obtains the reference inference results of the preset deep learning model in the reference hardware platform and the inference results of the preset deep learning model in each hardware platform to be tested; compares each inference result with the reference inference result to obtain the comparison error data corresponding to each hardware platform to be tested; generates a hardware operation deviation positioning report for each hardware platform to be tested for the preset deep learning model according to the comparison error data corresponding to each hardware platform to be tested; outputs the hardware operation deviation positioning report. By only comparing each inference result with the reference inference result, unnecessary intermediate tensor records are reduced, thus saving computing resources.
[0114] In addition, the hardware operation deviation positioning method provided in this embodiment specifies the module name or level in the preset deep learning model that needs to be monitored through the user configuration file or command line parameters to obtain the corresponding intermediate inference results, so that the monitoring range of the preset deep learning model can be flexibly controlled.
[0115] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus the necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method.
[0116] Figure 4The structural schematic diagram of the hardware operation deviation positioning device provided by the embodiment of the present application. As Figure 4 shown, the embodiment of the present application also provides a hardware operation deviation positioning device, including: a first determination module 401, a loading module 402, an input module 403, a collection module 404, a summarization module 405, a first acquisition module 406, a first comparison module 407, a first generation module 408, and a first output module 409.
[0117] The first determination module 401 is used to determine a plurality of hardware platforms, where the plurality of hardware platforms include a reference hardware platform and a plurality of hardware platforms to be tested;
[0118] The loading module 402 is used to load the same preset deep learning model on each hardware platform, where the preset deep learning model includes multiple operator nodes;
[0119] The input module 403 is used to input the preset test sample data into the preset deep learning model loaded on each hardware platform for each hardware platform, so that the preset deep learning model in each hardware platform starts to perform model inference processing;
[0120] The collection module 404 is used to collect the operation data of the multiple operator nodes of the preset deep learning model on each hardware platform during the process from the start of the model inference processing of the preset deep learning model to the completion of the model inference processing;
[0121] The summarization module 405 is used to summarize the operation data of each operator node collected for each hardware platform into model operation data;
[0122] The first acquisition module 406 is used to acquire the model operation data corresponding to the reference hardware platform and the model operation data corresponding to each hardware platform to be tested;
[0123] The first comparison module 407 is used to compare the model operation data corresponding to each hardware platform to be tested with the model operation data corresponding to the reference hardware platform to obtain the comparison error data corresponding to each hardware platform to be tested;
[0124] The first generation module 408 is used to generate a hardware operation deviation positioning report for each hardware platform to be tested for the preset deep learning model according to the comparison error data corresponding to each hardware platform to be tested;
[0125] The first output module 409 is used to output the hardware operation deviation positioning report.
[0126] In a possible implementation manner, the first comparison module 407 specifically includes:
[0127] A first acquisition unit, configured to acquire the operation data of each operator node corresponding to the model operation data of any hardware platform to be tested;
[0128] A second acquisition unit, configured to acquire the reference operation data of each operator node corresponding to the model operation data of a reference hardware platform;
[0129] A comparison unit, configured to compare the operation data of each operator node with the corresponding reference operation data layer by layer to obtain the comparison error value corresponding to each operator node;
[0130] An integration unit, configured to integrate each comparison error value into the comparison error data of the hardware platform to be tested;
[0131] A determination unit, configured to determine the comparison error data of the remaining hardware platforms to be tested.
[0132] In a possible implementation manner, the comparison error data includes the comparison error values of each operator node; correspondingly, the first generation module 408 specifically includes:
[0133] A judgment unit, configured to judge whether the comparison error value of each operator node of any hardware platform to be tested is greater than a preset error threshold;
[0134] A first generation unit, configured to generate a serious operation deviation positioning report of the operator node if it is determined that the comparison error value of any operator node is greater than the preset error threshold;
[0135] A second generation unit, configured to generate a minor operation deviation positioning report of the operator node if it is determined that the comparison error value of any operator node is less than or equal to the preset error threshold;
[0136] A first generation unit, configured to generate an operation deviation positioning report of the hardware platform to be tested for the preset deep learning model according to the serious operation deviation positioning report or the minor operation deviation positioning report corresponding to each operator node;
[0137] A second generation unit, configured to generate a hardware operation deviation positioning report of each hardware platform to be tested for the preset deep learning model according to the operation deviation positioning report corresponding to each hardware platform to be tested. [[ID=3B]]
[0138] In a possible implementation manner, the device further includes:
[0139] A second acquisition module, configured to acquire the reference inference result of the preset deep learning model in the reference hardware platform and the inference results of the preset deep learning model in each hardware platform to be tested;
[0140] A second comparison module, configured to compare each inference result with the reference inference result respectively to obtain the comparison error data corresponding to each hardware platform to be tested;
[0141] A second generation module, configured to generate a hardware operation deviation localization report for each hardware platform to be tested with respect to a preset deep learning model according to the comparison error data corresponding to each hardware platform to be tested;
[0142] A second output module, configured to output the hardware operation deviation localization report.
[0143] In a possible implementation manner, the second generation module specifically includes:
[0144] A determination unit, configured to determine whether the comparison error data corresponding to each hardware platform to be tested is greater than a preset error threshold;
[0145] A first generation unit, configured to generate a serious operation deviation localization report for the hardware platform to be tested if it is determined that the comparison error data corresponding to any hardware platform to be tested is greater than the preset error threshold;
[0146] A second generation unit, configured to generate a minor operation deviation localization report for the hardware platform to be tested if it is determined that the comparison error data corresponding to any hardware platform to be tested is less than or equal to the preset error threshold;
[0147] A third generation unit, configured to generate a hardware operation deviation localization report for each hardware platform to be tested with respect to the preset deep learning model according to the serious operation deviation localization report or the minor operation deviation localization report corresponding to each hardware platform to be tested.
[0148] In a possible implementation manner, the apparatus further includes:
[0149] A second determination module, configured to determine one or more operator nodes corresponding thereto in response to a localization operation on the preset deep learning model in the corresponding hardware platform to be tested according to the hardware operation deviation localization report;
[0150] An extraction module, configured to extract the operation data of each operator node and the pre-stored log information;
[0151] A third generation module, configured to generate a repaired deep learning model in response to a repair operation on the preset deep learning model according to the operation data of each operator node and the pre-stored log information.
[0152] In a possible implementation manner, the operation data of the operator node includes one or more of the following: input tensor, output tensor, operator node name, hierarchical path, and corresponding hardware platform information.
[0153] For the description of the features in the embodiments corresponding to the hardware operation deviation localization device, reference may be made to the relevant description of the embodiments corresponding to the hardware operation deviation localization method, which will not be elaborated herein one by one.
[0154] Figure 5 A schematic structural diagram of the electronic device provided by this application. As Figure 5 shown, the electronic device provided in this embodiment includes: at least one processor 501 and a memory 502. Optionally, the electronic device further includes a communication component 503. Among them, the processor 501, the memory 502, and the communication component 503 are connected through a bus.
[0155] In the specific implementation process, at least one processor 501 executes the computer-executable instructions stored in the memory 502, so that at least one processor 501 executes the above-mentioned embodiment of the hardware operation deviation positioning method.
[0156] For the specific implementation process of the processor 501, reference can be made to the above method embodiment, and its implementation principle and technical effect are similar, which will not be elaborated here in this embodiment.
[0157] In the above embodiment, it should be understood that the processor may be a central processing unit (Central Processing Unit, abbreviated as: CPU), or other general-purpose processors, digital signal processors (Digital Signal Processor, abbreviated as: DSP), application specific integrated circuits (Application Specific Integrated Circuit, abbreviated as: ASIC), etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the application can be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor.
[0158] The memory may include a high-speed memory (Random Access Memory, RAM), and may also include a non-volatile memory (Non-volatile Memory, NVM), such as at least one disk memory.
[0159] The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, the bus in the drawings of this application is not limited to only one bus or one type of bus.
[0160] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. Wherein, the computer program is configured to execute the steps in any of the above embodiments of the hardware operation deviation positioning method when running.
[0161] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media that can store computer programs such as USB flash drives, read-only memories (ROM for short), random access memories (RAM for short), mobile hard disks, magnetic disks or optical discs.
[0162] An embodiment of the present application further provides a computer program product. The above computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above embodiments of the hardware operation deviation positioning method.
[0163] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above embodiments of the hardware operation deviation positioning method.
[0164] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0165] The above has introduced in detail a hardware operation deviation positioning method, device, equipment and storage medium provided by this application. Specific examples are used in this article to elaborate on the principle and implementation manner of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the protection scope of the claims of this application.
Claims
1. A method for locating hardware operation deviation, characterized in that including: determine a plurality of hardware platforms, where the plurality of hardware platforms includes a reference hardware platform and a plurality of hardware platforms to be tested; load the same preset deep learning model on each hardware platform, where the preset deep learning model includes multiple operator nodes; for each hardware platform, input the preset test sample data into the preset deep learning model loaded on each hardware platform respectively, so that the preset deep learning model in each hardware platform starts to perform model inference processing; on each hardware platform, during the process from the preset deep learning model starting to perform model inference processing to completing the model inference processing, collect the operation data of the multiple operator nodes of the preset deep learning model; for each hardware platform, summarize the operation data of each operator node collected into model operation data; obtain the model operation data corresponding to the reference hardware platform and the model operation data corresponding to each hardware platform to be tested; compare the model operation data corresponding to each hardware platform to be tested with the model operation data corresponding to the reference hardware platform to obtain the comparison error data corresponding to each hardware platform to be tested; generate a hardware operation deviation localization report for each hardware platform to be tested for the preset deep learning model according to the comparison error data corresponding to each hardware platform to be tested; output the hardware operation deviation localization report.
2. The hardware operation deviation positioning method according to claim 1, wherein The comparing the model operation data corresponding to each hardware platform to be tested with the model operation data corresponding to the reference hardware platform to obtain the comparison error data corresponding to each hardware platform to be tested includes: obtain the operation data of each operator node corresponding to the model operation data of any hardware platform to be tested; obtain the reference operation data of each operator node corresponding to the model operation data of the reference hardware platform; compare the operation data of each operator node with the corresponding reference operation data layer by layer to obtain the comparison error value corresponding to each operator node; integrate each comparison error value into the comparison error data of the hardware platform to be tested; determine the comparison error data of the remaining hardware platforms to be tested.
3. The hardware operation deviation positioning method according to claim 1, characterized in that where the comparison error data includes the comparison error values of each operator node; Accordingly, the generating a hardware operation deviation localization report for each hardware platform to be tested for the preset deep learning model according to the comparison error data corresponding to each hardware platform to be tested includes: judge whether the comparison error value of each operator node of any hardware platform to be tested is greater than a preset error threshold; if it is determined that the comparison error value of any operator node is greater than the preset error threshold, generate a serious operation deviation localization report for the operator node; if it is determined that the comparison error value of any operator node is less than or equal to the preset error threshold, generate a minor operation deviation localization report for the operator node; generate an operation deviation localization report for the hardware platform to be tested for the preset deep learning model according to the serious operation deviation localization report or minor operation deviation localization report corresponding to each operator node; generate the hardware operation deviation localization report for each hardware platform to be tested for the preset deep learning model according to the operation deviation localization reports corresponding to each hardware platform to be tested.
4. The hardware operation deviation positioning method according to claim 1, characterized in that After inputting the preset test sample data into the preset deep learning models loaded on each hardware platform respectively for each hardware platform, such that the preset deep learning models in each hardware platform start to perform model inference processing, it further includes: Obtaining the reference inference results of the preset deep learning model in the reference hardware platform, and the inference results of the preset deep learning models in each hardware platform to be tested; Comparing each inference result with the reference inference result respectively to obtain the comparison error data corresponding to each hardware platform to be tested; Generating a hardware operation deviation localization report for each hardware platform to be tested with respect to the preset deep learning model according to the comparison error data corresponding to each hardware platform to be tested; Outputting the hardware operation deviation localization report.
5. The hardware operation deviation positioning method according to claim 4, characterized in that The generating a hardware operation deviation localization report for each hardware platform to be tested with respect to the preset deep learning model according to the comparison error data corresponding to each hardware platform to be tested includes: Judging whether the comparison error data corresponding to each hardware platform to be tested is greater than a preset error threshold; If it is determined that the comparison error data corresponding to any hardware platform to be tested is greater than the preset error threshold, generating a serious operation deviation localization report for the hardware platform to be tested; If it is determined that the comparison error data corresponding to any hardware platform to be tested is less than or equal to the preset error threshold, generating a minor operation deviation localization report for the hardware platform to be tested; Generating a hardware operation deviation localization report for each hardware platform to be tested with respect to the preset deep learning model according to the serious operation deviation localization reports or minor operation deviation localization reports corresponding to each hardware platform to be tested.
6. The hardware operation deviation positioning method according to claim 1, wherein After the outputting the hardware operation deviation localization report, it further includes: Responding to the localization operation on the preset deep learning model in the corresponding hardware platform to be tested according to the hardware operation deviation localization report, and determining one or more corresponding operator nodes; Extracting the operation data and pre-stored log information of each operator node; Responding to the repair operation on the preset deep learning model according to the operation data of each operator node and the pre-stored log information, and generating a repaired deep learning model.
7. The hardware operation deviation positioning method according to any one of claims 1 to 6, characterized in that The operation data of the operator node includes one or more of the following: input tensor, output tensor, operator node name, hierarchical path, and corresponding hardware platform information.
8. A hardware operation deviation positioning device, characterized in that It includes: A first determination module, configured to determine a plurality of hardware platforms, where the plurality of hardware platforms includes a reference hardware platform and a plurality of hardware platforms to be tested; A loading module, configured to load the same preset deep learning model on each hardware platform, where the preset deep learning model includes multiple layers of operator nodes; An input module, configured to, for each hardware platform, input the preset test sample data into the preset deep learning model loaded on each hardware platform respectively, such that the preset deep learning models in each hardware platform start to perform model inference processing; A collection module, configured to, on each hardware platform, collect the operation data of the multiple layers of operator nodes of the preset deep learning model during the process from the start of the preset deep learning model performing model inference processing to the completion of model inference processing; A summarization module, configured to summarize the operation data of each operator node collected for each hardware platform into model operation data; A first acquisition module, configured to acquire the model operation data corresponding to the reference hardware platform and the model operation data corresponding to each hardware platform to be tested; A first comparison module, configured to compare the model operation data corresponding to each hardware platform to be tested with the model operation data corresponding to the reference hardware platform to obtain comparison error data corresponding to each hardware platform to be tested; A first generation module, configured to generate a hardware operation deviation positioning report for each hardware platform to be tested for the preset deep learning model according to the comparison error data corresponding to each hardware platform to be tested; A first output module, configured to output the hardware operation deviation positioning report.
9. An electronic device, characterized in that, Comprising: A memory, configured to store a computer program; A processor, configured to implement the steps of the hardware operation deviation positioning method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein the computer program implements the steps of the hardware operation deviation positioning method according to any one of claims 1 to 7 when executed by a processor.
Citation Information
Patent Citations
Model parameter determination method and device and storage medium
CN111143148A
Method and device for determining hardware computing platform distribution mode
CN112988372A
Model processing method and device, electronic equipment and computer storage medium
CN114330668A
Cited By
Cross-frame model reasoning precision anomaly positioning and repairing method and system based on intelligent agent
CN120872784A