A method, device, storage medium and program product for locating an abnormal operator
By comparing the operation results of the target model on different computing devices, dynamically generate a list of operators to be verified and fall back to the verification computing device one by one, the problem of low manual layer-by-layer verification efficiency in the prior art is solved, and efficient and accurate abnormal operator positioning is achieved.
Patent Information
- Application Number
- CN202510344835.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-03-24
AI Technical Summary
In the prior art, the manual layer-by-layer verification method is inefficient and has the risk of manual operation errors, making it difficult to efficiently locate and solve the problem of operator calculation abnormalities in large models.
By obtaining the operation results of the target model on at least two different computing devices, the abnormal operator is determined based on the degree of difference in the operation results, a list of operators to be verified is generated, and one by one falls back to the verification computing device to verify the results, and positioning the abnormal operator.
It improves the efficiency of abnormal detection, avoids the risk of manual operation errors, can quickly lock the candidate range of abnormal operators, and accurately locate abnormal operators.
Smart Images

Figure CN119862064B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technologies, and in particular, to a method, device, storage medium, and program product for locating abnormal operators. Background Art
[0002] With the rapid development of large models in the field of artificial intelligence, the number of large models is increasing continuously, and the types of hardware devices running these models are also becoming increasingly rich. Different hardware devices need to be adapted to different kernels (parallel computing functions), and different large models correspond to different operators. Even for the same operator, its inputs may vary. During the training or inference process of large models, model abnormal problems such as abnormal training loss or incorrect inference results may occur, and these abnormalities are usually caused by abnormal operator calculations. Therefore, screening out the wrong operators and their corresponding input and output is crucial for locating and solving model abnormal problems. Currently, the more common method in related technologies is to perform layer-by-layer verification on large models from the front end of PyTorch (an open-source deep learning framework), and find the problem by comparing the results of each operator. This method has the problem of low efficiency. Summary of the Invention
[0003] This application provides a method, device, storage medium, and program product for locating abnormal operators, which at least solves the technical problem of low efficiency in the manual layer-by-layer verification method in related technologies, and achieves the technical effects of improving the troubleshooting efficiency and avoiding the risk of manual operation errors.
[0004] This application provides a method for locating abnormal operators, including: obtaining the running results of a target model on at least two different computing devices; when determining that there are abnormal operators according to the difference degree of the running results, generating a list of operators to be verified according to the operators called during the running process of the target model; the abnormal operators are at least one operator called during the running process of the target model; rolling back the operators in the list of operators to be verified one by one to a verification computing device, so that the verification computing device runs the target model according to the rolled-back operators to obtain verification results; and locating the abnormal operators according to the verification results and the running results.
[0005] This application also provides an electronic device, including: a memory for storing a computer program; a processor for implementing the steps of any of the above methods for locating abnormal operators when executing the computer program.
[0006] This application also provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the steps of any of the above methods for locating abnormal operators are implemented.
[0007] The present application also provides a computer program product, including a computer program which, when executed by a processor, implements the steps of any of the above-mentioned abnormal operator location methods.
[0008] Through the present application, by comparing the operation result differences of the target model on different computing devices, the candidate range of the abnormal operator is quickly locked, avoiding the waste of resources in the full-scale operator investigation; according to the operators called by the target model, a set of operators to be verified is dynamically generated, and the investigation range is compressed from the entire operator set to the set corresponding to the operators called by the target model; by successively rolling back to the verification computing device and combining with the automated result comparison, the abnormal operator is accurately located. Therefore, the technical problem of low efficiency existing in the manual layer-by-layer verification method can be solved, and the technical effects of improving the investigation efficiency and avoiding the risk of manual operation errors are achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] To more clearly illustrate the embodiments of the present application, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0010] Figure 1 It is a flowchart of a method for locating an abnormal operator provided by an embodiment of the present application;
[0011] Figure 2 It is a schematic diagram of modules of a method for locating an abnormal operator provided by an embodiment of the present application;
[0012] Figure 3 It is a scheduling schematic diagram of a method for locating an abnormal operator provided by an embodiment of the present application;
[0013] Figure 4 It is a specific flowchart of a method for locating an abnormal operator provided by an embodiment of the present application;
[0014] Figure 5 It is a schematic diagram of an electronic device provided by an embodiment of the present application;
[0015] Figure 6 It is a schematic diagram of a computer-readable storage medium provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0016] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.
[0017] It should be noted that in the description of this application, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. The terms "first", "second", etc. in this application are used to distinguish similar objects and not to describe a particular order or sequence.
[0018] To enable those skilled in the art of this technology to better understand the solution of this application, the following further detailed description of this application will be given in conjunction with the accompanying drawings and specific embodiments.
[0019] The embodiments of this application provide a method for locating abnormal operators. In combination with the execution process of the method for locating abnormal operators, the method will be described in detail. Among them, in computer science, a kernel refers to a parallel computing function running on a GPU (Graphics Processing Unit, graphics processing unit). PyTorch is an open-source deep learning framework designed specifically for deep learning and artificial intelligence research. In deep learning, an operator is like an operator in mathematics and is a basic building block for performing specific operations. An operator can be imagined as a "black box" that receives some input data, then processes these data according to predefined rules, and finally outputs the processed result; an operator is one of the core components for constructing a neural network. Different operators can be combined together to form a complex computational graph, thereby realizing the functions of various deep learning models. For example, a simple neural network layer may be composed of a matrix multiplication operator and an activation function operator. Common types of operators include: matrix multiplication, convolution, addition, subtraction, multiplication, division, activation function, pooling, normalization, etc. By combining these operators, deep learning models can perform complex tasks, such as image recognition, natural language processing, and speech recognition, etc.
[0020] As Figure 1 shown, the method for locating abnormal operators includes:
[0021] S11: Obtain the running results of the target model on at least two different computing devices.
[0022] In this step, potential computational anomalies are detected by comparing the running results of the target model on different computing devices. In practical applications, different hardware platforms (such as GPUs, CPUs (Central Processing Units), or other computing devices) may produce different computational results due to differences in computational architectures, precision, optimization, etc. By running the target model on at least two different computing devices (for example, a GPU and a TPU (Tensor Processing Unit), where the TPU is a computing device other than GPUs and CPUs)) and obtaining the computational results under the same input, it can provide strong evidence for subsequent operator anomaly localization. If there are significant differences in the results of different computing devices, then this difference may indicate that there are computational anomalies in some operators.
[0023] S12: When it is determined that there are abnormal operators based on the difference degree of the running results, generate a list of operators to be verified according to the operators called during the running of the target model; the abnormal operators are at least one of the operators called during the running of the target model.
[0024] When the target model is running on different computing devices, if it is detected that the difference in the output results exceeds a predetermined threshold (i.e., the difference degree is large), it indicates that there may be an anomaly during the running of the target model. At this time, it is necessary to locate the operators that may cause this anomaly.
[0025] To achieve this, it is necessary to list all the operators called during the running of the target model. These operators are the basic components of the target model, and each operator has a direct or indirect impact on the output of the target model. Generate a list of operators to be verified based on all the called operators, and these operators will be further verified and investigated to determine which operator causes the model anomaly.
[0026] S13: Step by step, roll back the operators in the list of operators to be verified to the verification computing device, so that the verification computing device runs the target model according to the rolled-back operators and obtains the verification results.
[0027] After determining the list of operators to be verified, step by step roll back these operators in the list of operators to be verified to the verification computing device. Rolling back to the verification computing device means switching the operator to be verified from the current computing device to another device and running the target model on this verification computing device. The verification computing device will execute the calculation of the target model based on the rolled-back operator and generate corresponding verification results. The purpose of this step is to confirm whether there are problems with the calculation of the operator by having the verification computing device execute the same operator under different hardware conditions.
[0028] S14: Locate the abnormal operator according to the verification results and the running results.
[0029] Based on the comparison between the verification result and the original running result, determine which operators are abnormal. During the verification phase, the verification result obtained by the verification computing device will be compared with the result when the target model runs for the first time. In this way, the operator that causes the abnormal operation of the target model can be located. This automatic location process can quickly and efficiently identify the abnormal operator, thereby reducing the time and cost of manual debugging and improving the efficiency of anomaly detection.
[0030] In an exemplary embodiment, obtaining the running results of the target model on at least two different computing devices includes: running the target model on the first computing device to generate a reference result; running the target model on the second computing device to generate a candidate result; the running results include the reference result and the candidate result; determining the existence of an abnormal operator according to the difference degree of the running results includes: determining whether the difference degree between the candidate result and the reference result meets a preset difference condition; if it meets, it is determined that there is an abnormal operator when the second computing device runs the target model.
[0031] In this embodiment, the location of the abnormal operator is achieved by comparing the differences in the running results of the target model on heterogeneous computing devices. Specifically, run the target model on the first computing device (such as a GPU with a verified correct result) to generate a reference result as a reliable reference for subsequent comparison; run the same target model on the second computing device (such as a TPU to be verified) to generate a candidate result, and quantify the difference degree between the two based on metrics such as element-wise error of tensors, statistical distribution differences, or consistency of model functional outputs; the differences can be errors in output values, differences in calculation accuracy, or more complex structural differences. If the difference meets the preset difference condition, it is determined that there is an abnormality in the second computing device due to an error in the implementation of a specific operator.
[0032] It can be seen that by detecting and quantifying this difference, this embodiment can effectively locate the operator that is abnormal when executed on the second computing device, providing a basis for further anomaly analysis and troubleshooting.
[0033] In an exemplary embodiment, the first computing device is a reference computing device with a verified running result, and the second computing device is a computing device to be verified.
[0034] Specifically, the first computing device (such as a verified GPU) serves as a reference device. Its operator implementation has been fully verified and the results are trustworthy, and it is used to generate the benchmark results of the target model. The second computing device (such as a TPU to be verified) serves as the target device. There may be potential defects in its operator implementation, and it runs the same target model to generate candidate results. The difference between the two stems from the specific implementation of the operator by the hardware (such as numerical precision, parallel optimization, or memory management). By comparing the differences between the benchmark results and the candidate results, it can be determined whether there are abnormal operators in the second computing device. Specifically, if the difference degree between the candidate results and the benchmark results exceeds the preset threshold (such as the per-element error exceeds the preset threshold or the functional outputs are inconsistent), it indicates that there are computational anomalies in the operator implementation of the second computing device, and it is necessary to further isolate the specific problematic operator through a dynamic fallback mechanism, thereby converging the complex model anomalies to local defects adaptable to the hardware.
[0035] In an exemplary embodiment, determining whether the difference degree between the candidate results and the benchmark results meets the preset difference condition includes: determining whether the difference degree between the candidate results and the benchmark results is greater than the preset threshold; if not, it is determined that the preset difference condition is met.
[0036] In this embodiment, the process of determining whether the difference degree between the candidate results and the benchmark results meets the preset difference condition is mainly to judge whether there are anomalies by comparing the differences in the running results on the two computing devices. Specifically, calculate the difference between the candidate results and the benchmark results. This difference can be manifested as numerical errors, result deviations, or other forms of computational differences, and then compare this difference with the preset threshold.
[0037] If the difference between the candidate results and the benchmark results is less than the preset threshold, then it can be determined that the computational results of the two are very close, that is, the difference on the computing device is acceptable and within the expected computational accuracy range. Therefore, there is no need to further locate the anomalies of the operator. At this time, the judgment result is that the preset difference condition is not met, that is, the difference degree is small and no anomaly is considered.
[0038] Conversely, if the difference between the candidate results and the benchmark results is greater than the preset threshold, it indicates that there is an obvious deviation between the two, indicating that at least one operator on the second computing device has an anomaly, resulting in incorrect computational results. Therefore, it is determined that the preset difference condition is met, and it is prompted that further investigation of the abnormal operator is required.
[0039] In this way, it is possible to effectively screen out which operators may cause computational result anomalies and conduct targeted further verification and repair.
[0040] In an exemplary embodiment, based on the verification result and the running result, locating abnormal operators includes: locating abnormal operators according to the difference between the verification result and the benchmark result. For example, in a specific embodiment, locating abnormal operators according to the difference between the verification result and the benchmark result includes: determining whether the verification result is the same as the benchmark result; if the verification result is the same as the benchmark result, determining the currently rolled-back operator as an abnormal operator; if the verification result is different from the benchmark result, determining the currently rolled-back operator as a normal operator.
[0041] In an exemplary embodiment, the specific implementation of locating abnormal operators is to roll back each operator in the list to be verified to the verified third computing device (such as a CPU) for execution one by one (other operators remain on the TPU); if the verification result is consistent with the benchmark result after rolling back an operator, it indicates that the implementation of this operator on the original device (TPU) is abnormal (because the result returns to normal after rolling back to the CPU), so this operator is directly locked as an abnormal operator; conversely, if the verification result is still abnormal, it is excluded that the currently rolled-back operator is an abnormal operator, and the next operator is continued to be tested.
[0042] In this embodiment, through the single-variable isolation strategy, the complex model problem is converged to the hardware adaptation defect of a specific operator, realizing efficient and accurate anomaly location.
[0043] In an exemplary embodiment, after generating the list of operators to be verified according to the operators called during the running of the target model, it further includes: generating a configuration file according to the list of operators to be verified; rolling back the operators in the list of operators to be verified to the verification computing device one by one, including: parsing the configuration file and rolling back the operators in the configuration file to the verification computing device one by one.
[0044] In this embodiment, after generating the list of operators to be verified according to the operators called during the running of the target model, the abnormal operators are further located by generating a configuration file and rolling back the operators to the verification computing device one by one.
[0045] Generating a configuration file (such as a yaml configuration file) is to save the list of operators to be verified in a structured form and ensure that the subsequent operations have clear steps. The configuration file can include detailed information of the operators to be verified, such as the string name of each operator, related parameters, information of the verification computing device, and the detailed process of how to roll back these operators to the verification computing device one by one. The configuration file helps to associate the operators to be verified with specific rollback steps or verification computing devices, ensuring the operability and traceability during the verification process.
[0046] Next, the operators in the operator list to be verified will be rolled back to the verification computing device one by one. This process requires parsing the configuration file and performing specific operations according to the definitions therein. The operator information in the configuration file is parsed out, and according to the order indicated in the configuration file and the requirements of the verification computing device, the operators are rolled back from the current computing device to the specified verification computing device one by one. The rollback process is actually to transfer a certain operator in the target model from the current running platform (such as TPU) to a known stable platform (such as CPU). By executing the operator on the verification computing device and comparing the difference between the verification result and the benchmark result, it can be determined whether there are problems in the implementation of the operator on different computing devices.
[0047] Among them, the format of the configuration file can be:
[0048] yaml;
[0049] name: UNREGISTER_OPERS;
[0050] value: add, mm, conv, …;
[0051] The name of the yaml configuration file item is UNREGISTER_OPERS, and the value of the item is the operator list to be verified, such as add, mm, conv. When registering operators, the UNREGISTER_OPERS information in the yaml configuration file is read dynamically to determine which operators need to be rolled back to the third computing device for execution. Based on this configuration, conditional judgment and the RegisterPlugin function are used to control the registration and rollback of operators.
[0052] The core principle of this process is to manage the complex operator rollback steps structurally through the configuration file and perform rollback verification one by one according to the configuration. In this way, the potential problems of each operator can be effectively isolated, avoiding omissions in large-scale model verification, and ensuring the consistency and stability of each operator when running on different hardware platforms.
[0053] In an exemplary embodiment, parsing the configuration file and rolling back the operators in the configuration file to the verification computing device one by one includes: parsing the configuration file to obtain the string name of the operator in the configuration file; during the process of registering operators, skipping the registration of the operator in the configuration file according to the string name of the operator in the configuration file, so that when running the target model, the operators in the configuration file are rolled back to the verification computing device one by one.
[0054] In this embodiment, the process of parsing the configuration file and rolling back each operator in the configuration file to the verification computing device is mainly described. Specifically, the purpose of parsing the configuration file can be to extract the string names of the operators contained therein, and these string names are the unique identifiers for identifying the operators to be verified. The configuration file lists the names of these operators and their related information in a structured form. By parsing these string names, it can be known which operators need to be rolled back to the verification computing device. During the process of registering operators, according to the operator string names in the configuration file, the scheduling layer will skip the registration of the operators in the configuration file according to predefined rules. Usually, operator registration is an important step in model running, which registers each operator in the model into the execution environment to ensure that they can execute as expected. However, in this step, since the goal is to roll back each operator in the configuration file to the verification computing device instead of executing on the current computing device, the registration of these operators needs to be skipped. By skipping the registration of the operators in the configuration file, the operators in the configuration file will not be executed on the current device but will be rolled back to the verification computing device for execution. In this process, the configuration file plays a guiding role to ensure that each operator is rolled back to the verification computing device one by one in order.
[0055] This mechanism ensures that the verification process of operators can be carried out one by one in a strictly controlled environment, not only avoiding redundant calculations of operators on the current device, but also accurately detecting potential problems of operators, and finally achieving efficient exception location.
[0056] In an exemplary embodiment, it further includes: a predefined operator registry table, which is used to store the mapping relationship between the string names of operators and the corresponding execution logics; during the process of registering operators, according to the string names of the operators in the configuration file, skipping the registration of the operators in the configuration file includes: when filling the operator registry table with a preset registration function, skipping the filling of the operator corresponding to the target string name, and filling the operator registry table according to other operators except the operators in the configuration file; the target string name is the string name of the operator in the configuration file; rolling back each operator in the configuration file to the verification computing device when running the target model includes: when running the target model, calling the filled operator registry table; the filled operator registry table does not include the operators in the configuration file; according to the filled operator registry table, rolling back each operator in the configuration file to the verification computing device.
[0057] In this embodiment, the purpose is to flexibly control the registration and fallback operations of operators through the cooperation of an operator registry and a configuration file, so as to achieve the verification of the target model. Specifically, first, an operator registry is predefined. This operator registry is used to store the string names of operators and their corresponding execution logics. The role of the operator registry is to quickly find and execute operator operations through a mapping relationship. During the process of registering operators, first, the configuration file is parsed to extract the string names of the operators listed therein, and based on these names, it is determined which operators need to be skipped from registration. Specifically, when calling the preset registration function to populate the operator registry, according to the operator string names in the configuration file, these operators in the configuration file are avoided from being registered into the operator registry, that is, their registration operations are skipped. In this way, only the operators not listed in the configuration file will be normally registered into the operator registry. Among them, the target string name refers to the name of the operator listed in the configuration file. During the registration process, these operators will be skipped and not filled into the registry.
[0058] When running the target model, the populated operator registry is called. Since the operator registry does not contain the operators skipped in the configuration file, these skipped operators will not directly participate in the normal calculation process (that is, they will not be calculated on the second computing device), but will be individually fallback to the verification computing device. The purpose of the fallback is to migrate the operators specified in the configuration file to the verification computing device for further verification. On the verification computing device, these operators can be independently executed, tested, and verified to ensure the computational correctness in a specific environment.
[0059] Among them, the preset registration function can be implemented by defining a RegisterPlugin function to populate the operator registry. In this function, by traversing the operators and their corresponding execution logics, the string names of the operators are used as keys (key) with the map data structure, and the function pointers of the corresponding execution logics are used as values (value) to be populated into the registry. The RegisterPlugin function dynamically registers all operators and their execution logics into the system in this way, so that during the subsequent running of the target model, the corresponding logic can be efficiently found and executed according to the operator name. This mechanism simplifies the operator registration process and can be flexibly extended.
[0060] In an exemplary embodiment, the operator registry is configured as a key-value pair structure; among them, the key is the string name of the operator, and the value is the function corresponding to the execution logic of the operator.
[0061] In this embodiment, the operator registry is designed as a key-value pair structure (for example, using std::map to store the mapping between the operator name and the corresponding execution function, and populating it through the RegisterPlugin function). The key is the string name of the operator, and the value is the function pointer of the execution logic corresponding to the operator. Specifically, the operator string name serves as the key of the registry. In this way, each operator can be identified by its string name, and the execution logic associated with the operator can be quickly accessed. The value is the function pointer pointing to the operator execution logic, ensuring that the corresponding function can be called to execute the operator operation when needed.
[0062] When an operator needs to be registered, the string name of the operator and the corresponding function pointer are added to the operator registry through a registration process. The advantage of this structure is that it enables efficient lookup and invocation. When an operator needs to be executed during the operation of the target model, the execution logic (function pointer) corresponding to it can be quickly obtained by looking up the string name (key) in the registry, thus realizing the execution of the operator. This design method avoids defining the same logic code multiple times in the program, improving the maintainability and extensibility of the code.
[0063] For example, assume there are multiple operators such as addition, multiplication, activation functions, etc. Each operator has its own corresponding function. The operator registry associates the string names of these operators (such as "Add", "Multiply", "ReLU", etc.) as keys with the function pointers corresponding to these operators (such as add_func, multiply_func, relu_func). When running the target model, simply querying the operator registry can find and call the corresponding operator execution logic.
[0064] Through this operator registry organized in a key-value pair structure, not only is the efficient management and lookup of operators achieved, but also operators can be flexibly registered or unregistered as needed, enhancing the flexibility and extensibility of operator management. In addition, combined with the configuration switch and the mechanism to skip specific operators, it is possible to precisely control which operators should participate in the calculation and which operators should be excluded during runtime, further optimizing the operator management process.
[0065] It can be seen that this embodiment ensures that the operators in the configuration file can be fully verified without interfering with the normal operation of the target model. It can flexibly control the registration and rollback of operators during the verification process, avoiding unnecessary interference and improving the verification efficiency.
[0066] In an exemplary embodiment, after the predefined operator registry, it further includes: defining a configuration switch corresponding to the operator, where the state of the configuration switch represents whether the corresponding operator is filled in the operator registry, and the state of the configuration switch includes a closed state and an open state; when filling the operator registry using a preset registration function, skipping the filling of the operator corresponding to the target string name, including: setting the configuration switch corresponding to the target string name to the closed state; when filling the operator registry using the preset registration function, skipping the filling of the operator corresponding to the closed state of the configuration switch, and filling the operator registry according to the operator with the open state of the configuration switch.
[0067] In this embodiment, in order to further enhance the control and flexibility of the operator registration process, the concept of a configuration switch is introduced, and each operator corresponds to a configuration switch. The state of this configuration switch is used to indicate whether the operator will be filled into the operator registry. The state of the configuration switch has two types, the open state and the closed state. The open state means that the operator will be registered in the operator registry, while the closed state means that the operator will be skipped and not registered in the operator registry. In this way, it is possible to flexibly control which operators need to participate in the operation of the target model and which operators need to be excluded for specific verification operations.
[0068] Specifically, when filling the operator registry using the preset registration function, first check the state of the configuration switch of each operator. If the state of the configuration switch is the open state, then the operator will be registered in the operator registry according to the normal process, mapping its string name and the corresponding execution logic. If the state of the configuration switch is the closed state, then the operator will be skipped, and the preset registration function will ignore its registration. Specifically, when encountering the operator corresponding to the target string name listed in the configuration file, set the state of the configuration switch corresponding to the operator to the closed state, so that the operator will not be filled into the operator registry and will be prevented from participating in the execution process of the target model on the second computing device.
[0069] It can be seen that in this application, through the configuration switch, it is possible to accurately control which operators should be registered and which operators should be excluded, improving flexibility and configurability. When actually running the target model, only those operators with the open state of the configuration switch will be used and executed, while the operators with the closed configuration switch will be skipped, ensuring that the model can run as expected. In addition, through this configuration switch control, the operator registration process can be combined with the fallback operation of the verification computing device. Without affecting the overall computing process, the operators to be verified can be gradually fallback to the verification computing device one by one, further improving the verification efficiency.
[0070] In an exemplary embodiment, it further includes: recording the input and output parameter information of the abnormal operator and the string name of the abnormal operator; the input and output parameter information includes the tensor shape, stride, and data type of the input and output; generating a corresponding test script for the operator based on the input and output parameter information, the string name, and a preset script; the test script is used to reproduce the calculation process of the abnormal operator.
[0071] In this embodiment, by recording the input and output parameter information of the abnormal operator and the string name of the operator, the runtime behavior of the operator can be obtained. The input and output parameter information includes the shape, stride, and data type of the tensor, and these parameters are important attributes for describing the input and output data of the operator and are crucial for locating and analyzing abnormal behaviors. By collecting this information, all key factors related to the operator can be considered when generating the test script. Next, an operator is generated based on the preset script, and combined with the collected input and output parameter information and the string name, a corresponding test script is automatically generated. The test script will simulate the execution of the abnormal operator under various input conditions, verify whether it works as expected, and be able to capture potential errors or inconsistencies.
[0072] It can be seen that through this automated script generation method in this embodiment, not only the efficiency of operator verification is improved, but also the comprehensiveness and accuracy of the test are ensured, thereby effectively detecting and fixing potential defects of the operator.
[0073] Such as Figure 2As shown in the figure, the automatic troubleshooting model exception program corresponding to the positioning of the exception operator can be divided into the following virtual modules: the model operation module, the operator list processing module, the result processing module, and the operator test script generation module. Among them, the model operation module executes the inference or training task of the target model on heterogeneous computing devices (such as GPUs and TPUs) to generate operation results. For example, when running the target model on a verified device (such as a GPU), the output result is saved as the correct benchmark result; when running the same target model on the device to be verified (such as a TPU), the candidate result is recorded, and the operator call tracking function is enabled to generate the list of unverified operators actually called. The operator list processing module dynamically manages the list of unverified operators called during the operation of the target model, generates and parses the configuration file to control the operator fallback logic. For example, it obtains the unverified list corresponding to the operator name from the model operation module, automatically generates a yaml configuration file, marks the operators that need to fallback to the CPU, and updates the configuration file cyclically according to the verification progress to fallback each operator in the configuration file one by one. The result processing module compares the operation results on different computing devices, quantifies the difference degree, and locates the exception operator. The operator test script generation module automatically generates a reproducible test script according to the input and output parameters of the exception operator. The automatic troubleshooting model exception program is equivalent to the overall control module, coordinating other modules to complete the automatic troubleshooting process of the exception operator, calling the model operation module in sequence to generate results, triggering the operator list processing module to update the configuration, and executing the fallback verification cyclically. According to the difference analysis result of the result processing module, it decides whether to continue the test or terminate the loop.
[0074] The module interaction and overall process are as follows: the model operation module generates the benchmark result and the candidate result on the GPU and TPU respectively; the operator list processing module extracts the list of unverified operators called by the TPU and generates the configuration file; the automatic troubleshooting program sequentially falls back the operators in the configuration file to the CPU, triggering the model operation module to re-execute the target model; the result processing module compares the verification result with the benchmark result to determine the exception operator; if the exception operator is located, the test script generation module generates the debug script and exits, if not, it updates the configuration file and enters the next loop.
[0075] Such as Figure 3 , the front end of the deep learning framework communicates with the back-end scheduling layer through the interface module. The back-end scheduling layer determines whether it is necessary to register operators (such as operator 1, operator 2... operator n) according to the configuration file, and then determines whether it is necessary to fallback the corresponding operators to the verification computing device.
[0076] Such as Figure 4As shown below, the process of a specific embodiment is as follows: Run the target model on the first computing device to generate a baseline result; Generate a list of operators to be verified based on the operators called during the running of the target model; Generate a configuration file according to the list of operators to be verified; Run the target model, and roll back each operator in the configuration file to the verification computing device one by one; Save the verification result; Verify whether the verification result is consistent with the baseline result; If not, re-enter the step of generating a configuration file according to the list of operators to be verified; If so, determine it as an abnormal operator, and automatically generate a test script corresponding to the abnormal operator according to the input and output parameter information and string name, and also re-enter the step of generating a configuration file according to the list of operators to be verified again. It should be understood that re-entering the step of generating a configuration file according to the list of operators to be verified is to verify the next operator to be verified in the list of operators to be verified.
[0077] Specifically, the specific implementation of this specific embodiment is as follows:
[0078] Step 1: Generate baseline results and candidate results: Run the target model on the GPU (the first computing device), input a preset question, and save the output to llama_gpu.txt (the file corresponding to the baseline result); Run the same target model on the TPU (the second computing device), enable the operator name recording function, save the output to llama_tpu.txt (the file corresponding to the candidate result), and generate a list of operators to be verified, opers.txt, according to the difference between the candidate result and the baseline result.
[0079] Step 2: Dynamically generate a yaml configuration file: Read the list of operators to be verified, opers.txt, and extract all operator names (such as add, mm, silu); Create fallback_config.yml and mark the operators that need to be rolled back.
[0080] Step 3: Roll back operators one by one and verify: The scheduling layer reads fallback_config.yml and rolls back each operator to the CPU (the verification computing device) in turn; The verification process is as follows: When rolling back the add operator for the first time, modify the configuration file to only contain add, run the model and save the result to llama_tpu_fallback_add.txt (the file corresponding to the verification result); Compare llama_tpu_fallback_add.txt (the file corresponding to the verification result) with llama_gpu.txt (the file corresponding to the baseline result). If the results are consistent, it means that the add implementation on the TPU is abnormal. If the results are inconsistent, continue to roll back the next operator (such as mm); Repeat the execution until all the operators in the list of operators to be verified are tested.
[0081] Step 4: Abnormal operator localization and test script generation: Assume that the result is consistent with the baseline result after falling back to mm, and determine mm as the abnormal operator; Enable the logging function, save the input and output shapes, strides, and data types of the mm operator to mm_info.yaml (input and output parameter information of the abnormal operator), and automatically generate test_mm.py (test script) according to mm_info.yaml for reproducing the abnormality.
[0082] Among them, for the scheduling layer and the operator registry, the registry structure uses std::map, with the key being the string name of the operator (such as mm) and the value being a pointer to the function in the adaptation layer. The registry is populated through the RegisterPlugin function, skipping the operators marked in the configuration file, registering the TPU operators to the PyTorch scheduling layer, and triggering CPU fallback for unregistered operators. The fallback function calls the native CPU implementation of PyTorch.
[0083] The automated implementation process of this embodiment can quickly locate abnormal operators, directly lock the implementation defects of operators on the TPU through the single-variable isolation strategy, and is fully automated from result comparison to test script generation, reducing manual intervention.
[0084] This embodiment realizes efficient and accurate abnormal operator localization by dynamically controlling operator registration in the scheduling layer and managing the fallback logic with a yaml configuration file, combined with result comparison and automated script generation. It is particularly suitable for the debugging scenario when large models are deployed on heterogeneous hardware (such as GPU / TPU), significantly improving development efficiency and system stability.
[0085] In an exemplary embodiment, it further includes: obtaining the execution process of the abnormal operator in the target model according to the running result of the test script; analyzing its performance characteristics based on the execution process of the abnormal operator; adjusting the execution strategy of the abnormal operator according to the input parameters and calculation conditions, and the adjustment strategies include: dynamically adjusting the parallelism of the operator according to the size of the input data, the load of the hardware resources, and the complexity of the calculation task; dynamically adjusting the batch processing size of the operator according to the characteristics of the input data and the resource usage of the calculation device; dynamically allocating or scheduling the hardware resources of the calculation device according to factors such as the current load, memory usage, and bandwidth requirements of the calculation device; using machine learning algorithms to predict the optimal execution strategy based on historical operator execution data and performance metrics, and real-time optimizing the execution strategy of the target operator; automatically triggering adjustment operations when abnormalities are detected during the operator execution process; tracking the actual effect of the adjusted execution strategy on the operator through a real-time monitoring system, and further optimizing the adjustment strategy according to the feedback; and automatically recording the adjusted execution strategy for future execution processes of the same operator, improving the execution efficiency of the operator and reducing the possibility of abnormalities.
[0086] In this embodiment, through a comprehensive analysis of the execution process of abnormal operators, combined with the performance characteristics and input / output parameters of the operators, by dynamically adjusting the parallelism, batch processing size, and hardware resource allocation of the operators, it is possible to better adapt to different input data and computing conditions, optimize the execution efficiency of the operators, avoid excessive or insufficient resource allocation, and improve the overall running performance of the model. Based on real-time monitoring and performance characteristic analysis, it is possible to adjust the execution strategy in a timely manner when an anomaly is detected, reducing the occurrence of anomalies caused by mismatched hardware resources or unbalanced computing task loads, thereby improving the stability of the model. By using machine learning algorithms to analyze historical operator execution data and performance metrics, it is possible to predict the optimal execution strategy of the operators and perform real-time optimization. This adaptive optimization mechanism can continuously improve the execution strategy as the usage time increases, enabling the operators to exhibit better computing capabilities in different environments. Dynamically allocating resources according to factors such as the current load, memory usage, and bandwidth requirements of the computing device can effectively avoid resource bottlenecks, balance the computing requirements and resource usage of each operator, thereby reducing the overload situation and runtime errors of the computing device. By tracking the actual effects of the adjusted execution strategy on the operators through a real-time monitoring system, adjustments and optimizations can be made in a timely manner based on the feedback. This makes the entire adjustment process more targeted and flexible, capable of coping with changing computing requirements. Automatically record the adjusted execution strategy for reuse during the execution process of the same operator in the future. This can not only provide experience and reference for the execution process of the same operator in the future, but also reduce the debugging and adjustment time during each execution, improving the overall efficiency and consistency of the system. Through the automated adjustment strategy, the dependence on manual intervention is reduced. The detection, analysis, and optimization of abnormal operators can all be carried out automatically, improving the debugging efficiency and reducing the possibility of human errors.
[0087] Generally speaking, this embodiment provides a more efficient and stable execution environment for abnormal operators by combining performance analysis, dynamic adjustment, and intelligent optimization, not only optimizing the execution efficiency of the operators, but also enhancing the adaptability.
[0088] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.
[0089] As Figure 5 shown, the embodiment of the present application also provides an electronic device, including a memory 51 and a processor 52. The memory 51 stores a computer program, and the processor 52 is configured to run the computer program to execute the steps in any of the above-described embodiments of the method for locating abnormal operators.
[0090] For the description of the features in the embodiments corresponding to the electronic device, reference can be made to the relevant descriptions in the embodiments corresponding to the method for locating abnormal operators, which will not be elaborated herein one by one.
[0091] As Figure 6 shown, an embodiment of the present application further provides a computer-readable storage medium 61, in which a computer program 62 is stored. Among them, the computer program 62 is configured to execute the steps in any of the above-described embodiments of the method for locating abnormal operators when running.
[0092] In an exemplary embodiment, the above computer-readable storage medium 61 may include but is not limited to: various media such as USB flash drives, read-only memories (ROM for short), random access memories (RAM for short), mobile hard disks, magnetic disks, or optical discs that can store the computer program 62.
[0093] For the description of the features in the embodiments corresponding to the computer-readable storage medium 61, reference can be made to the relevant descriptions in the embodiments corresponding to the method for locating abnormal operators, which will not be elaborated herein one by one.
[0094] An embodiment of the present application further provides a computer program product. The above computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-described embodiments of the method for locating abnormal operators.
[0095] An embodiment of the present application further provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-described embodiments of the method for locating abnormal operators.
[0096] For the description of the features in the embodiments corresponding to the computer program product, reference can be made to the relevant descriptions in the embodiments corresponding to the method for locating abnormal operators, which will not be elaborated herein one by one.
[0097] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0098] The above has introduced in detail a method, device, storage medium and program product for locating an abnormal operator provided by the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.
Claims
1. A method for locating an abnormal operator, characterized in that: include: Obtaining running results of a target model on at least two different computing devices, the running results comprising a benchmark result obtained by running the target model on a first computing device and a candidate result obtained by running the target model on a second computing device; When it is determined that an abnormal operator exists according to the difference of the running results, a list of operators to be verified is generated according to the operators called during the running of the target model; The abnormal operator is at least one operator called during the operation of the target model; Returning the operators in the list of operators to be verified to the verification computing device one by one, so that the verification computing device runs the target model according to the returned operators to obtain a verification result; the verification computing device is a verified and stable third computing device; Determining whether the verification result is the same as the benchmark result; If the verification result is the same as the benchmark result, the currently rolled back operator is determined to be the abnormal operator; If the verification result is different from the benchmark result, the currently rolled back operator is determined to be a normal operator; Returning the operators in the list of operators to be verified to the verification computing device one by one, so that the verification computing device runs the target model according to the returned operators to obtain a verification result, including: Each operator in the list of operators to be verified is returned to the verification computing device in turn to run the target model, and other operators except the operators in the list of operators to be verified are retained on the second computing device to run the target model, and the result obtained by the verification computing device running the target model is determined as the verification result.
2. The method for locating an abnormal operator according to claim 1, characterized in that: Determining the existence of an abnormal operator according to the difference of the running results includes: Determine whether the difference between the candidate result and the benchmark result meets a preset difference condition; If satisfied, it is determined that an abnormal operator exists when the second computing device runs the target model.
3. The method for locating an abnormal operator according to claim 2, characterized in that: The first computing device is a reference computing device whose running results have been verified, and the second computing device is a computing device to be verified.
4. The method for locating an abnormal operator according to claim 2, characterized in that: Determining whether the difference between the candidate result and the benchmark result meets a preset difference condition includes: Determine whether the difference between the candidate result and the benchmark result is greater than a preset threshold; If so, it is determined that the preset difference condition is met.
5. The method for locating an abnormal operator according to claim 1, characterized in that: After generating a list of operators to be verified according to the operators called during the operation of the target model, the method further includes: Generate a configuration file according to the list of operators to be verified; Returning the operators in the list of operators to be verified to the verification computing device one by one includes: The configuration file is parsed, and the operators in the configuration file are rolled back to the verification computing device one by one.
6. The method for locating an abnormal operator according to claim 5, characterized in that: Parsing the configuration file and returning the operators in the configuration file to the verification computing device one by one includes: Parse the configuration file to obtain the string name of the operator in the configuration file; In the process of registering operators, registration of operators in the configuration file is skipped according to the string names of the operators in the configuration file, so that the operators in the configuration file are rolled back one by one to the verification computing device when the target model is run.
7. The method for locating an abnormal operator according to claim 6, characterized in that: Also includes: A predefined operator registry, which is used to store a mapping relationship between an operator string name and a corresponding execution logic; In the process of registering an operator, according to the string name of the operator in the configuration file, the registration of the operator in the configuration file is skipped, including: When filling the operator registration table using the preset registration function, skip filling the operator corresponding to the target string name, and fill the operator registration table according to other operators other than the operator in the configuration file; the target string name is the string name of the operator in the configuration file; When running the target model, the operators in the configuration file are rolled back to the verification computing device one by one, including: When running the target model, calling the populated operator registry; the populated operator registry does not include the operator in the configuration file; According to the populated operator registry, the operators in the configuration file are rolled back to the verification computing device one by one.
8. The method for locating an abnormal operator according to claim 7, characterized in that: After the predefined operator registry, it also includes: Define a configuration switch corresponding to the operator, the state of the configuration switch represents whether the corresponding operator is filled in the operator registration table, and the state of the configuration switch includes a closed state and an open state; When the operator registration table is filled using the preset registration function, the filling of the operator corresponding to the target string name is skipped, including: Setting the configuration switch corresponding to the target string name to the closed state; When the operator registration table is filled using the preset registration function, the filling of the operator corresponding to the closed state of the configuration switch is skipped, and the operator registration table is filled according to the operator whose state of the configuration switch is the open state.
9. The method for locating an abnormal operator according to claim 7, characterized in that: The operator registry is configured as a key-value pair structure, wherein the key is the string name of the operator and the value is the function corresponding to the execution logic of the operator.
10. The method for locating an abnormal operator according to any one of claims 1 to 9, characterized in that: Also includes: Record the input and output parameter information of the abnormal operator and the string name of the abnormal operator; The input and output parameter information includes input and output tensor shape, step size and data type; Generate a corresponding test script according to the input and output parameter information, the string name, and a preset script generation operator; The test script is used to reproduce the calculation process of the abnormal operator.
11. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the method for locating an abnormal operator according to any one of claims 1 to 10 when executing the computer program.
12. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the method for locating an abnormal operator according to any one of claims 1 to 10.
13. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method for locating an abnormal operator according to any one of claims 1 to 10 are implemented.
Citation Information
Patent Citations
Anomaly detection apparatus, anomaly detection method and program
US20220284332A1
System, method, and computer program product for location aware device fault detection
US20230214287A1