Distributed training framework upgrading verification method and related device

By converting the weak alignment operator in the distributed training framework into a strong alignment operator and using hash value comparison, the accuracy alignment problem during the version of the distributed training framework is solved, and the accuracy and reliability of upgrade verification are achieved.

CN120012875APending Publication Date: 2025-05-16BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411900213.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-20
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

When the version of the existing distributed training framework is changed, it is difficult to effectively verify the accuracy alignment of the training framework, resulting in the failure of upgrade verification.

Method used

By converting the weak alignment operator into an equivalent strong alignment operator, and determining the upgrade verification result of the distributed training framework based on the comparison results of the first and second hash values ​​of each operator.

Benefits of technology

It effectively solves the problem of accuracy misalignment caused by inconsistent operator results, and ensures the accuracy and reliability of the version upgrade verification of the distributed training framework.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120012875A_ABST
    Figure CN120012875A_ABST
Patent Text Reader

Abstract

The invention discloses a distributed training framework upgrading verification method and a related device. The method comprises the following steps: converting a weak alignment operator in a distributed training framework into an equivalent strong alignment operator; the weak alignment operator is an operator with inconsistent results when the same processing object is processed in the old version and the new version, and the strong alignment operator is an operator with consistent results when the same processing object is processed in the old version and the new version; determining an upgrade verification result of the distributed training framework according to a comparison result of the first hash value and the second hash value of each operator of the distributed training framework; wherein the first hash value is the hash value of the execution result of each operator under the old version, and the second hash value is the hash value of the execution result of each operator under the new version. According to the method and the device, the cross-version consistency of operators is ensured, and the version upgrading verification efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, specifically to the field of distributed training of artificial intelligence models, and more particularly to a distributed training framework upgrade verification method and device, electronic device, computer-readable storage medium, and computer program product. Background Art

[0002] As the scale of deep learning models has shown a rapid expansion trend, it has achieved remarkable results in many fields such as natural language, image and audio. However, such models often have hundreds of millions of parameters, and their training process requires a massive amount of data samples. Such a huge demand makes it difficult for a single-machine computing environment to bear it, which has spawned a distributed training framework based on data parallel synchronous communication mechanism.

[0003] When the current distributed training framework faces continuous version upgrades, it is necessary to verify that the training frameworks between different versions are completely consistent. This process is the precision alignment of the training framework. However, due to various reasons, the calculation results of the training frameworks between different versions may be different, which makes it impossible to align the precision of different versions of the distributed training framework, resulting in the inability to effectively verify whether the distributed training framework upgrade is successful. Summary of the invention

[0004] The present disclosure provides a distributed training framework upgrade verification method and device, an electronic device, a computer-readable storage medium, and a computer program product.

[0005] According to a first aspect, a distributed training framework upgrade verification method is provided, which is used in the process of upgrading an old version of the distributed training framework to a new version, and is characterized in that the processing object of the distributed training framework includes at least any one of the following: text, image, audio or video data; the method includes: converting a weak alignment operator in the distributed training framework into an equivalent strong alignment operator; the weak alignment operator is an operator whose results are inconsistent when processing the same processing object in the old version and the new version, and the strong alignment operator is an operator whose results are consistent when processing the same processing object in the old version and the new version; determining the upgrade verification result of the distributed training framework based on the comparison result of the first hash value and the second hash value of each operator of the distributed training framework; wherein the first hash value is the hash value of the execution result of each operator under the old version, and the second hash value is the hash value of the execution result of each operator under the new version.

[0006] According to the second aspect, a distributed training framework upgrade verification device is provided, which is used in the process of upgrading an old version of the distributed training framework to a new version, characterized in that the processing object of the distributed training framework includes at least any one of the following: text, image, audio or video data; the device includes: a strong alignment operator conversion unit, configured to convert a weak alignment operator in the distributed training framework into an equivalent strong alignment operator; the weak alignment operator is an operator whose results are inconsistent when processing the same processing object in the old version and the new version, and the strong alignment operator is an operator whose results are consistent when processing the same processing object in the old version and the new version; a framework upgrade verification unit, configured to determine the upgrade verification result of the distributed training framework based on the comparison result of the first hash value and the second hash value of each operator of the distributed training framework; wherein the first hash value is the hash value of the execution result of each operator under the old version, and the second hash value is the hash value of the execution result of each operator under the new version.

[0007] According to a third aspect, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method described in any implementation manner of the first aspect.

[0008] According to a fourth aspect, a non-transitory computer-readable storage medium storing computer instructions is provided, where the computer instructions are used to cause a computer to execute the method described in any implementation of the first aspect.

[0009] According to a fifth aspect, a computer program product is provided, comprising a computer program, which implements the method described in any implementation manner of the first aspect when executed by a processor.

[0010] The technical solution disclosed in the present invention specifically solves the key problems of precision alignment and upgrade verification faced by the distributed training framework during version updates, which helps to improve the quality and reliability of version upgrades of the distributed training framework and promote its better application in many fields such as natural language, images, audio, and video.

[0011] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.

[0013] Figure 1is a schematic diagram of a distributed training framework provided by an embodiment of the present disclosure;

[0014] Figure 2 is a flow chart of a distributed training framework upgrade verification method provided by an embodiment of the present disclosure;

[0015] Figure 3 is a flow chart of a distributed training framework upgrade verification method provided by an embodiment of the present disclosure;

[0016] Figure 4 is a flow chart of a distributed training framework upgrade verification method provided by an embodiment of the present disclosure;

[0017] Figure 5 is a flow chart of a distributed training framework upgrade verification method provided by an embodiment of the present disclosure;

[0018] Figure 6 is a flow chart of a distributed training framework upgrade verification method provided by an embodiment of the present disclosure;

[0019] Figure 7 is a flow chart of a distributed training framework upgrade verification method provided by an embodiment of the present disclosure;

[0020] Figure 8 It is a structural diagram of a distributed training framework upgrade verification device provided by an embodiment of the present disclosure;

[0021] Fig. 9 It is a block diagram of an electronic device used to implement the distributed training framework upgrade verification method of the embodiment of the present disclosure. DETAILED DESCRIPTION

[0022] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0023] The present disclosure proposes a distributed training framework upgrade verification method and related devices.

[0024] It is understandable that the method can be integrated into various computing devices, including but not limited to terminal devices with text processing capabilities (such as smart phones, tablets, personal computers, etc.) and servers (such as application servers, Web servers, whether single servers or cluster-deployed server systems). The method does not rely on a specific hardware platform or software architecture. During the execution of the method, both independently running terminal devices and terminal-server architectures that work collaboratively through a network can effectively utilize the method disclosed in the present invention. When the terminal device is executed independently, the method can be independent of the external network. In scenarios where higher processing performance or wider data resource support is required, the method of the present invention also supports communication between the terminal device and the server, utilizing the powerful computing power and rich data resources of the server to jointly complete the method. The method can adapt to different operating systems and platform environments, including mobile operating systems such as iOS, Android, and Hongmeng, desktop operating systems such as Windows and macOS, and server operating systems such as Linux and Unix.

[0025] It can be understood that the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved in the technical solution of the present disclosure are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0026] When the distributed training framework in the prior art is upgraded, there will be inconsistent precision caused by different underlying implementation methods of the underlying communication library of tensor communication. Within the scope of the communication mode of distributed training, the communication modes of some communication operators all involve tensor accumulation operations. Different underlying communication libraries use different methods to implement these communication modes. Even if these protocol communications are equivalently convertible from a logical level, in the actual communication process, due to the different tensor accumulation methods, slight numerical differences will occur, resulting in precision misalignment problems. These weakly aligned operators produce inconsistent results when processing the same processing object in the old version and the new version, which makes the upgrade verification of the distributed training framework unsuccessful.

[0027] In the distributed training framework, there are also precision misalignments caused by non-deterministic operators. The output results of some operators during tensor calculations are uncertain, which leads to differences in tensor calculation precision under the same model and hyperparameter conditions. These non-deterministic operators have uncertain results when processing the same processing objects in the old and new versions, resulting in the failure of the upgrade verification of the distributed training framework.

[0028] Embodiment 1

[0029] For ease of understanding, Figure 1 A schematic diagram of a scenario of an embodiment of the present disclosure is shown.

[0030] The processing objects of the distributed training framework include at least any one of the following: text, image, audio or video data.

[0031] In the distributed training framework, there are two important components: computing operators and communication operators. Among them, computing operators can be further divided into two categories: deterministic computing operators and non-deterministic computing operators. The non-deterministic operator is an operator whose result is uncertain when the same operation is performed in the old version and the new version, and the deterministic operator is an operator whose result is determined when the same operation is performed in the old version and the new version.

[0032] like Figure 1 As shown, in the upgrade verification method 100 of the distributed training framework, the input data is input into the deterministic calculation operator 103 to obtain the output data of the operator 103; the output data of the operator 103 is input into the strong alignment operator 106 to obtain the output data of the operator 106. The above calculation process is calculated once in the old framework and the new framework respectively. If for the same input data, the final output data of the old framework and the new framework are exactly the same, it can be considered that the old framework is equivalent to the new framework (that is, the upgrade from the old framework to the new framework is successful). If for the same input data, the final output data of the old framework and the new framework are inconsistent, it can be considered that the old framework is not equivalent to the new framework (that is, the upgrade from the old framework to the new framework is unsuccessful).

[0033] If the old framework or the new framework contains non-deterministic operators or misaligned operators, the calculation results will be uncertain or the precision will not be aligned. At this time, for the same input data, the old framework and the new framework will calculate it once respectively, and the final output data will be different. In other words, the old framework and the new framework are not equivalent. Even if the old framework and the new framework are actually equivalent, due to the uncertainty of the operator results or the misalignment of precision, the wrong conclusion that the old framework and the new framework are not equivalent will be drawn.

[0034] The present disclosure adopts a specific processing method, that is, the non-deterministic operator 101 is converted into an equivalent deterministic calculation operator 103 through a deterministic operator mapper 102. Then, the deterministic calculation operator 103 is used to perform processing operations on the relevant processing objects, and a certain output result will be obtained.

[0035] On this basis, with the help of a specific detection method, the weak alignment operator 104 is accurately detected from the operators of the distributed training framework. Then, through the communication mapper 105, the detected weak alignment operator 104 is further mapped to a precision-aligned communication operator, which we call a strong alignment operator 106. The weak alignment operator is an operator that produces inconsistent results when processing the same processing object in the old version and the new version, and the strong alignment operator 106 is an operator that produces consistent results when processing the same processing object in the old version and the new version.

[0036] After the above series of operations, factors that interfere with the framework upgrade verification are eliminated as much as possible, and the framework upgrade verification of different versions can be continued. For each operator of the distributed training framework, obtain the first hash value of the execution result of each operator in the old version, and obtain the second hash value of the execution result of each operator in the new version; obtain the comparison result of the first hash value and the second hash value of each operator, and determine the upgrade verification result of the distributed training framework based on all the comparison results.

[0037] The distributed training framework upgrade verification method provided in this embodiment has the following technical effects:

[0038] First, solve the problem of inconsistent precision. In the prior art, due to the non-deterministic algorithms and different implementation methods of the underlying communication library of tensor communication, the precision of different versions of distributed training frameworks will be inconsistent. The present invention converts non-deterministic operators into equivalent deterministic operators through deterministic operator conversion, eliminating the precision difference caused by the uncertainty of the output results of some operators; at the same time, the weak alignment operator is converted into a strong alignment operator, which solves the precision problem caused by the inconsistent results of the same operation performed by the operator in the old version and the new version, thereby effectively improving the situation of inconsistent precision of training frameworks between different versions.

[0039] Secondly, implement framework upgrade verification. The present disclosure obtains the hash values ​​of the execution results of each operator of the distributed training framework under the old version and the new version and compares them to determine the upgrade verification results of the distributed training framework. This enables effective verification of the upgrade process of the distributed training framework from the old version to the new version, overcoming the difficulty of verifying whether the correctness of the new version is consistent with the old version due to the inability to obtain completely consistent deterministic calculation results from different previous versions, providing a reliable verification mechanism for the accurate upgrade of the distributed training framework, and ensuring that the upgraded training framework meets the expected requirements in terms of function and accuracy.

[0040] Finally, based on the above framework upgrade verification mechanism, this paper uses a comprehensive accuracy comparison process to carefully compare the output of the new framework with the output of the old framework as a benchmark, layer by layer and operator by operator. Through this high-precision comparison, developers can quickly and accurately locate operators with accuracy problems, so as to carry out targeted repairs and optimizations, greatly improving development efficiency and model accuracy stability.

[0041] Embodiment 2

[0042] Figure 2 A distributed training framework upgrade verification method 200 according to an embodiment of the present disclosure is shown, which is used for the process of upgrading an old version of the distributed training framework to a new version. The processing object of the distributed training framework includes at least any one of the following: text, image, audio or video data.

[0043] The method 200 includes: step 201, converting the weak alignment operator in the distributed training framework into an equivalent strong alignment operator; the weak alignment operator is an operator that produces inconsistent results when processing the same processing object in the old version and the new version, and the strong alignment operator is an operator that produces consistent results when processing the same processing object in the old version and the new version.

[0044] Weakly aligned operators: In both the old and new versions of the distributed training framework, weakly aligned operators refer to operators that produce inconsistent output results when performing the same operation on the same processing object.

[0045] This inconsistency may be due to adjustments made to the internal implementation details of the operator in different versions, such as changes in the optimization strategy of the algorithm, differences in data processing methods, or changes in the calls to the underlying library functions. Even if the functions of the operators seem to be the same logically, the actual execution results are different. Taking the image feature extraction operator as an example, the old version may use a simple feature extraction algorithm to implement it, while the new version uses a more complex optimization algorithm to improve efficiency. This may result in the extracted feature vectors being numerically different when processing the same image. Although these features may be semantically similar, from the perspective of operator result consistency, it is a weak alignment operator.

[0046] Another manifestation of this inconsistency is in the communication link of the distributed training framework. In distributed training, multiple processes need to communicate frequently to synchronize data. Typical communication modes include Broadcast, Scatter, Gather, AllGather, Reduce, AllReduce, and ReduceScatter. Among them, the three communication modes of Reduce, AllGather, and ReduceScatter involve the accumulation of tensors. Because the implementation methods of the underlying communication library are not identical, the protocol communication that can be converted to each other logically will produce slight differences in values ​​in actual communication due to different tensor accumulation methods. This difference is the reason why the communication results in the old and new architectures of the deep learning framework are different.

[0047] Multiple alignment operators can be converted to equivalent strong alignment operators. Strong alignment operators are operators that can ensure that the execution results are completely consistent when the old and new versions process the same processing object. This means that both the functional logic of the operator and its specific implementation details have been specially processed or adjusted between the two versions to ensure the consistency of the results.

[0048] Taking the widely used GPU accelerator card as an example, in the deep learning training process, the implementation of communication operators such as Reduce, AllGather, and ReduceScatter will cause slight deviations in accuracy due to algorithm differences, which are weakly aligned operators. In order to eliminate these potential accuracy inconsistencies, the communication mapper will map all these operators that may produce accuracy differences to the AllReduce operator. As a strongly aligned communication operation, AllReduce can ensure that exactly the same communication results are obtained on all nodes participating in the calculation, thereby effectively avoiding the accuracy problems caused by differences in communication operators.

[0049] Converting weak alignment operators to strong alignment operators can effectively solve the problem caused by inconsistent results between operators in different versions. In distributed training, especially when it comes to large-scale data and complex model training, the consistency of the calculation results of each operator is crucial to the accuracy and stability of the entire training process. For example, in the training of deep learning models for audio signal processing, if a weakly aligned audio feature transformation operator in different versions of the training framework leads to differences in feature data, then subsequent model training and parameter updates based on these features will be affected, which may eventually lead to a decrease in model performance or instability. After the strong alignment operator is converted, whether the old or new version of the framework is used, the same audio data will have the same result after being processed by the operator, thus ensuring the consistency and comparability of the entire training process between different versions.

[0050] In practical applications, upgrading the training framework is a common requirement. When migrating a model trained on an old version of a distributed training framework to a new version of the framework for continued training or application, if there is a weak alignment operator that causes inconsistent results, the model may not work properly or its performance may drop significantly. The strong alignment operator conversion makes the migration of models between different versions of the framework smoother, reducing the workload of making a lot of adjustments to the model due to operator differences. For example, a text classification model trained on an old version of the framework in a natural language processing task can be more easily deployed on a new version of the framework for online updates or extended training, because the strong alignment operator ensures the consistency of the old and new versions of the framework when processing the same text data, which can eliminate the influence of the weak alignment operator, making the comparison between the old and new frameworks simpler and more direct.

[0051] The method 200 also includes: step 202, determining the upgrade verification result of the distributed training framework based on the comparison result of the first hash value and the second hash value of each operator of the distributed training framework; wherein the first hash value is the hash value of the execution result of each operator under the old version, and the second hash value is the hash value of the execution result of each operator under the new version.

[0052] The operators in the above distributed training framework refer to all operators involved in the framework operation, including but not limited to the aforementioned non-deterministic operators, deterministic operators, weak alignment operators, and strong alignment operators. In addition, other operators outside the above categories are also included.

[0053] Hash value: For each operator in the distributed training framework, after executing the operator to process a specific processing object under the framework, the execution result is calculated through a hash algorithm (such as common MD5, SHA-1 and other hash algorithms) to obtain a hash value of fixed length. This hash value is a characteristic representation of the execution result of the old version of the operator. It is unique and irreversible, that is, different execution results usually produce different hash values, and the original execution result data cannot be deduced from the hash value.

[0054] Considering the huge size of tensors in deep learning calculations, directly storing all intermediate results not only takes up a lot of storage space, but also is not convenient for subsequent comparative analysis. Therefore, the PaddlePaddle framework adopts an efficient data summary method: instead of directly saving tensor data, its MD5 hash value is calculated for storage. This method significantly reduces the burden on the storage system and speeds up the result comparison speed during the precision alignment process.

[0055] The same processing object is processed, and the hash value of the execution result of the old framework is extracted as the first hash value. The hash value of the execution result of the new framework is extracted as the second hash value. The two are two hash values ​​of the execution results obtained after the same operator processes the same processing object. They reflect the calculation characteristics and result characteristics of the operator in the old and new versions of the framework. By comparing the first hash value and the second hash value, it can be determined whether the calculation results of the operator in the old and new versions of the framework have changed.

[0056] By obtaining the hash values ​​of the execution results of each operator in the old version and the new version and comparing them, it is possible to accurately detect whether the calculation results of each operator have changed during the framework upgrade process. This hash value-based verification method is efficient and accurate because the hash algorithm can quickly convert complex execution result data into fixed-length feature values ​​for comparison. If the first hash value and the second hash value of an operator are the same, it means that the calculation results of the operator in the old and new versions of the framework are consistent, otherwise it means there is a difference.

[0057] Framework upgrade verification does not only focus on the changes of a single operator, but also comprehensively verifies all operators in the entire distributed training framework. This makes it possible to evaluate the quality and effect of the framework upgrade as a whole. If the calculation results of a large number of operators have changed, it is necessary to deeply analyze whether these changes are caused by functional improvements, performance optimizations, or potential errors. For example, in a deep learning model training framework that contains multiple neural network layers, hash value verification of the operators of each layer can clearly understand which layers' calculations have changed after the upgrade, providing a detailed basis for further analysis of the impact of framework upgrades on model training.

[0058] In summary, the two steps in the above method 200 each have clear technical connotations and important technical effects. They cooperate and work synergistically with each other to provide a complete verification solution for the safe, stable and efficient upgrade of the distributed training framework.

[0059] Embodiment 3

[0060] Non-deterministic operators and deterministic operators are classified from the perspective of determinism, and the classification criteria are whether the execution result of the operator is certain. The classification criteria of weakly aligned operators and strongly aligned operators are whether the execution result of the operator is precisely aligned.

[0061] Figure 3 A distributed training framework upgrade verification method 300 according to an embodiment of the present disclosure is shown, which is used for the process of upgrading an old version of the distributed training framework to a new version. The processing object of the distributed training framework includes at least any one of the following: text, image, audio or video data.

[0062] The method 300 includes: step 301, converting the non-deterministic operator in the distributed training framework into an equivalent deterministic operator; the non-deterministic operator is an operator whose result is uncertain when processing the same processing object in the old version and the new version, and the deterministic operator is an operator whose result is determined when processing the same processing object in the old version and the new version.

[0063] Non-deterministic operators: In the distributed training framework, non-deterministic operators refer to those operators whose output results after executing operations cannot be guaranteed to be consistent even when the old and new versions face exactly the same processing objects (such as the same images, audio clips, text data, etc.). This uncertainty may come from the random factors within the algorithm. In the implementation of the computing card, there may be several implementation algorithms for the same computing operation, and the results of some of the implementation algorithms are non-deterministic, which will cause differences in the results of the two calculations.

[0064] Non-deterministic operators can be converted to deterministic operators. In contrast to non-deterministic operators, deterministic operators can stably give the same output results when processing the same processing object, whether in the old version or the new version of the distributed training framework. In order to ensure that the upgrade verification of the distributed framework is not affected by non-deterministic operators, the non-deterministic operators in the framework can be converted to equivalent deterministic operators. For example, in the implementation of a computing card, there may be several implementation algorithms for the same computing operation, some of which have non-deterministic results and others are deterministic. In this case, the non-deterministic implementation algorithms can be converted to deterministic implementation algorithms for the same computing operation.

[0065] By converting non-deterministic operators into deterministic operators, the differences in calculation results between different versions of training frameworks caused by the uncertainty of the operators themselves are effectively eliminated. This is especially important in distributed training, because a large number of computing tasks are distributed across multiple computing nodes. If the results of some operators are unstable, the entire training process will be difficult to reproduce, and the performance and effects of different versions of the framework cannot be accurately compared. For example, in an image classification training task, if a non-deterministic operator causes the model parameters obtained in each training to be slightly different, the accuracy and generalization ability of the model will become unreliable. After the conversion, the same input data will always produce the same output, making the training results repeatable and comparable.

[0066] In the process of upgrading the distributed training framework from the old version to the new version, deterministic operator conversion is a key link to ensure the functional correctness and computational accuracy consistency of the upgraded framework. It provides a stable foundation for subsequent framework upgrade verification, avoids misjudgments caused by non-deterministic operators, and enables the verification results to truly reflect the improvements or changes in the functions and performance of the new version of the framework relative to the old version, which helps to improve the quality and credibility of the entire distributed training framework upgrade process.

[0067] The method 300 also includes: step 302, converting the weak alignment operator in the distributed training framework into an equivalent strong alignment operator; the weak alignment operator is an operator that produces inconsistent results when processing the same processing object in the old version and the new version, and the strong alignment operator is an operator that produces consistent results when processing the same processing object in the old version and the new version.

[0068] Step 303, determine the upgrade verification result of the distributed training framework based on the comparison result of the first hash value and the second hash value of each operator of the distributed training framework; wherein the first hash value is the hash value of the execution result of each operator under the old version, and the second hash value is the hash value of the execution result of each operator under the new version.

[0069] The order of step 301 and step 302 can be interchanged. Step 301 only needs to be performed before step 303.

[0070] Embodiment 4

[0071] Figure 4 A distributed training framework upgrade verification method 400 according to an embodiment of the present disclosure is shown. Figure 3 Step 301 in is expanded to obtain step 401.

[0072] Taking the case where the processing object contains text as an example, the distributed training framework upgrade verification method in this embodiment is used for the process of upgrading the distributed training framework from an old version to a new version.

[0073] In this embodiment, before converting the non-deterministic operator in the distributed training framework into an equivalent deterministic operator, it also includes: acquiring the non-deterministic operator in the distributed training framework according to a preset non-deterministic operator feature library.

[0074] That is, step 401 is to obtain the non-deterministic operator in the distributed training framework according to the preset non-deterministic operator feature library; convert the non-deterministic operator in the distributed training framework into an equivalent deterministic operator; the non-deterministic operator is an operator with uncertain results when processing the same processing object in the old version and the new version, and the deterministic operator is an operator with certain results when processing the same processing object in the old version and the new version.

[0075] The non-deterministic operators in the distributed training framework are obtained through the preset non-deterministic operator feature library. Taking the random number generation operator as an example of a non-deterministic operator, the random number sequence generated by the random number generation operator may be different in different versions of software or different operating environments. The operator features such as the operator name and operator type features of these random number generation operators are stored in the non-deterministic operator feature library so that the non-deterministic operators can be matched therefrom.

[0076] For the acquisition of general non-deterministic operator feature libraries, it is necessary to collect relevant information of each operator involved in the distributed training framework. This includes obtaining the basic properties of the operator, such as the operator name, input and output parameter format, and the computing module in which it is located, from different running instances and training task execution records under different versions of software. At the same time, the collected data is standardized to meet the unified format requirements of subsequent matching analysis, such as unified case specifications for operator names and standardized expressions of parameter formats, to ensure data consistency and accuracy, and to prepare for comparison with features in the non-deterministic operator feature library.

[0077] For the matching of non-deterministic operators, the operators to be matched are compared one by one with the corresponding features pre-stored in the non-deterministic operator feature library according to key dimensions such as operator name and operator type features. For example, for an operator to be detected, check whether its operator name matches the name of non-deterministic operators such as random number generation operators recorded in the feature library; then analyze other operator type features to determine whether it is a non-deterministic operator.

[0078] The method 400 also includes: step 402, converting the weak alignment operator in the distributed training framework into an equivalent strong alignment operator; the weak alignment operator is an operator that produces inconsistent results when processing the same processing object in the old version and the new version, and the strong alignment operator is an operator that produces consistent results when processing the same processing object in the old version and the new version.

[0079] Step 403, determine the upgrade verification result of the distributed training framework based on the comparison result of the first hash value and the second hash value of each operator of the distributed training framework; wherein the first hash value is the hash value of the execution result of each operator under the old version, and the second hash value is the hash value of the execution result of each operator under the new version.

[0080] The order of step 401 and step 402 can be interchanged. Step 401 only needs to be performed before step 403.

[0081] The technical effect of the above scheme is that the operators in the distributed training framework can be screened and matched through the preset non-deterministic operator feature library. By accurately identifying the non-deterministic operators in the distributed training framework, it is possible to clearly understand which operators have uncertain operating results, and then when analyzing whether the old framework and the new framework are equivalent, the impact of these non-deterministic operators can be avoided in a targeted manner.

[0082] Embodiment 5

[0083] The non-deterministic operators on the distributed training framework need to be mapped to deterministic operators to avoid the inconsistent accuracy caused by the non-deterministic algorithm. In addition, when running on different computing devices such as GPU, NPU, XPU, etc., due to their unique collaborative working methods, resource scheduling mechanisms, instruction sets and other characteristics, it is necessary to consider establishing corresponding deterministic operator mappers for different computing device types.

[0084] Figure 5 A distributed training framework upgrade verification method 500 according to an embodiment of the present disclosure is shown. Figure 3 Step 301 in is expanded to obtain step 501.

[0085] Taking the case where the processing object contains text as an example, the distributed training framework upgrade verification method in this embodiment is used for the process of upgrading the distributed training framework from an old version to a new version.

[0086] In this embodiment, the conversion of the non-deterministic operator in the distributed training framework into an equivalent deterministic operator includes: according to a preset deterministic operator mapper, obtaining a deterministic operator equivalent to the non-deterministic operator in the distributed training framework, and converting the non-deterministic operator into the deterministic operator.

[0087] The method 500 includes step 501, according to a preset deterministic operator mapper, obtaining a deterministic operator equivalent to a non-deterministic operator in the distributed training framework, and converting the non-deterministic operator into the deterministic operator; the non-deterministic operator is an operator whose result is uncertain when processing the same processing object in the old version and the new version, and the deterministic operator is an operator whose result is determined when processing the same processing object in the old version and the new version.

[0088] According to the preset deterministic operator mapper, a deterministic operator equivalent to the non-deterministic operator is obtained. Prior to this, corresponding deterministic operator mappers are established in advance for different computing device types, and the deterministic operator is an operator whose result is determined when processing the same processing object in the old version and the new version.

[0089] Deterministic operators, such as pseudo-random number generators, start the random number generation process with an initial seed value (seed). When the seed value is fixed, the generator will produce the same "random" number sequence according to a specific algorithm and steps. This is because each step of the algorithm is deterministic, but the generated number sequence looks random. The operator features of these pseudo-random number generation operators, such as operator name, operator type, seed-related features, etc., are stored in the deterministic operator feature library so that deterministic operators can be matched from it.

[0090] The method 500 also includes: step 502, converting the weak alignment operator in the distributed training framework into an equivalent strong alignment operator; the weak alignment operator is an operator that produces inconsistent results when processing the same processing object in the old version and the new version, and the strong alignment operator is an operator that produces consistent results when processing the same processing object in the old version and the new version.

[0091] Step 503, determine the upgrade verification result of the distributed training framework based on the comparison result of the first hash value and the second hash value of each operator of the distributed training framework; wherein the first hash value is the hash value of the execution result of each operator under the old version, and the second hash value is the hash value of the execution result of each operator under the new version.

[0092] The technical effect of the above scheme is: by finding an equivalent deterministic operator based on a preset deterministic operator mapper, for example, by utilizing the characteristic that a pseudo-random number generator can generate the same "random" number sequence based on a fixed seed value, it is possible to obtain a certain result when processing the same object in both the old version and the new version, thereby greatly improving the consistency of the training results and facilitating repeated experiments, comparative analysis, and subsequent model verification.

[0093] In this embodiment, before converting the non-deterministic operators in the distributed training framework into equivalent deterministic operators, it also includes: establishing corresponding deterministic operator mappers for different computing device types.

[0094] GPU has large-scale parallel computing capabilities, has many CUDA cores, and is good at processing large-scale data parallel tasks, such as matrix operations in deep learning. Its memory architecture and data transmission method are significantly different from traditional CPUs. For example, there are dedicated video memory and cache to accelerate data reading. NPU is a chip designed specifically for neural network computing. It has unique hardware acceleration units for specific neural network operations such as convolution and fully connected layer calculations. Its computing accuracy and data processing flow are optimized for the characteristics of neural networks. As a comprehensive computing device, XPU integrates multiple computing units (may include CPU, GPU, FPGA, etc.), its internal architecture is more complex, and the collaborative working method and resource scheduling mechanism between each computing unit are unique. Due to the essential differences in these hardware architectures, the execution efficiency, resource consumption and result determinism of the same operator on different devices may vary greatly. Therefore, it is necessary to establish corresponding deterministic operator mappers to adapt to the characteristics of different devices to ensure that stable and efficient computing results can be obtained on different hardware platforms.

[0095] Different computing devices such as GPU, NPU, and XPU have different instruction sets, so it is necessary to establish corresponding deterministic operator mappers for different computing device types. GPU and NPU have their own neural network instruction sets. XPU may support multiple instruction sets because it integrates multiple computing units, and the device type needs to be considered when mapping operators.

[0096] Embodiment 6

[0097] This embodiment Figure 1 Step 101 in the above is extended to obtain the weak alignment operator.

[0098] In this embodiment, before converting the weak alignment operator in the distributed training framework into an equivalent strong alignment operator, it also includes: setting a hook function in the distributed training framework, extracting operator features of each operator of the distributed training framework based on the hook function, the operator features including code features and / or communication transmission mode features; identifying weak alignment operators among the operators based on the operator features.

[0099] In the distributed training framework, different types of communication operators have different code logics. For example, for the common AllReduce operator, its code involves the logic of data aggregation and distribution between multiple nodes. Through the hook function, these key code snippets are accurately located to extract code features such as data block size processing logic, data merging and splitting code patterns, etc.

[0100] The communication transmission mode characteristics cover aspects such as the data transmission path, transmission protocol, transmission rate limit, and data caching mechanism during transmission. Taking the distributed training framework based on MPI (Message Passing Interface) as an example, some communication operators may use specific MPI communication functions, which determine how data is transmitted between different processes, whether it is blocking transmission or non-blocking transmission, and whether there is a data verification mechanism during the transmission process. The hook function needs to go deep into the calling level of these communication functions to extract these key transmission mode feature information, so as to comprehensively characterize the characteristics of the communication operator.

[0101] In the process of finding weak alignment operators based on the preset weak alignment operator feature library, the feature library needs to be carefully constructed and maintained. The weak alignment operator feature library should not only contain the typical features of common known weak alignment operators, such as the data processing difference features of a communication operator under a specific version due to protocol updates, but also have a dynamic update mechanism. When the distributed training framework is upgraded or a new communication library is introduced, the newly emerged weak alignment operator features can be incorporated in a timely manner.

[0102] By setting a hook function to extract the communication operator, the communication operator in the distributed training framework can be accurately identified, and then the weak alignment operator can be accurately screened out and converted into a strong alignment operator based on the weak alignment operator feature library. This process effectively solves the upgrade problem that may be caused by inconsistent results when the operator processes the same image data in the old and new versions, ensures the stability and consistency of the framework for image data processing between different versions, and improves the success rate and reliability of the upgrade.

[0103] Embodiment 7

[0104] In the distributed training framework, communication operators such as Reduce, AllGather, and ReduceScatter all involve tensor accumulation operations, which may lead to inconsistent precision due to different underlying implementation methods of the underlying communication library of tensor communication. As a result, the upgrade verification of the distributed training framework cannot be successful. Communication operators such as Reduce, AllGather, and ReduceScatter are weakly aligned operators. The AllReduce operator can achieve precision alignment and is a strong alignment operator. Therefore, in the distributed training framework, the conversion of weakly aligned operators such as Reduce, AllGather, and ReduceScatter is particularly important.

[0105] Figure 6 A distributed training framework upgrade verification method 600 of an embodiment of the present disclosure is shown. This embodiment expands step 101 to obtain step 601.

[0106] In this embodiment, the conversion of the weak alignment operator in the distributed training framework into an equivalent strong alignment operator includes: according to a preset communication mapper, obtaining a strong alignment operator equivalent to the weak alignment operator in the distributed training framework, and converting the weak alignment operator into the strong alignment operator.

[0107] The method 600 includes step 601, according to a preset communication mapper, obtaining a strong alignment operator equivalent to a weak alignment operator in the distributed training framework, and converting the weak alignment operator into the strong alignment operator; the weak alignment operator is an operator that produces inconsistent results when processing the same processing object in the old version and the new version, and the strong alignment operator is an operator that produces consistent results when processing the same processing object in the old version and the new version.

[0108] In this embodiment, the weak alignment operator includes any one of the following: Reduce operator, ReduceScatter operator, AllGather operator; the strong alignment operator includes AllReduce operator.

[0109] In this embodiment, in response to the weak alignment operator being a Reduce operator, converting the weak alignment operator into an equivalent strong alignment operator is mapping the Reduce operator to an AllReduce operator; in response to the weak alignment operator being a ReduceScatter operator, converting the weak alignment operator into an equivalent strong alignment operator is mapping the ReduceScatter operator to an AllReduce operator; in response to the weak alignment operator being an AllGather operator, converting the weak alignment operator into an equivalent strong alignment operator is mapping the AllGather operator to an AllReduce operator.

[0110] In some cases, it is difficult to achieve precision alignment for the Reduce operator, ReduceScatter operator, and AllGather operator, but the AllReduce operator has the feature of precision alignment.

[0111] The working principle of the Reduce operator is to perform specified reduction operations (such as summation, product, etc.) on the data elements at the same position on each process in a group of processes, and save the results in one of the processes. However, since there may be slight differences in the data on different processes, for example, when the data representation precision is limited, rounding errors will accumulate during the reduction process. This cumulative effect is particularly obvious when the data involved in the reduction is unevenly distributed. In addition, the Reduce operator only saves the final result in one process, and other processes cannot obtain complete calculation process information. This makes it difficult to ensure that the results of each execution are completely consistent in different versions or different operating environments, and it is impossible to achieve precision alignment.

[0112] The ReduceScatter operator first performs local reduction on the data of each process, and then distributes the reduced results to each process. In the local reduction stage, the problem of rounding error accumulation will also be faced, because the local calculation of each process depends on its own data accuracy and calculation order. In the process of dispersing the results, due to the differences in the order of data transmission and distribution and the storage method of data in different processes, these factors will make it difficult to accurately reproduce the final results in different operating environments, making it impossible to achieve precision alignment.

[0113] The AllGather operator is used to collect data from each process to all processes, so that each process has a copy of the data of all processes. Although its purpose is data collection, during the data transmission and combination process, the differences in data representation accuracy between different processes and slight changes in the order and synchronization mechanism of data transmission will affect the consistency of the final collected data in each process. For example, during the transmission process, if the data of a process changes slightly due to network delays or differences in the version of the transmission protocol, then when the complete data set is combined, the results obtained by different processes will be biased, and precision alignment cannot be guaranteed.

[0114] In contrast, the AllReduce operator can achieve precision alignment. The AllReduce operator performs reduction operations simultaneously in the entire distributed process group and broadcasts the final reduction results to all processes. During the reduction process, since all processes participate in the same calculation steps and the calculation results are synchronized and broadcasted among all processes, any slight error changes will be uniformly processed and propagated throughout the process group. For example, when processing floating-point calculations, even if there are rounding errors, since all processes perform calculations and propagate results according to the same rules and order, the results obtained by each process are consistent in the end, thus achieving precision alignment. This feature gives the AllReduce operator a clear advantage in distributed training scenarios that require high precision and result consistency, and can effectively avoid fluctuations and deviations in training results caused by operator precision issues.

[0115] The technical effect of the above scheme is that when the weak alignment operator is converted into an equivalent strong alignment operator according to the preset communication mapper, the Reduce operator, ReduceScatter operator, and AllGather operator are respectively mapped to the AllReduce operator, utilizing the precision alignment feature of the AllReduce operator. Through this mapping conversion, it is possible to ensure that the calculation results in different versions and different operating environments have higher consistency and accuracy in the distributed training framework, thereby improving the stability and reliability of the entire training process.

[0116] Embodiment 8

[0117] The weak alignment operator on the distributed training framework needs to be mapped to a strong alignment operator to avoid the inconsistent accuracy caused by the different communication modes of the underlying communication library. In addition, when running on different computing devices such as GPU, NPU, XPU, etc., due to their unique collaborative working methods, resource scheduling mechanisms, instruction sets and other characteristics, it is necessary to consider establishing corresponding communication mappers for different computing device types.

[0118] In this embodiment, before converting the weak alignment operator in the distributed training framework into an equivalent strong alignment operator, it also includes: establishing corresponding communication mappers for different computing device types.

[0119] As described in the previous embodiments, different computing devices such as GPU, NPU, and XPU have different instruction sets, so corresponding communication mappers need to be established for different types of computing devices.

[0120] In this embodiment, corresponding communication mappers are established for different computing device types, including: setting a target numerical accuracy standard according to the accuracy requirements of the computing device; based on the target numerical accuracy standard, converting a weak alignment operator into a strong alignment operator with numerical accuracy alignment, and establishing a communication mapper based on the correspondence between the weak alignment operator and the strong alignment operator.

[0121] In this embodiment, after converting the weak alignment operator in the distributed training framework into an equivalent strong alignment operator, it also includes: optimizing and adjusting the parameters of the strong alignment operator during the mapping process to adapt to the current computing device and communication mode.

[0122] Suppose that in a distributed framework, after analyzing the model training data type and computing device (such as a specific model of GPU), it is determined to use single-precision floating point numbers (32 bits) as the target numerical accuracy standard to ensure the consistency of calculation accuracy throughout the training process and meet the model's accuracy requirements.

[0123] In the gradient calculation phase of a model parameter update, the old distributed training framework uses the Reduce operator to reduce the gradients calculated on each computing node. Due to factors such as the initial representation accuracy of data on different nodes and rounding errors in the calculation process, the results obtained by the Reduce operator may be slightly different in different training rounds or different versions of the software environment, affecting the stability and reproducibility of model training.

[0124] Now convert the Reduce operator to the AllReduce operator according to the set single-precision floating-point precision standard. The AllReduce operator will perform gradient reduction operations on all computing nodes simultaneously, and each node will participate in the entire reduction calculation process. For example, for the weight gradient of a certain layer in the model, all nodes first organize and prepare their local gradient data according to the format and precision requirements of single-precision floating-point numbers, and then perform sum reduction operations at the same time. In this process, the calculation error of any node will be synchronously propagated and processed in the entire node group, instead of being reflected only in the node where the final result is located like the Reduce operator. In the end, each node can obtain exactly the same reduced gradient result with aligned precision, which can be used for subsequent unified model parameter update operations.

[0125] Through such a conversion, the Reduce operator, which may have inconsistent results due to precision issues, is converted to the AllReduce operator that can achieve numerical precision alignment, adapting to the current computing devices and communication modes, and building a communication mapper based on this. In subsequent distributed training, when similar gradient reduction operations are required, the AllReduce operator can be used directly to replace the original Reduce operator based on the communication mapper, thereby improving the consistency and stability of the results of the entire distributed training process under different computing devices and different versions, and ensuring the accuracy and repeatability of model training.

[0126] Embodiment 9

[0127] In this embodiment, before determining the upgrade verification result of the distributed training framework based on the comparison result of the first hash value and the second hash value of each operator of the distributed training framework, it also includes: obtaining each operator according to the logical order of each layer of the distributed training framework; the logical order is from the input layer to the middle layer and then to the output layer.

[0128] In this embodiment, the step of determining the upgrade verification result of the distributed training framework according to the comparison result of the first hash value and the second hash value of each operator of the distributed training framework includes:

[0129] Compare the first hash value and the second hash value of each operator in sequence according to the logical order of each layer of the distributed training framework to obtain a comparison result; the logical order is from the input layer to the middle layer and then to the output layer;

[0130] The upgrade verification result of the distributed training framework is determined according to the comparison results of the various operators.

[0131] The operators in the above distributed training framework refer to all operators involved in the framework operation, including but not limited to the aforementioned non-deterministic operators, deterministic operators, weak alignment operators, and strong alignment operators. In addition, other operators outside the above categories are also included.

[0132] Taking the case where the processing object includes video as an example, the distributed training framework upgrade verification method in this embodiment is applied during the upgrade of the distributed training framework from an old version to a new version.

[0133] Obtain each operator in the logical order of the distributed training framework from the input layer to the middle layer and then to the output layer. First obtain the first hash value of the operator's execution result under the old version, and then obtain the second hash value of its execution result under the new version. When obtaining the comparison result of the first hash value and the second hash value, the comparison is also carried out in sequence according to the logical order from the input layer to the middle layer and then to the output layer, and finally determine the upgrade verification result of the distributed training framework based on the comparison result.

[0134] Operators are obtained and hash values ​​are compared in a logical order from the input layer to the middle layer and then to the output layer, establishing a clear process framework for the entire upgrade verification process. When processing complex video data, since operators at each layer are interrelated and executed sequentially, sequential operations can avoid omissions or mismatches caused by a chaotic verification order, thereby significantly improving the accuracy of the verification results. For example, operators at different levels such as the feature extraction layer and encoding layer of the video can be accurately verified to ensure that each link can run correctly after the version upgrade and maintain the quality and stability of video processing. In addition, operator processing and verification are performed using layer-by-layer logic. When the upgrade verification fails, the level where the problem is located can be quickly located.

[0135] Embodiment 10

[0136] Figure 7 A distributed training framework upgrade verification method 700 according to an embodiment of the present disclosure is shown. Figure 1 Step 102 is expanded to obtain step 702.

[0137] Step 701 of the method 700 converts the weak alignment operator in the distributed training framework into an equivalent strong alignment operator; the weak alignment operator is an operator that produces inconsistent results when processing the same processing object in the old version and the new version, and the strong alignment operator is an operator that produces consistent results when processing the same processing object in the old version and the new version.

[0138] Step 702 of the method 700 is to obtain a first hash value of each operator of the distributed training framework; obtain a second hash value of each operator of the distributed training framework; wherein the first hash value is a hash value of an execution result of each operator under the old version, and the second hash value is a hash value of an execution result of each operator under the new version; compare the first hash value and the second hash value of each operator to obtain a comparison result of each operator; in response to the comparison results of all the operators being the same, determine that an upgrade verification result of the distributed training framework is successful; in response to the comparison results of at least one of the operators being different, determine that an upgrade verification result of the distributed training framework is unsuccessful, and output the upgrade unsuccessful operator whose comparison result is different.

[0139] Assume that there is a distributed training framework for training a simple handwritten digit recognition model. The following example illustrates the upgrade verification process in detail.

[0140] (I) The distributed training framework has three main operators in the old version (assuming version 1.0): data loading operator, distributed convolution operator, and full connection operator. Among the three operators in the old version, the distributed convolution operator involves the Reduce operator of tensor communication, which has precision misalignment and is a weakly aligned operator.

[0141] 1. Data loading operator

[0142] In the old version, the training data set is a data set containing 60,000 handwritten digital images. The data loading operator reads the image data and the corresponding digital labels from the local storage in batches (batch size is set to 100). After traversing the entire data set once, all the read image data and label information are integrated, and the hash value is calculated using the MD5 hash algorithm. The first hash value obtained is assumed to be "Hash1_DataLoad_1.0".

[0143] 2. Distributed Convolution Operator

[0144] After the data loading operator passes each batch of data to the distributed convolution operator, in the old version, the operator uses distributed computing to perform convolution operations. In a cluster environment consisting of 5 computing nodes, the convolution operator is configured with 3 convolution layers, with convolution kernel sizes of 3×3, 5×5, and 3×3, respectively, with a step size of 1, and the padding method is the same padding. Each computing node is responsible for processing the convolution calculation task of a part of the input data, and then the local feature maps calculated by each node are aggregated and summarized through the Reduce operator. However, in this old version, the Reduce operator may have precision misalignment, such as slight differences in the data type or precision setting after the decimal point of different computing nodes, which may cause slight errors during aggregation. After processing the batch data of the entire data set multiple times, all the generated feature map data are aggregated, and the hash value is calculated using the MD5 hash algorithm. This first hash value is recorded as "Hash1_Convolution_1.0".

[0145] 3. Fully connected operator

[0146] Based on the feature graph output by the distributed convolution operator, the fully connected operator builds two fully connected layers in the old version, with 128 and 10 neurons respectively (corresponding to 10 digital categories), and converts the features into the final predicted probability distribution through corresponding calculations. After processing all batches of data, all prediction results are collected, and the first hash value calculated by the MD5 hash algorithm is "Hash1_FullyConnected_1.0".

[0147] (II) Now we need to upgrade the distributed training framework to a new version (assuming version 2.0), and the functions and implementation details of each operator have changed. Among the three operators in the new version, the distributed convolution operator involves the AllReduce operator of tensor communication, and there is no precision misalignment, so it is a strong alignment operator.

[0148] 1. Data loading operator

[0149] In the new version, although it is still for the 60,000 handwritten digital image dataset, and the batch size is still 100, the underlying data reading has optimized the cache mechanism, and the reading speed has been improved. Similarly, after reading the dataset completely, the data is integrated and the hash value is calculated using the MD5 hash algorithm. The second hash value is assumed to be "Hash2_DataLoad_2.0".

[0150] 2. Distributed Convolution Operator

[0151] The new version of the distributed convolution operator is adjusted to 4 convolution layers, with convolution kernel sizes of 3×3, 3×3, 5×5, and 3×3, respectively, and step sizes of 1, 2, 1, and 1, respectively. The filling method has also changed accordingly, and the task allocation strategy of distributed computing has been optimized. The number of computing nodes involved in the calculation has increased to 8, aiming to extract richer and more accurate features. After each computing node performs convolution calculations according to the new rules, the local feature maps are aggregated and summarized through the precision-aligned AllReduce operator (the operator converted from the previous strong alignment operator is used, and the new version is further adapted to the new distributed computing scenario). After the entire data set is processed in batches, the feature map data is summarized, and the second hash value obtained by the MD5 hash algorithm is recorded as "Hash2_Convolution_2.0".

[0152] 3. Fully connected operator

[0153] In the new version of the fully connected operator, the number of neurons in the fully connected layer is adjusted to 256 and 10, and a new activation function is used to improve the model's expressiveness and prediction accuracy. After processing all batches of data, the prediction results are collected and the second hash value calculated using the MD5 hash algorithm is "Hash2_FullyConnected_2.0".

[0154] (III) Strong alignment operator conversion steps

[0155] Before performing the upgrade verification, the key step of strong alignment operator conversion must be performed. For the Reduce operators with unaligned precision in the distributed convolution operator in the old version, all of them are replaced with the precision-aligned AllReduce operator. The AllReduce operator can ensure that the data precision of each node remains consistent when data is aggregated on multiple computing nodes, avoiding result deviations caused by precision differences. For example, the data types participating in the aggregation of each computing node will be unified as high-precision floating point types (such as 64-bit floating point numbers), and the precision-related settings such as the effective digits after the decimal point will be standardized. After such a conversion, the distributed convolution operator can perform more accurate feature map aggregation and summary operations based on the precision-aligned AllReduce operator in subsequent execution, providing a more reliable data basis for subsequent upgrade verification. After its hash value is recalculated, the first hash value obtained by the MD5 hash algorithm is recorded as "Hash2_Convolution_1.0A".

[0156] (IV) Compare hash values ​​and determine upgrade verification results

[0157] 1. Compare the hash values ​​of the data loading operators

[0158] Compare "Hash1_DataLoad_1.0" and "Hash2_DataLoad_2.0". If the two are consistent, it means that although the internal implementation of the data loading operator is optimized before and after the upgrade, the actual content of the data loaded is unchanged, and the comparison result is the same; if the two are inconsistent, the comparison result is different.

[0159] 2. Comparing the hash values ​​of distributed convolution operators

[0160] Comparing "Hash2_Convolution_1.0A" and "Hash2_Convolution_2.0", if the hash values ​​of the two are exactly the same, it means that although the structure and parameters of the distributed convolution operator have been adjusted, they are equivalent in terms of the overall execution results, and the comparison results are the same; if the two are different, the comparison results are different.

[0161] 3. Compare the hash values ​​of the fully connected operators

[0162] Similarly, compare "Hash1_FullyConnected_1.0" with "Hash2_FullyConnected_2.0". If they are the same, the comparison results of the fully connected operators are the same, otherwise they are different.

[0163] After comparison, it is found that the hash values ​​of the distributed convolution operator and the fully connected operator are the same, but the hash values ​​of the data loading operator are different. Since at least one operator has a different comparison result, the upgrade verification result of the distributed training framework is determined to be unsuccessful, and the data loading operator with a different comparison result is output as the upgrade unsuccessful operator, which is convenient for developers to further investigate the reasons for the difference in the operator in the new version, and then adjust and optimize the upgrade process.

[0164] The upgrade verification of the entire distributed training framework can only be determined to be successful if the upgrade verification results of all operators in the distributed training framework are successful. As long as it is found that the upgrade verification result of at least one operator is unsuccessful, the upgrade verification result of the entire distributed training framework will be judged as a failure. In addition, the system will automatically output the information of those operators whose upgrade verification results are unsuccessful. This information will provide a critical basis for subsequent in-depth investigation of the root cause of the problem and targeted optimization and improvement work.

[0165] This embodiment records in detail the hash values ​​of the execution results of each operator in the old version and the new version, and performs comparative analysis. Once the upgrade verification fails, the specific operator that failed to upgrade can be quickly and accurately located. This provides developers with a clear debugging direction and avoids the dilemma of blindly troubleshooting problems in large-scale distributed training frameworks. By analyzing and optimizing these problem operators in a targeted manner, compatibility or functional problems that arise during the upgrade process can be quickly resolved, effectively shortening the framework upgrade cycle, reducing the time and resource costs consumed by upgrade debugging, and improving the maintenance efficiency of the framework.

[0166] The beneficial effect of this embodiment is that by obtaining the hash value of the execution results of each operator in the old version and the new version and comparing them, it is possible to accurately determine whether the output of each operator has changed before and after the upgrade, thereby ensuring functional consistency. If the comparison results of all operators are the same, it means that from a functional perspective, the entire distributed training framework can still maintain consistent behavior with the old version after the upgrade. This can effectively avoid some potential functional anomalies caused by the upgrade, and ensure that the framework can continue to perform training tasks stably as expected after the upgrade. For example, when training the model, the data processing flow, feature extraction, and final prediction generation will not deviate due to the upgrade, providing users with a stable and reliable training environment.

[0167] In addition, the problematic operator can be quickly located. When the comparison results of at least one operator are different, the upgrade verification result is determined to be unsuccessful, and the unsuccessful upgrade operator with different comparison results will be output. This design eliminates the need for developers or operation and maintenance personnel to conduct a comprehensive and complicated investigation of the entire distributed training framework when problems arise during the upgrade. You can focus directly on these specific, different operators, greatly narrowing the scope of troubleshooting, saving time and effort, and helping to quickly locate which operators were affected during the upgrade process, thereby more efficiently analyzing the causes, making timely adjustments to the upgrade plan or repairing problems with related operators, and promoting the smooth completion of the upgrade work.

[0168] Embodiment 11

[0169] In this embodiment, after obtaining the first hash value of each operator of the distributed training framework, it also includes: storing the first hash value in a first log file;

[0170] After obtaining the second hash value of each operator of the distributed training framework, the method further includes: storing the second hash value in a second log file;

[0171] Before comparing the first Hash value and the second Hash value of each operator, the method further includes: obtaining the first Hash value in the first log file; and obtaining the second Hash value in the second log file.

[0172] The operators in the above distributed training framework refer to all operators involved in the framework operation, including but not limited to the aforementioned non-deterministic operators, deterministic operators, weak alignment operators, and strong alignment operators. In addition, other operators outside the above categories are also included.

[0173] Take image data in image classification tasks as an example, which is a common processing object in this field.

[0174] For any operator in the distributed training framework, we must first determine the corresponding image data as its processing object. In the old version of the distributed training framework environment, start the operator to process the image data. When the operator completes the processing of the image data, it will produce a specific execution result. We use the hash algorithm to process this execution result, extract its hash value, define this hash value as the first hash value, and store its record in a special first log file.

[0175] Subsequently, the same image data is obtained again as the processing object of the operator, and this time the operator is run in the new version environment of the distributed training framework. After the processing is completed, a new execution result is obtained, and a hash value is extracted from the new execution result, set as the second hash value, and then the second hash value is saved in the second log file.

[0176] When determining the final result of the upgrade verification of the distributed training framework, the following operations must be performed for each operator in the framework: read the corresponding first hash value from the first log file, then read the corresponding second hash value from the second log file, and then conduct a detailed comparison analysis of the two hash values ​​to obtain the comparison result. If the comparison result shows that the two hash values ​​are completely consistent, then the upgrade verification of this operator has passed successfully, that is, the upgrade verification result is successful; conversely, if the comparison result shows that there is a difference between the two hash values, then the upgrade verification of the operator has failed, that is, the upgrade verification result is unsuccessful.

[0177] The beneficial effect of this embodiment is to ensure the accuracy and integrity of the data. The log file has a certain storage stability and persistence. Once the hash value is stored, as long as the file storage system operates normally, the data can be completely preserved. Compared with temporary records or simply storing data in memory, log files can better ensure that hash value data is not lost or tampered with, and provide an accurate and reliable data basis for the comparison of hash values ​​of each operator during the upgrade verification process, ensuring the authenticity of the verification results and avoiding deviations in verification due to data loss or errors.

[0178] Embodiment 12

[0179] like Figure 8 As shown, the distributed training framework upgrade verification device 800 provided in this embodiment includes: a strong alignment operator conversion unit 801 and a framework upgrade verification unit 802.

[0180] Among them, the strong alignment operator conversion unit 801 is configured to convert the weak alignment operator in the distributed training framework into an equivalent strong alignment operator; the weak alignment operator is an operator whose results are inconsistent when processing the same processing object in the old version and the new version, and the strong alignment operator is an operator whose results are consistent when processing the same processing object in the old version and the new version;

[0181] The framework upgrade verification unit 802 is configured to determine the upgrade verification result of the distributed training framework based on the comparison result of the first hash value and the second hash value of each operator of the distributed training framework; wherein the first hash value is the hash value of the execution result of each operator under the old version, and the second hash value is the hash value of the execution result of each operator under the new version.

[0182] In this embodiment, the specific processing of each unit in the distributed training framework upgrade verification device 800 and the technical effects it brings can be referred to the relevant descriptions of each step in the aforementioned embodiment, and will not be repeated here.

[0183] Embodiment 13

[0184] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.

[0185] Fig. 9 A schematic block diagram of an example electronic device 900 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.

[0186] like Fig. 9 As shown, the device 900 includes a computing unit 901, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the device 900 can also be stored. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0187] A number of components in the device 900 are connected to the I / O interface 905, including: an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a disk, an optical disk, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the device 900 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0188] The computing unit 901 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 901 performs the various methods and processes described above, such as a distributed training framework upgrade verification method. For example, in some embodiments, the distributed training framework upgrade verification method may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 908. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the computing unit 901, one or more steps of the distributed training framework upgrade verification method described above may be performed. Alternatively, in other embodiments, the computing unit 901 may be configured to perform the distributed training framework upgrade verification method in any other appropriate manner (e.g., by means of firmware).

[0189] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0190] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable information processing device, so that the program code, when executed by the processor or controller, implements the functions / operations specified in the flow chart and / or block diagram. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.

[0191] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0192] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0193] The systems and techniques described herein may be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an information server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communications network). Examples of communications networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0194] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.

[0195] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.

[0196] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A distributed training framework upgrade verification method, used in the process of upgrading an old version of the distributed training framework to a new version, characterized in that: The processing object of the distributed training framework includes at least any one of the following: text, image, audio or video data; the method includes: Converting a weak alignment operator in the distributed training framework into an equivalent strong alignment operator; the weak alignment operator is an operator that produces inconsistent results when processing the same processing object in the old version and the new version, and the strong alignment operator is an operator that produces consistent results when processing the same processing object in the old version and the new version; The upgrade verification result of the distributed training framework is determined based on the comparison result of the first hash value and the second hash value of each operator of the distributed training framework; wherein the first hash value is the hash value of the execution result of each operator under the old version, and the second hash value is the hash value of the execution result of each operator under the new version.

2. The method according to claim 1, characterized in that Before determining the upgrade verification result of the distributed training framework according to the comparison result of the first hash value and the second hash value of each operator of the distributed training framework, the method further includes: The non-deterministic operators in the distributed training framework are converted into equivalent deterministic operators; the non-deterministic operators are operators whose results are uncertain when processing the same processing object in the old version and the new version, and the deterministic operators are operators whose results are determined when processing the same processing object in the old version and the new version.

3. The method according to claim 2, characterized in that Before converting the non-deterministic operators in the distributed training framework into equivalent deterministic operators, it also includes: acquiring the non-deterministic operators in the distributed training framework according to a preset non-deterministic operator feature library.

4. The method according to claim 2, characterized in that: The converting of the non-deterministic operator in the distributed training framework into an equivalent deterministic operator includes: obtaining a deterministic operator equivalent to the non-deterministic operator in the distributed training framework according to a preset deterministic operator mapper, and converting the non-deterministic operator into the deterministic operator.

5. The method according to claim 2, characterized in that: Before converting the non-deterministic operators in the distributed training framework into equivalent deterministic operators, it also includes: establishing corresponding deterministic operator mappers for different computing device types.

6. The method according to claim 1, characterized in that Before converting the weak alignment operator in the distributed training framework into an equivalent strong alignment operator, it also includes: setting a hook function in the distributed training framework, extracting operator features of each operator of the distributed training framework based on the hook function, and the operator features include code features and / or communication transmission mode features; identifying weak alignment operators among the operators based on the operator features.

7. The method according to claim 1, characterized in that The converting of the weak alignment operator in the distributed training framework into an equivalent strong alignment operator includes: obtaining a strong alignment operator equivalent to the weak alignment operator in the distributed training framework according to a preset communication mapper, and converting the weak alignment operator into the strong alignment operator.

8. The method according to claim 7, characterized in that The weak alignment operator includes any one of the following: Reduce operator, ReduceScatter, AllGather operator; the strong alignment operator includes AllReduce operator.

9. The method according to claim 7, characterized in that: In response to the weak alignment operator being a Reduce operator, converting the weak alignment operator into an equivalent strong alignment operator is mapping the Reduce operator into an AllReduce operator; In response to the weak alignment operator being a ReduceScatter operator, converting the weak alignment operator into an equivalent strong alignment operator is mapping the ReduceScatter operator into an AllReduce operator; In response to the weak alignment operator being an AllGather operator, converting the weak alignment operator into an equivalent strong alignment operator is mapping the AllGather operator into an AllReduce operator.

10. The method according to claim 1, characterized in that Before converting the weak alignment operator in the distributed training framework into an equivalent strong alignment operator, the method further includes: establishing corresponding communication mappers for different computing device types.

11. The method according to claim 10, characterized in that The corresponding communication mappers are established for different computing device types, including: A target numerical precision standard is set according to the precision requirement of the computing device; based on the target numerical precision standard, a weak alignment operator is converted into a strong alignment operator with numerical precision alignment, and a communication mapper is established according to the corresponding relationship between the weak alignment operator and the strong alignment operator.

12. The method according to claim 10, characterized in that After converting the weak alignment operator in the distributed training framework into an equivalent strong alignment operator, it also includes: optimizing and adjusting the parameters of the strong alignment operator during the mapping process to adapt to the current computing device and communication mode.

13. The method according to claim 1, characterized in that Before determining the upgrade verification result of the distributed training framework according to the comparison result of the first hash value and the second hash value of each operator of the distributed training framework, the method further includes: Each operator is obtained according to the logical order of each layer of the distributed training framework; the logical order is from the input layer to the middle layer and then to the output layer.

14. The method according to claim 1, characterized in that The step of determining the upgrade verification result of the distributed training framework according to the comparison result of the first hash value and the second hash value of each operator of the distributed training framework comprises: Compare the first hash value and the second hash value of each operator in sequence according to the logical order of each layer of the distributed training framework to obtain a comparison result; the logical order is from the input layer to the middle layer and then to the output layer; The upgrade verification result of the distributed training framework is determined according to the comparison results of the various operators.

15. The method according to claim 1, characterized in that The step of determining the upgrade verification result of the distributed training framework according to the comparison result of the first hash value and the second hash value of each operator of the distributed training framework comprises: Obtaining a first hash value of each operator of the distributed training framework; obtaining a second hash value of each operator of the distributed training framework; comparing the first hash value and the second hash value of each operator to obtain a comparison result of each operator; In response to the comparison results of all the operators being the same, it is determined that the upgrade verification result of the distributed training framework is successful; in response to the comparison results of at least one of the operators being different, it is determined that the upgrade verification result of the distributed training framework is unsuccessful, and the upgrade unsuccessful operators with different comparison results are output.

16. The method according to claim 15, characterized in that The obtaining of the first hash value of each operator of the distributed training framework includes: Obtaining a processing object of each operator; processing the processing object based on each operator under the old version to obtain an execution result of each operator; extracting a hash value of the execution result as a first hash value; The obtaining of the second hash value of each operator of the distributed training framework includes: Acquire a processing object of each operator; process the processing object based on each operator under the new version to obtain an execution result of each operator; and extract a hash value of the execution result as a second hash value.

17. The method according to claim 15, characterized in that After obtaining the first hash value of each operator of the distributed training framework, the method further includes: Storing the first hash value in a first log file; After obtaining the second hash value of each operator of the distributed training framework, the method further includes: storing the second hash value in a second log file; Before comparing the first Hash value and the second Hash value of each operator, the method further includes: Obtain a first hash value in the first log file; Obtain a second hash value in the second log file.

18. A distributed training framework upgrade verification device, used in the process of upgrading an old version of the distributed training framework to a new version, characterized in that: The processing object of the distributed training framework includes at least any one of the following: text, image, audio or video data; the device includes: A strong alignment operator conversion unit is configured to convert a weak alignment operator in the distributed training framework into an equivalent strong alignment operator; the weak alignment operator is an operator that produces inconsistent results when processing the same processing object in the old version and the new version, and the strong alignment operator is an operator that produces consistent results when processing the same processing object in the old version and the new version; The framework upgrade verification unit is configured to determine the upgrade verification result of the distributed training framework based on the comparison result of the first hash value and the second hash value of each operator of the distributed training framework; wherein the first hash value is the hash value of the execution result of each operator under the old version, and the second hash value is the hash value of the execution result of each operator under the new version.

19. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 17.

20. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 17.

21. A computer program product, comprising a computer program, which, when executed by a processor, implements the method of any one of claims 1 to 17.