Method, computing device, medium, and program product for generating target operator code
Patent Information
- Application Number
- CN202611260233.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-19
- Publication Date
- 2026-09-25
AI Technical Summary
[0003]综上,传统的算子代码生成的方法存在的不足之处在于:耗时较长、效率较低、难以支持非标准算子代码的生成,以及算子代码生成过程中缺乏系统化的验证
[0022]本发明根据获取算子的参考实现代码和算法摘要文档,自动地生成了结构化规范文档,并基于该结构化规范文档自动生成在目标处理器上运行的目标算子代码,从而避免了人工编写的高耗时和低效率,实现了从已有代码和文档到目标代码的自动转换,大幅提高了算子代码的生成效率。此外,由于规范文档是基于参考实现代码与摘要文档生成的,而非依赖于固定的内部模版,因此能够灵活适配于各种非标准算子的生成。最后,在生成目标算子代码后,基于参考实现代码对目标算子代码执行正确性验证,并根据验证结果输出最终代码,解决了算子代码生成过程中缺乏系统化的验证的问题。因此,本发明不仅能够显著提高算子代码的生成效率,而且适用于非标准算子代码的生成,并且在算子代码生成过程中提供了系统化的验证。
Smart Images

Figure CN122816645A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention generally relate to the field of artificial intelligence technology, and more specifically to a method, computing device, computer-readable storage medium, and computer program product for generating target operator code. Background Technology
[0002] Traditional methods for generating operator code include manual coding, automatic matching and optimization based on internal templates, and AI-assisted generation. However, manual coding requires highly skilled engineers, is time-consuming, and inefficient. Automatic matching and optimization based on internal templates is time-consuming and struggles to support the generation of non-standard operator code. AI-assisted generation lacks systematic verification during the operator code generation process.
[0003] In summary, the shortcomings of traditional operator code generation methods are: long processing time, low efficiency, difficulty in supporting the generation of non-standard operator codes, and lack of systematic verification during the operator code generation process. Summary of the Invention
[0004] This invention provides a method, computing device, medium, and program product for generating target operator code, which can not only significantly improve the efficiency of operator code generation, but also be applicable to the generation of non-standard operator code, and provide systematic verification during the operator code generation process.
[0005] According to a first aspect of the present invention, a method for generating target operator code for running on an artificial intelligence chip is provided. The method includes: obtaining reference implementation code of the operator and an algorithm summary document associated with the operator; generating a structured specification document of the operator on a target processor of the artificial intelligence chip based on the reference implementation code and the algorithm summary document; generating target operator code for running on the target processor based on the structured specification document; and performing correctness verification on the target operator code based on the reference implementation code and outputting the verified target operator code. Generating the structured specification document of the operator on the target processor of the artificial intelligence chip based on the reference implementation code and the algorithm summary document includes: performing structured parsing on the reference implementation code and the algorithm summary document to extract algorithmic feature information of the operator; and mapping each computational dimension of the operator to hardware parameters of the target processor based on the algorithmic feature information to generate the structured specification document.
[0006] In some embodiments, obtaining the reference implementation code of the operator and the algorithm summary document associated with the operator includes: matching and obtaining the corresponding reference implementation code and algorithm summary document from a preset operator code library based on the operator identifier of the operator.
[0007] In some embodiments, matching and obtaining the corresponding reference implementation code and algorithm summary document from a preset operator code library based on the operator identifier of the operator includes: matching the corresponding reference implementation code and algorithm summary document in the preset operator code library based on a multi-priority source search strategy; and generating the corresponding reference implementation code and algorithm summary document based on the algorithm expression corresponding to the operator in response to the failure to match the corresponding reference implementation code and algorithm summary document.
[0008] In some embodiments, mapping each computational dimension of the operator to the hardware parameters of the target processor based on the algorithm feature information includes: classifying the algorithm feature information into dimensions to determine the corresponding dimension categories; determining the parallel computing mode matching the operator based on the corresponding dimension categories, and obtaining the mapping template corresponding to the parallel computing mode, wherein the parallel computing mode indicates how the operator runs in parallel on the artificial intelligence chip; and generating the structured specification document based on the corresponding mapping template and the hardware parameters of the target processor, wherein the hardware parameters of the target processor include at least the amount of shared memory, the capacity of shared memory, the number of registers, and the memory bandwidth.
[0009] In some embodiments, the dimension categories include the following categories: batch dimension, spatial parallel dimension, feature parallel dimension, reduction dimension, and broadcast dimension.
[0010] In some embodiments, generating target operator code for the operator to run on a target processor based on the structured specification document includes: parsing the structured specification document to extract implementation decision information required for running on the target processor; converting the structured specification document into API call information adapted to the target processor based on the implementation decision information; and compiling the API call information in a compiler to generate target operator code.
[0011] In some embodiments, parsing the structured specification document to extract implementation decision information required for operation on the target processor includes one or more of the following: determining the execution mode and precision based on the overview information in the structured specification document, and selecting the corresponding kernel function keywords; determining each tensor information based on the tensor information in the structured specification document, and selecting the corresponding tensor type; determining the grid dimension and number of threads based on the startup configuration information in the structured specification document, and setting kernel startup parameters; determining the block size constant and the number of burst transfers based on the circular block information in the structured specification document, and deciding whether to use template parameters or compile-time constants; determining the size and number of bytes of the shared memory buffer based on the shared memory information in the structured specification document, and writing a shared memory declaration; and determining the synchronization point location and synchronization range based on the synchronization method information in the structured specification document, and selecting the corresponding synchronization API for the target processor.
[0012] In some embodiments, converting the structured specification document into API call information adapted to the target processor based on the implementation decision information includes: extracting primitive information describing computational logic from pseudo-kernel information in the structured specification document, wherein the pseudo-kernel information is a transformation mapping table for converting the primitive information into API call information adapted to multiple target processors; and converting the primitive information into API call information adapted to the target processor.
[0013] In some embodiments, compiling in the compiler based on the API call information to generate target operator code includes: in response to a compilation failure based on the API call information, performing automatic repair according to the corresponding error information in the compiler and recompiling.
[0014] In some embodiments, performing correctness verification on the target operator code based on the reference implementation code and outputting the verified target operator code includes: performing an element-wise floating-point comparison between the output of the reference implementation code and the output of the target operator code, calculating the maximum absolute error based on the comparison result, and determining whether the maximum absolute error is less than an error threshold; and outputting the target operator code in response to the maximum absolute error being less than the error threshold.
[0015] In some embodiments, performing correctness verification on the target operator code based on the reference implementation code includes: in response to the maximum absolute error being greater than or equal to an error threshold, performing an error distribution analysis based on the maximum absolute error; updating the structured specification document based on the error distribution analysis results; and regenerating the target operator code based on the updated structured specification document so as to perform correctness verification again.
[0016] In some embodiments, the method further includes: in response to the target operator code passing the correctness verification, acquiring hardware performance data; calculating performance characterization data of the target operator code based on the hardware performance data; in response to the performance characterization data being less than a target threshold, generating an automatic optimization strategy based on the performance characterization data; updating the structured specification document based on the automatic optimization strategy; and regenerating the target operator code based on the updated structured specification document to perform the correctness verification again.
[0017] In some embodiments, calculating the performance characterization data of the target operator code based on the hardware performance data includes: obtaining the actual execution cycle of the target operator code based on the hardware performance data; obtaining the current theoretical optimal cycle and the ideal optimal cycle of the target operator code; and determining the performance characterization data based on the actual execution cycle, the current theoretical optimal cycle and the ideal optimal cycle.
[0018] In some embodiments, the method may be performed by one or more intelligent agents.
[0019] According to a second aspect of the present invention, a computing device is also provided. The computing device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the computing device to perform the method of the first aspect of the present invention.
[0020] According to a third aspect of the present invention, a computer-readable storage medium is also provided. The computer-readable storage medium stores a computer program that, when executed by a machine, performs the method of the first aspect of the present invention.
[0021] According to a fourth aspect of the present invention, a computer program product is also provided, comprising a computer program that, when executed by a machine, performs the method of the first aspect of the present invention.
[0022] This invention automatically generates a structured specification document based on the obtained reference implementation code and algorithm summary document of the operator. Based on this document, it automatically generates the target operator code to run on the target processor, thus avoiding the time-consuming and inefficient manual coding. This achieves automatic conversion from existing code and documentation to target code, significantly improving the efficiency of operator code generation. Furthermore, since the specification document is generated based on the reference implementation code and summary document, rather than relying on a fixed internal template, it can flexibly adapt to the generation of various non-standard operators. Finally, after generating the target operator code, it performs correctness verification based on the reference implementation code and outputs the final code according to the verification results, solving the problem of lacking systematic verification during operator code generation. Therefore, this invention not only significantly improves the efficiency of operator code generation but is also applicable to the generation of non-standard operator code, and provides systematic verification during the operator code generation process.
[0023] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0024] The above and other features, advantages, and aspects of the various embodiments of the present invention will become more apparent from the accompanying drawings and the following detailed description. In the drawings, the same or similar reference numerals denote the same or similar elements.
[0025] Figure 1a A block diagram of a computing apparatus for generating target operator code according to some embodiments of the present invention is shown schematically.
[0026] Figure 1b A schematic diagram of a general-purpose graphics processor 100b for generating target operator code according to some embodiments of the present invention is shown.
[0027] Figure 2 A flowchart illustrating a method for generating target operator code according to some embodiments of the present invention is shown.
[0028] Figure 3 A schematic block diagram of an artificial intelligence chip according to some embodiments of the present invention is shown.
[0029] Figure 4 A flowchart illustrating a method for generating structured specification documents according to some embodiments of the present invention is shown.
[0030] Figure 5 The flowchart illustrating a method for verifying the correctness of target operator code according to some embodiments of the present invention is shown.
[0031] Figure 6 A flowchart illustrating a method for generating an automatic optimization strategy according to some embodiments of the present invention is shown.
[0032] Figure 7 The diagram illustrates a timing diagram of agent signaling interactions in multiple stages according to some embodiments of the present invention.
[0033] In the various figures, the same or corresponding reference numerals indicate the same or corresponding parts. Detailed Implementation
[0034] Preferred embodiments of the invention will now be described in more detail with reference to the accompanying drawings. While preferred embodiments of the invention are shown in the drawings, it should be understood that the invention can be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that the invention will be thorough and complete, and will fully convey the scope of the invention to those skilled in the art.
[0035] The term "comprising" and its variations as used herein signify open inclusion, i.e., "including but not limited to". Unless otherwise stated, the term "or" means "and / or". The term "based on" means "at least partially based on". The terms "one example embodiment" and "one embodiment" mean "at least one example embodiment". The term "another embodiment" means "at least one additional embodiment". The terms "first", "second", etc., may refer to different or the same objects.
[0036] To at least partially address one or more of the aforementioned problems and other potential issues, an exemplary embodiment of the present invention proposes a method for generating target operator code. In this method, reference implementation code of the operator and an algorithm summary document associated with the operator are obtained; based on the reference implementation code and the algorithm summary document, a structured specification document for the operator on a target processor of an artificial intelligence chip is generated; according to the structured specification document, target operator code for the operator running on the target processor is generated; and based on the reference implementation code, correctness verification is performed on the target operator code, and the verified target operator code is output. The present invention automatically generates a structured specification document by obtaining the reference implementation code and the algorithm summary document, and automatically generates target operator code for running on the target processor based on this document, thereby avoiding the high time consumption and low efficiency of manual coding, achieving automatic conversion from existing code to target code, and significantly improving generation efficiency. Furthermore, since the specification document is dynamically generated based on the reference implementation code and the summary document, rather than relying on a fixed internal template, it can flexibly adapt to the generation needs of various non-standard operators. Furthermore, in some embodiments, after generating the target operator code, its correctness is verified based on the reference implementation code, and the final code is output according to the verification results, thereby establishing a systematic verification and iterative closed loop in the operator code generation process. Therefore, the present invention can significantly shorten the time spent on operator code generation and improve efficiency, is applicable to the generation of non-standard operator code, and also provides systematic verification in the operator code generation process.
[0037] Figure 1a A schematic diagram of a computing apparatus 100 for implementing a method for generating target operator code according to an embodiment of the present invention is shown. Figure 1aAs shown, the computing device 100 may have one or more processing units and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor. The processing unit includes dedicated processing units such as graphics processing units (GPUs), field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), and general-purpose computing on graphics processing units (GPGPUs), as well as general-purpose processing units such as CPUs. The computing device 100 also includes at least: an operator reference implementation code and algorithm summary document acquisition unit 102, a structured specification document generation unit 104, a target operator code generation unit 106, and a target operator code verification unit 108.
[0038] Regarding the operator reference implementation code and algorithm summary document acquisition unit 102, it is used to acquire the operator reference implementation code and the algorithm summary document related to the operator.
[0039] The structured specification document generation unit 104 is used to generate a structured specification document for the operator on the target processor of the artificial intelligence chip based on the reference implementation code and the algorithm summary document. Generating the structured specification document for the operator on the target processor of the artificial intelligence chip based on the reference implementation code and the algorithm summary document includes: performing structured parsing on the reference implementation code and the algorithm summary document to extract the algorithm feature information of the operator; and mapping each computational dimension of the operator to the hardware parameters of the target processor based on the algorithm feature information, in order to generate the structured specification document.
[0040] Regarding the target operator code generation unit 106, it is used to generate target operator code for the operator to run on the target processor according to the structured specification document.
[0041] Regarding the target operator code verification unit 108, it is used to perform correctness verification on the target operator code based on the reference implementation code, and output the verified target operator code.
[0042] Figure 1bA schematic diagram illustrates a general-purpose graphics processor 100b for generating target operator code according to an embodiment of the present invention. In some embodiments, the computing device 100 further includes a general-purpose graphics processor 100b (such as...). Figure 1b As shown). Figure 1b As shown, the general-purpose graphics processor 100b includes multiple stream processing clusters (SPCs), and the SPCs share data through a global cache / global memory.
[0043] Taking streaming multiprocessor cluster 1 as an example, it includes: vector operation units (e.g. Figure 1b The diagram shows computational units 1 and 2, and a tensor operation unit. The tensor operation unit performs tensor computations, such as matrix multiplication (MMA), convolution, etc. Taking computational unit 1 as an example, it includes multiple computational cores and a shared cache. The multiple computational cores interact and collaborate with each other through the shared cache.
[0044] In one embodiment of the present invention, one or more computational tasks run on one or more computational cores of a computing unit and independently execute one or more subtasks in a method for generating target operator code. Further, taking a computational task running on any one computational core as an example, it may include, for example: Figure 1a The unit shown is one or more of the following: operator reference implementation code and algorithm summary document acquisition unit 102, structured specification document generation unit 104, target operator code generation unit 106, and target operator code verification unit 108. These units can be located on one or more computing cores, and processing instances on multiple computing cores exchange data through a shared cache to complete the generation of the target operator code.
[0045] It is understandable that the above Figure 1b The general-purpose graphics processor and its method for generating target operator code, as illustrated, can be widely applied in artificial intelligence and high-performance computing scenarios, significantly reducing the barrier and cost of developing operator libraries for new processor architectures. This invention is particularly suitable for handling complex tasks with a wide variety of operators and diverse target hardware architectures. For example, it can be applied to scenarios such as automatic operator generation and target processor architecture adaptation in large language model training and inference deployment, automated generation of customized operators in autonomous driving computing platforms, and automatic generation of computational operators in medical imaging and scientific computing, thereby promoting the application of deep learning in a wider range of industries.
[0046] In the above scenario, Figure 1bThe general-purpose graphics processing unit (GPU) described herein executes the method described in this invention through processing instances on one or more computing cores. Based on the reference implementation code and algorithm summary document, it generates a structured specification document and target operator code through multi-stage automated processing, and outputs verified code based on a built-in gated closed-loop verification mechanism. This method supports the automatic generation of non-standard operators, significantly reducing the workload of manual optimization and search time, enabling developers to develop high-performance operators at a lower cost, improving development efficiency and code correctness. Furthermore, the multi-stage collaboration and closed-loop verification architecture adopted in this invention is not only applicable to GPU operator development but can also be extended to other software engineering tasks requiring multi-stage collaboration, demonstrating universal value.
[0047] The following will combine Figure 2 and Figure 3 A method 200 for generating target operator code, according to embodiments of the present invention, is described. It should be understood that method 200 can, for example, be implemented in... Figure 1a The described computing device 100 performs the operation. Method 200 may also include additional actions not shown and / or the actions shown may be omitted; the scope of the invention is not limited in this respect.
[0048] At step 202, the computing device 100 obtains the reference implementation code of the operator and the algorithm summary document associated with the operator.
[0049] Regarding the computing device 100, it is configured, for example, to run, invoke, or schedule one or more intelligent agents. The multiple intelligent agents include, for example, a first intelligent agent and a second intelligent agent, such that different intelligent agents correspond to different task processing functions. In some embodiments, the computing device may, for example, execute stored instructions to determine the cooperative relationship between the multiple intelligent agents, so as to provide input information to at least one intelligent agent, receive intermediate results generated by at least one intelligent agent, and generate processing results based on one or more intermediate results.
[0050] The operators include, for example, computational rules that perform specific mathematical operations or data processing logic. These computational rules are independent of the underlying hardware platform and include, for example, the transformation methods and computational steps of the input data. In some embodiments, the reference implementation code of the operator and the algorithm summary document associated with the operator form the computational logic basis for the subsequent generation of structured specification documents and target operator code.
[0051] Regarding the reference implementation code, it may indicate, for example, a verified implementation code executed by a central processing unit. The reference implementation code is a concrete software code instantiation of the operator on a general-purpose processor (such as a CPU) platform, which transforms the abstract computational logic of the operator into executable program code; simultaneously, the reference implementation code serves as the basis for comparison and verification of the target operator code generated in subsequent steps.
[0052] Regarding the algorithm summary document, it includes, for example, a structured description of the algorithmic logic characteristics and interface specifications of the operator, including, for example, mathematical formulas, tensor definitions, and algorithm steps, in order to generate a subsequent structured specification document. In some embodiments, the method for obtaining the algorithm summary document and the reference implementation code includes, for example, matching and obtaining the algorithm summary document from a preset operator code library based on the operator identifier of the operator, where the algorithm summary document is a pre-stored document corresponding to the reference implementation code; and / or, in response to not finding a matching reference implementation code and algorithm summary document, generating reference implementation code based on the algorithm expression corresponding to the operator, and extracting original information from the generated reference implementation code to constitute the algorithm summary document.
[0053] In some embodiments, an operator may correspond to a kernel function, which is executed simultaneously on multiple parallel threads of the target processor. In some embodiments, the target operator includes, for example, at least one of a convolution operator, a matrix multiplication operator, a normalization operator, and an activation function operator. Specifically, the convolution operator includes, for example, Conv2D; the matrix multiplication operator includes, for example, GEMM; the normalization operator includes, for example, BatchNorm and / or LayerNorm; and the activation function operator includes, for example, Softmax, ReLU, and / or GELU.
[0054] In some embodiments, the computing device 100 uses the reference implementation code and algorithm summary document of the operator as input information and processing objects for the entire operation process. After generating the structured specification document, it generates the target operator code to run on the target processor and outputs the correctness-verified target operator code. The overall operation process described above can be broken down into one or more pipelines. The coordination between these pipelines can be implemented using a pipeline orchestrator, the specific implementation of which will be described in detail in method 700.
[0055] A method for obtaining reference implementation code of an operator and an algorithm summary document related to the operator may include, for example, a computing device 100 matching and obtaining the corresponding reference implementation code and algorithm summary document from a preset operator code library based on the operator identifier of the operator.
[0056] Operator identifiers for operators include, for example, operator name, tensor shape, data type, and forward / backward identifiers. The operator name indicates, for example, the specific computational operation performed by the operator (e.g., convolution, pooling, activation function, etc.). The tensor shape indicates, for example, the size of the input or output tensor in each dimension (e.g., batch size, number of channels, height, width). The data type indicates, for example, the numeric type of each element in the tensor (e.g., floating-point, integer, half-precision, etc.). The forward identifier indicates, for example, that the operator is executed during the forward propagation phase to compute the model's output. The backward identifier indicates, for example, that the operator is executed during the back propagation phase to compute gradients to update the model parameters.
[0057] In some embodiments, the operator identifier also includes stride, dilation, and padding. The stride, for example, indicates the distance the convolutional kernel or sliding window moves across the input tensor with each step, used to control the size of the output feature map. The dilation, for example, indicates the spacing between adjacent elements in the convolutional kernel, used to expand the receptive field without increasing parameters. Padding, for example, indicates the number of zero-valued layers added around the boundaries of the input tensor, used to control the spatial size of the output feature map and preserve edge information.
[0058] Regarding the method by which the computing device 100 matches and obtains the corresponding reference implementation code and algorithm summary document from a preset operator code library based on the operator identifier of the operator, the method includes, for example, the computing device 100 matching the corresponding reference implementation code and algorithm summary document in the preset operator code library based on a multi-priority source search strategy; and in response to the failure to match the corresponding reference implementation code and algorithm summary document, generating the corresponding reference implementation code and algorithm summary document based on the algorithm expression corresponding to the operator.
[0059] Regarding the reference implementation code, it may indicate, for example, a verified operator algorithm implementation running on a CPU, serving as a standard for correctness comparison (e.g., the Golden reference implementation). In some embodiments, the reference implementation code may be, for example, a templated C / C++ function, such as a hostXxx() function, which may be of the form template.<typename E> In some embodiments, templated C / C++ functions may use internal precision (e.g., float precision) for numerical calculations, thereby facilitating the embedding of subsequent test code.
[0060] The predefined operator codebase includes, for example, version branch libraries, third-party dependency libraries, and C++ interface libraries. Version branch libraries are, for example, independent copies of versions derived from the main development line within the PyTorch code repository (e.g., the PyTorch branch or the BrPyTorch branch). Third-party dependency libraries are, for example, independent collections of code maintained by external organizations and integrated into other projects (e.g., the PyTorch ATen and torchvision / mmcv libraries). C++ interface libraries are, for example, C++ codebases for torch, including many APIs available in PyTorch (e.g., the Libtorch library).
[0061] Furthermore, regarding the method for matching corresponding reference implementation code and algorithm summary documents in a preset operator code library based on a multi-priority source search strategy, it includes, for example, starting with a PyTorch branch code library as the starting point, and if no corresponding reference implementation code and algorithm summary document are found in the preset operator code library, searching in third-party dependency libraries, and if no corresponding reference implementation code and algorithm summary document are found in the third-party dependency library, continuing the search in the Libtorch library.
[0062] In some embodiments, in response to finding available source code in a preset operator code library, the core algorithm logic is extracted from the source code, and a C / C++ template function that can run without relying on an external database is generated based on the extracted core algorithm logic, which is equivalent to the reference implementation code in step 202.
[0063] In some embodiments, a method for generating corresponding reference implementation code and algorithm summary document based on the algorithm expression corresponding to the operator in response to the failure to find a matching reference implementation code and algorithm summary document includes, for example, the following: in response to the failure to find a corresponding reference implementation code and algorithm summary document in a preset operator code library, the computing device 100 generates corresponding reference implementation code through an agent based on the algorithm expression, thereby ensuring the availability of the reference implementation code.
[0064] For example, when the computing device 100 performs a search based on a multi-priority source search strategy, if it fails to find the corresponding reference implementation code and algorithm summary document in the preset operator code library, it generates the corresponding reference implementation code and algorithm summary document based on manually input C / C++. In this way, the multi-priority source search ensures the accuracy of the matched reference implementation code: the multi-level source search strategy, from PyTorch branches to handwritten C++, maximizes the use of existing verified algorithm implementations and only degrades to less reliable reference implementation code sources when necessary.
[0065] In step 204, the computing device 100 generates a structured specification document of the operator on the target processor of the artificial intelligence chip based on the reference implementation code and the algorithm summary document.
[0066] A method for generating a structured specification document of the operator on the target processor of an artificial intelligence chip based on the reference implementation code and the algorithm summary document includes, for example, the following steps: A computing device 100 extracts mapping features from the reference implementation code and categorizes the extracted mapping features to match the operators in the reference implementation code to known parallel computing modes; the computing device 100 also obtains the hardware parameters of the target GPU, maps them according to the GPU's hardware parameters, and generates the corresponding structured specification document. Further details will be described in detail in method 400 and will not be repeated here.
[0067] The structured specification document, for example, contains a standardized template for generating target operator code for a target processor, comprising one or more standard document sections. This standardized template includes hints about the operator implementation details for the target processor. These details include, for example, overview information, tensor information (e.g., tensor contracts), startup configuration information, loop block information, shared memory information, synchronization method information, and pseudo-kernel information. In some embodiments, the standardized template also includes information such as data flow, synchronization scheme, occupancy model, performance model, pseudo-kernel, boundary handling, and optimization hints.
[0068] The standard documentation section contains pseudo-kernel information described by abstract primitives, which indicates specific hardware operations to be performed on a target processor (e.g., a GPU). These specific hardware operations are, for example, hardware-independent operations written using an abstract primitive framework, including, for example, one or more operations targeting the target processor (e.g., a GPU), such as parallel indexing, memory operations, synchronization barriers, reduction operations, type conversions, and mathematical operations. The abstract primitives describe the semantic intent of the operations and are independent of the API names of any specific processor platform, allowing the same standard documentation section to be flexibly adapted to code on different processors or platforms. For example, the same standard documentation section can be flexibly translated into code for multiple platforms such as CUDA, SUPA, and HIP. This standard documentation section can be configured as a pluggable translation map to adapt to multiple platforms after translation.
[0069] In step 206, the computing device 100 generates target operator code for the operator to run on the target processor, based on the structured specification document.
[0070] A method for generating target operator code for running on a target processor based on the structured specification document includes, for example: a computing device 100 parses the structured specification document to extract implementation decision information required for running on the target processor; based on the implementation decision information, converts the structured specification document into API call information adapted to the target processor; and based on the API call information, compiles it in a compiler to generate target operator code.
[0071] In some embodiments, a method for parsing the structured specification document to extract implementation decision information required for operation on the target processor includes, for example, one or more of the following: determining the execution mode and precision based on overview information in the structured specification document, and selecting the corresponding kernel function keyword; determining each tensor information based on tensor information in the structured specification document, and selecting the corresponding tensor type; determining the grid dimension and number of threads based on startup configuration information in the structured specification document, and setting kernel startup parameters; determining the block size constant and the number of burst transfers based on the circular block information in the structured specification document, and deciding whether to use template parameters or compile-time constants; determining the size and number of bytes of the shared memory buffer based on the shared memory information in the structured specification document, and writing a shared memory declaration; and determining the synchronization point location and synchronization range based on the synchronization method information in the structured specification document, and selecting the corresponding synchronization API for the target processor.
[0072] In some embodiments, a method for determining the execution mode and precision based on overview information in the structured specification document, and selecting the corresponding kernel function keyword, includes, for example, the computing device 100 selecting a precision keyword based on the execution mode. Exemplarily, the corresponding kernel function keyword (e.g., global_ or _global_mega) is selected according to (e.g., G-mode or T-mode), where G-mode represents the general SIMT execution mode, used only by vcores, and employs regular grid / block scheduling. T-mode represents the tensor persistence execution mode, where all SPCs are resident and a fixed grid configuration (NUM_SPCS, die_num) is used.
[0073] In some embodiments, the method for determining each tensor information and selecting the corresponding tensor type based on the tensor information in the structured specification document includes, for example, the computing device 100 selecting the corresponding tensor type (e.g., DynMatrix2D, NumaDynMatrix2D, or raw pointer) based on the tensor name, shape, data type, memory type, and layout. Here, the tensor name represents the identifier name of the tensor, the shape represents the size of each dimension of the tensor, the data type represents the type of elements in the tensor (e.g., float, int), the memory type represents the type of memory region where the tensor resides (e.g., global memory, shared memory), the layout represents the arrangement of the tensor in memory (e.g., row-major, column-major), DynMatrix2D represents a dynamic two-dimensional matrix type, NumaDynMatrix2D represents a non-uniform memory access-aware dynamic two-dimensional matrix type, and raw pointer represents a pointer to the underlying memory address.
[0074] In some embodiments, a method for determining the grid dimension and thread count, and setting kernel startup parameters based on startup configuration information in the structured specification document, includes, for example, the following: The computing device 100, based on the grid dimension and thread count per thread block extracted from the document, passes the grid dimension as a parameter to the grid configuration portion of a kernel startup function (e.g., the `suLaunchKernel` function), and passes the thread count as a parameter to a compilation instruction (e.g., the `_launch_bounds_` parameter) used to specify the maximum number of threads and the minimum number of resident blocks in the kernel function. Here, the grid dimension represents the size of the thread grid in three directions, the thread count per thread block represents the number of threads contained in each thread block, `suLaunchKernel` represents the kernel startup function, and `_launch_bounds_` represents a compilation instruction used to specify the maximum number of threads and the minimum number of resident blocks in the kernel function.
[0075] In some embodiments, the method for determining a chunk size constant and the number of burst transfers based on the tiling information in the structured specification document, and deciding whether to use template parameters or compile-time constants, includes, for example, the computing device 100 determining whether to use template parameters or a constantpr constant based on the chunk size constant and the number of burst transfers. Here, the chunk size constant represents the size of each chunk in tiling, the number of burst transfers represents the number of data units transferred in a single burst transfer, the template parameter represents a compile-time variable parameterization type, and the constantpr constant represents a compile-time constant. The tiling information indicates that large-scale data is divided into smaller chunks of a fixed size, and only one chunk is processed at a time to accommodate the limited storage capacity of the GPU. It should be understood that the shared memory and register file capacity of the GPU are limited and cannot accommodate the entire tensor at once. Through tiling information, large tensors can be divided into smaller chunks that can fit into shared memory, thereby achieving block-level data reuse, reducing the number of accesses to global memory, and thus improving bandwidth utilization. In some embodiments, the structured specification document generated in step 204 includes a description of the chunking strategy, which may include, for example, chunk size in various dimensions, register chunking, and data reuse ratio analysis.
[0076] In some embodiments, a method for determining the size and number of bytes of a shared memory buffer based on shared memory information in the structured specification document, and for writing a shared memory declaration, exemplarily includes: the computing device 100 writing a declaration keyword (e.g., "_shared_") based on the buffer definition and number of bytes in the shared memory information. The buffer definition includes a data type and a number of elements, the number of bytes represents the size of the memory bytes occupied by the buffer, and "_shared_" represents a keyword used to declare shared memory within a thread block.
[0077] In some embodiments, a method for determining the location and range of a synchronization point based on synchronization method information in the structured specification document, and selecting the corresponding synchronization API for the target processor, exemplarily includes: the computing device 100 selecting the corresponding synchronization API based on the synchronization point and the synchronization range. Here, the synchronization point represents the location where the synchronization operation is inserted, and the synchronization range represents intra-block synchronization or global synchronization. The synchronization API could be, for example, _syncthreads(), which represents the API in CUDA used for synchronization of all threads within a thread block.
[0078] In some embodiments, a method for converting the structured specification document into API call information adapted to the target processor based on the implementation decision information includes, for example: a computing device 100 extracting primitive information describing computational logic from pseudo-kernel information in the structured specification document, wherein the pseudo-kernel information is a transformation mapping table for converting the primitive information into API call information adapted to multiple target processors; and converting the primitive information into API call information adapted to the target processor.
[0079] Regarding pseudo-kernel information, it indicates, for example, the portion of the structured specification document that describes the complete kernel logic body, which contains several abstract primitives used to describe computational logic.
[0080] Regarding the primitive information, it indicates, for example, abstract primitives extracted from the pseudo-kernel information that are used to describe the operator computation logic.
[0081] Regarding the API call information, it indicates, for example, information in the form of API calls adapted to the target processor, converted from the primitive information. For example, the API call information is used for API calls of the target platform (e.g., the hardware on which the target operator code ultimately runs and its accompanying compiler).
[0082] A method for converting primitive information into API call information adapted to the target processor based on the implementation decision information includes, for example, the computing device 100 translating each abstract primitive in the pseudo-kernel information into a specific API call for the target processor, while simultaneously generating host-side memory allocation, data transfer, and kernel startup code. The translation process maintains the original specification's loop structure, stage decomposition, memory hierarchy allocation, synchronization point location, and reduction hierarchy unchanged.
[0083] In some embodiments, a method for compiling in a compiler based on the API call information to generate target operator code includes, for example,: in response to a compilation failure in the compiler based on the API call information, the computing device 100 performs automatic repair and recompiles according to the corresponding error information in the compiler.
[0084] In some embodiments, the method for generating a structured specification document of the operator on the target processor of the artificial intelligence chip further includes, for example, generating a source file after successful compilation by the computing device 100. The source file contains kernel functions, host code, and compilation artifacts. Kernel functions refer to parallel computing functions running on the target processor (such as a GPU or MLU), responsible for performing computational tasks; host code refers to control code running on the CPU, responsible for memory allocation, data transfer, kernel startup, and synchronization operations; the compilation artifacts are, for example, binary executable files that can be directly run on the target platform.
[0085] At step 208, the computing device 100 performs a correctness verification on the target operator code based on the reference implementation code and outputs the verified target operator code.
[0086] In some embodiments, a method for performing correctness verification on the target operator code based on the reference implementation code and outputting the verified target operator code includes, for example: a computing device 100 performing an element-wise floating-point comparison between the output of the reference implementation code and the output of the target operator code, calculating the maximum absolute error based on the comparison result, and determining whether the maximum absolute error is less than an error threshold; and outputting the target operator code in response to the maximum absolute error being less than the error threshold.
[0087] The method for element-wise floating-point comparison includes, for example, subtracting the output of the reference implementation code from the output of the target operator code element by element at each corresponding position, and calculating the difference between the floating-point numbers at each corresponding position.
[0088] Methods for calculating the maximum absolute error include, for example, determining the absolute values of all differences obtained from element-wise floating-point comparisons and obtaining the largest absolute value to obtain the maximum absolute error.
[0089] Regarding the error threshold, it can be, for example, a floating-point precision tolerance, i.e., the upper limit of the allowed floating-point numerical deviation, used to determine whether two floating-point outputs are consistent within an acceptable range. For example, for a 32-bit single-precision floating-point format (FP32), the error threshold can be set to 1e. -4 (i.e., 0.0001); for example, for 16-bit Brain Floating Point 16 (BF16) data, the error threshold can be set to 1e. -2 (i.e., 0.01).
[0090] The foregoing embodiments mainly describe how the steps of the present invention are executed by the computing device 100. In some embodiments, the aforementioned execution entity can be specifically implemented by an intelligent agent with execution capabilities. Through the collaborative cooperation of multiple intelligent agents, target operator code running on the target processor can be generated efficiently. This intelligent agent can be used to represent corresponding task processing logic, functional instances, software modules, model instances, service instances, processes, threads, workflow nodes, or combinations thereof. The present invention does not limit the physical hardware structure in which the intelligent agents are deployed. Each intelligent agent in the multiple intelligent agents can be logically distinguished from each other, but its program instructions, model parameters, running data, or calculation processes can be deployed, stored, or executed in the same physical component, or they can be deployed, stored, or executed in different physical components. For example, at least two intelligent agents in the multiple intelligent agents can share the same processor, the same memory, the same artificial intelligence chip, or the same computing device; or, at least two intelligent agents in the multiple intelligent agents can be supported by different processors, different memories, different artificial intelligence chips, or different computing devices, respectively. This disclosure does not limit the one-to-one correspondence between multiple intelligent agents and physical hardware.
[0091] In some embodiments, the method 200 for generating the target operator code is performed, for example, by one or more intelligent agents. Specifically, steps 202 to 208 of method 200, as well as subsequent methods 400, 500, and 600, can all be implemented collaboratively by one or more intelligent agents. For example, step 202 is performed by a first intelligent agent, step 204 by a second intelligent agent, step 206 by a third intelligent agent, step 208 by a fourth intelligent agent, and method 600 by a fifth intelligent agent.
[0092] In the above scheme, by obtaining the reference implementation code and algorithm summary document of the target operator, a structured specification document is automatically generated, and the target operator code running on the target processor is automatically generated based on this structured specification document. Therefore, this invention avoids the high time consumption and low efficiency of manual coding, achieving automatic conversion from existing code and documents to target code, significantly improving the efficiency of operator code generation. Furthermore, since the specification document is generated based on the reference implementation code and summary document, rather than relying on a fixed internal template, it can flexibly adapt to the generation of various non-standard operators, expanding the scope of application of this invention. Finally, after generating the target operator code, the correctness of the target operator code is verified based on the reference implementation code, and the final code is output according to the verification result, solving the problem of lack of systematic verification in the operator code generation process. Therefore, this invention can significantly shorten the time consumption and improve the efficiency of operator code generation, is applicable to the generation of non-standard operator code, and also provides systematic verification in the operator code generation process.
[0093] In some embodiments, the computing device 100 may further include an artificial intelligence chip 300 (such as...). Figure 3 (As shown). Regarding the artificial intelligence chip 300, it is configured, for example, to perform at least some computational operations related to at least one of a plurality of intelligent agents. In some embodiments, the artificial intelligence chip may, for example, perform corresponding computational operations based on relevant model data of at least one intelligent agent to generate corresponding intermediate results.
[0094] like Figure 3 As shown, the operator reference implementation code and algorithm summary document acquisition unit 102 is configured, for example, to be communicatively connected to the structured specification document generation unit 104, for transmitting the generated reference implementation code and algorithm summary document to the structured specification document generation unit 104; in some embodiments, the operator reference implementation code and algorithm summary document acquisition unit 102 is also configured, for example, to be communicatively connected to the target operator code verification unit 108, for transmitting the reference implementation code to the target operator code verification unit 108.
[0095] Regarding the structured specification document generation unit 104, it is configured, for example, to be communicatively connected to the operator reference implementation code and algorithm summary document acquisition unit 102, for receiving the reference implementation code and algorithm summary document output from the operator reference implementation code and algorithm summary document acquisition unit 102, and generating a structured specification document; in some embodiments, the structured specification document generation unit 104 is also configured, for example, to be communicatively connected to the target operator code generation unit 106, for sending the generated structured specification document to the target operator code generation unit 106.
[0096] Regarding the target operator code generation unit 106, it is configured, for example, to be communicatively connected to the structured specification document generation unit 104, for receiving the structured specification document output by the structured specification document generation unit 104, and generating target operator code for operation on the target processor based on the structured specification document; in some embodiments, the target operator code generation unit 106 is also configured, for example, to be communicatively connected to the target operator code verification unit 108, for receiving error distribution analysis results from the target operator code verification unit 108, and updating the structured specification document based on the error distribution analysis results; in other embodiments, the target operator code generation unit 106 is also configured, for example, to receive an automatic optimization strategy, and update the structured specification document based on the automatic optimization strategy, so as to regenerate the target operator code.
[0097] Regarding the target operator code verification unit 108, it is configured, for example, to be communicatively connected to the operator reference implementation code and algorithm summary document acquisition unit 102, for receiving reference implementation code output by the operator reference implementation code and algorithm summary document acquisition unit 102; in some embodiments, the target operator code verification unit 108 is also configured, for example, to be communicatively connected to the target operator code generation unit 106, for performing correctness verification on the target operator code generated by the target operator code generation unit 106 based on the reference implementation code, and sending error distribution analysis results to the target operator code generation unit 106; in other embodiments, the target operator code verification unit 108 is also configured, for example, to issue the target operator code that has passed the correctness verification.
[0098] In some embodiments, the computing device 100 is configured, for example, to be communicatively connected to the target operator code verification unit 108, for receiving target operator code from the target operator code verification unit 108; in some embodiments, the computing device 100 is further configured, for example, to be communicatively connected to the target operator code generation unit 106, for generating an automatic optimization strategy based on the performance data of the target operator code, and sending the automatic optimization strategy to the target operator code generation unit 106, so that the target operator code generation unit 106 updates the structured specification document and regenerates the target operator code based on the automatic optimization strategy.
[0099] In some embodiments, multiple processing stages correspond to multiple agents. Each agent (e.g., the operator reference implementation code and algorithm summary document acquisition unit 102, the structured specification document generation unit 104, etc.) relies only on its own reference materials and the output of the preceding agent, without depending on the internal states of other agents. Furthermore, there is no state sharing or resource competition among the agents. Any agent can independently change the AI model version, update the reference document, or modify its internal algorithm without affecting the behavior of other agents.
[0100] In the above scheme, a multi-agent collaborative system is constructed to generate an algorithm summary document and reference implementation code, and a structured specification document is generated based on the algorithm summary document. Furthermore, target operator code is generated based on the structured specification document, and the correctness and performance of the target operator code are verified against the reference implementation code. Thus, this invention automates the entire operator development process—from summary generation and specification generation to code generation, correctness verification, and performance verification—within a continuous chain, completing the entire operator development and verification process without manual intervention, significantly improving operator development efficiency and code quality.
[0101] In some embodiments, the agents included in the above modules interact with each other based on a standardized data format specification. This standardized data format specification, for example, is an artifact contract, configured to define one or more of the following: the type of artifact output by each agent, required fields, and semantic constraints. In other words, agents interact with each other based on the artifact contract, independent of the agent's internal data processing.
[0102] This approach imbues the agents with a highly modular character. Each agent only needs to fulfill the contract requirements, and its internal implementation can be freely adjusted. This gives the pipeline a highly modular nature, making it suitable for the applications described in this invention. By introducing data output contracts as a standardized specification for interaction between agents, this invention significantly improves the decoupling and collaborative efficiency between system modules. Each agent only needs to follow a unified output contract, and its internal implementation can evolve and optimize independently, thereby enhancing the system's scalability and maintainability. Furthermore, the standardized data interaction mechanism reduces integration complexity, contributing to improved overall pipeline operational stability and development efficiency.
[0103] As described above, computing device 100 can generate a structured specification document of the operator on the target processor of the artificial intelligence chip. Therefore, method 200 may also include, for example, a method 400 for generating a structured specification document of the operator on the target processor of the artificial intelligence chip based on the reference implementation code and the algorithm summary document.
[0104] The following will combine Figure 3 and Figure 4 A method 400 for generating structured specification documents, according to embodiments of the present invention, is described. It should be understood that method 400 can, for example, be used in... Figure 1a The described computing device 100 performs the operation. Method 400 may also include additional actions not shown and / or the actions shown may be omitted; the scope of the invention is not limited in this respect.
[0105] At step 402, the computing device 100 performs structured parsing on the reference implementation code and the algorithm summary document in order to extract the algorithm feature information of the operator.
[0106] Table 1 shows the structured definitions of reference implementation code in some embodiments, including field names, field types, meanings, and value ranges. The field function name (e.g., function_name) is a string. Its value range is in the format hostXxx, where Xxx represents the specific operator name, such as hostAdd. The field data type (e.g., template_type) is a type parameter specifying the data type of the elements in the tensor. Its value range is E = float, bf16, or fp16, where E represents the element type, float is a single-precision floating-point number, bf16 is a 16-bit floating-point number, and fp16 is a half-precision floating-point number. The field parameter list (e.g., parameters) is a list of pointers representing the memory addresses of the input and output tensors. Its value range is E... The raw pointer is a pointer to data of type E. The shape parameter field (e.g., `shape_params`) is a list of integers describing the tensor's dimensions and shape, taking values of type int such as N, C, H, and W, where N represents the batch size, C the number of channels, H the height, and W the width. The algorithm summary field (e.g., `algorithm_summary`) is text containing mathematical formulas and steps, used as an algorithm summary document input into the reference implementation code to explain the operator's specific computational logic. The source field (e.g., `source_attribution`) is a string containing the path to a framework such as BrPyTorch, PyTorch, or mmcv, used to identify the reference implementation source of the operator.
[0107] Table 1
[0108] Regarding algorithmic feature information, it may indicate, for example, all tensor information obtained through structured parsing (e.g., tensor name, shape, data type), computational volume information (e.g., including loop structures and arithmetic operations), and data dependencies.
[0109] For example, a method for parsing the reference implementation code (e.g., the structured definition shown in Table 1) includes: the computing device 100 extracts the tensor name from the parameters field in Table 1, the tensor shape from the shape_params field, and the data type from the template_type field to extract complete tensor information; regarding the extraction of computational body information, for example, the loop structure and arithmetic operation are extracted from the algorithm_summary field (i.e., the algorithm summary document part) in Table 1, and the corresponding loop structure and arithmetic operation are generated based on the mathematical formula and step description to extract the computational logic of the operator; as another example, regarding the extraction of data dependencies, the read-write dependencies between tensors are extracted jointly from the parameters field and the algorithm_summary field in Table 1.
[0110] In some embodiments, the method for generating a structured specification document of the operator on a target processor of an artificial intelligence chip further includes, for example, the computing device 100 mapping each computational dimension of the operator to the hardware parameters of the target processor based on the algorithm feature information, in order to generate the structured specification document. The method further includes the steps described in steps 404 to 408. In other words, after performing structured parsing on the reference implementation code and the algorithm summary document to obtain the algorithm feature information of the operator, steps 404 to 408 can be executed to generate a structured specification document of the operator on the target processor of the artificial intelligence chip.
[0111] At step 404, the computing device 100 performs dimensional classification on the algorithm feature information to determine the corresponding dimensional category.
[0112] Regarding the aforementioned dimension categories, they include, for example, the following categories: batch dimension, spatial parallel dimension, feature parallel dimension, reduction dimension, and broadcast dimension. Batch dimension represents the size of the sample batch; spatial parallel dimension represents the parallel partitioning of data in a spatial dimension; feature parallel dimension represents the parallel partitioning in a feature channel dimension; reduction dimension represents a dimension that requires reduction operations such as summation and averaging; broadcast dimension represents a dimension that needs to be automatically expanded during computation to match other tensor shapes.
[0113] A method for classifying the algorithm feature information by dimension to determine the corresponding dimension category includes, for example, the computing device 100 obtaining the shape and dimension name of the tensor from the tensor information of the algorithm feature information, then analyzing the loop structure and arithmetic operations on each dimension from the computational volume information, and finally determining the read / write dependencies between tensors from the data dependency relationship.
[0114] For example, if data in a certain dimension involves cross-element accumulation or summation operations in the computational volume information, it is classified as a reduction dimension; if data in this dimension is read and reused simultaneously by multiple output locations based on data dependencies, it is classified as a broadcast dimension; if data instances in this dimension are completely independent and have no cross-element dependencies, further judgment is made: if this dimension corresponds to a sliding or traversal operation in spatial location in the computational volume information, it is classified as a spatial parallel dimension; if this dimension corresponds to a linear transformation or mapping operation on feature channels in the computational volume information, it is classified as a feature parallel dimension; if this dimension represents a batch of samples and the instances are completely independent, it is classified as a batch dimension. Through the above comprehensive analysis of tensor information, computational volume information, and data dependencies, each dimension involved in the algorithm is classified into its corresponding dimension category.
[0115] Table 2 illustrates a method for classifying algorithm dimensions based on tensor information, computational volume information, and data dependencies in one embodiment of the present invention, and its mapping method on a GPU. As shown in Table 2, the batch dimension represents complete independence between instances; in other words, instances in different batches are independent of each other, and their computational logic is completely identical. A fully parallel mapping method is adopted at the grid dimension so that different grids process different batches, thereby achieving maximum parallelism. For example, matrix multiplication does not involve the batch dimension.
[0116] Regarding the spatial parallel dimension, it indicates, for example, that computations at different spatial locations within the data are independent of each other and do not depend on results from other locations. A block-based parallel mapping method is used in the grid dimension to divide the output matrix into multiple sub-blocks, each computed by a different thread block, thus achieving spatial parallelism. For example, the M rows and N columns of the output matrix in matrix multiplication both belong to the spatial parallel dimension.
[0117] Regarding the parallel-feature dimension, it indicates, for example, that the computation between feature channels is a fine-grained transformation or mapping, with no dependency between them, employing a parallel mapping method at the thread level so that different threads can process different feature channels, thereby achieving fine-grained parallelism. For example, the vector width in vectorization operations belongs to the parallel-feature dimension.
[0118] Regarding the reduction dimension, it indicates, for example, that multiple elements on this dimension need to be accumulated or summed to a single result, that there are data dependencies between the elements, and that a mapping method using loops and shared memory (SM) is employed so that by looping through all elements and performing partial accumulation in shared memory, the reduction result is finally obtained. For example, the K dimension in matrix multiplication is a reduction dimension.
[0119] Regarding broadcast dimensions, for example, it indicates that data in this dimension is repeatedly read from multiple computation locations, exhibiting a one-to-many data reuse relationship. It employs shared memory or register mapping to load data into high-speed storage for multiple threads to share, avoiding repeated readings from global memory. For instance, the mean and variance in layer normalization belong to broadcast dimensions.
[0120] Table 2
[0121] At step 406, the computing device 100 determines the parallel computing mode matching the operator based on the corresponding dimension category, and obtains the mapping template corresponding to the parallel computing mode, wherein the parallel computing mode indicates the way the operator runs in parallel on the artificial intelligence chip.
[0122] Regarding parallel computing modes, this refers, for example, to how operators run in parallel on an artificial intelligence chip. Examples of parallel computing modes include, for instance, Element-wise operators, Reduction operators, GEMM (Generalized Matrix Multiplication), Conv2D (Two-Dimensional Convolution), Softmax (Softmax function), LayerNorm (Layer Normalization), Attention (Attention Mechanism), Scatter / Gather (Scatter / Gather operators), and many other operators.
[0123] Regarding the mapping template, it indicates, for example, a pre-stored structured operator execution scheme for the aforementioned parallel computing mode. In some embodiments, the mapping template includes, for example, a general parallel strategy on the target processor and input parameters.
[0124] At step 408, the computing device 100 generates the structured specification document based on the corresponding mapping template and the hardware parameters of the target processor. The hardware parameters of the target processor include at least the amount of shared memory, the capacity of shared memory, the number of registers, and the memory bandwidth.
[0125] Among them, the number of computing units (SMs) represents the number of streaming multiprocessors in the target processor; the shared memory per SM represents the initial size of the high-speed temporary storage inside the SM (streaming multiprocessor); the number of registers per SM represents the total number of registers available for threads in each SM; and the memory bandwidth represents the amount of data that can be transferred between the target processor and the computing units per second.
[0126] The method for generating the structured specification document based on the corresponding mapping template and the hardware parameters of the target processor includes, for example, the computing device 100 using the hardware parameters of the target processor as constraints to calculate the input parameters that meet the hardware limitations of the target processor and have the best performance.
[0127] Using the above method, based on the algorithm's feature information, each computational dimension of the operator is mapped to the hardware parameters of the target processor, enabling the structured specification document to accurately reflect the correspondence between algorithm features and hardware resources. After dimensional classification, corresponding parallel computing modes and mapping templates are matched based on the corresponding dimension categories, thus achieving automatic alignment between algorithm computation requirements and hardware parameters, thereby accelerating the generation speed from algorithm analysis to hardware mapping. Furthermore, by adopting standardized mapping templates and replacing natural language descriptions with precise field definitions, the mapping process for different operators follows a unified specification format, eliminating ambiguities and omissions caused by inconsistent document formats in traditional processes, ensuring the accuracy and reusability of the mapping results. Simultaneously, this mapping template is built based on abstract primitives, allowing the same structured specification document to be adapted to code generation processes on multiple hardware platforms, achieving the technical effect of "analyze once, reuse multiple times." Therefore, the generation efficiency and reliability of the operator's structured specification document are significantly improved.
[0128] As described above, the computing device 100 performs correctness verification on the target operator code based on the reference implementation code and outputs the verified target operator code. Therefore, the method 200 may, for example, include a method 500 for regenerating the target operator code based on error distribution analysis. The following will be combined with... Figure 3 and Figure 5 This invention describes a method 500 for performing correctness verification on target operator code, according to embodiments of the present invention. It should be understood that method 500 can, for example, be implemented in... Figure 1a The described computing device 100 performs the operation. Method 500 may also include additional actions not shown and / or the actions shown may be omitted; the scope of the invention is not limited in this respect.
[0129] At step 502, the computing device 100 performs an element-wise floating-point comparison between the output of the reference implementation code and the output of the target operator code, calculates the maximum absolute error based on the comparison result, and determines whether the maximum absolute error is less than the error threshold.
[0130] In some embodiments, the output of the reference implementation code and the output of the target operator code are, for example, floating-point tensors with the same dimensions and shape.
[0131] Furthermore, the method for calculating the maximum absolute error based on the comparison results includes, for example, the computing device 100 traversing each pair of elements at the same position in the two tensors based on the index position of each element in the tensor, calculating the absolute difference of each pair of elements, and taking the maximum value of the absolute difference as the maximum absolute error of the current round of comparison. If the maximum absolute error is less than the error threshold, the verification is deemed to have passed; otherwise, the verification is deemed to have failed.
[0132] Regarding the error threshold, it is configured, for example, to be set based on the data type. For instance, the error threshold for a 32-bit single-precision floating-point format (FP32) could be set to 1e. -4 The error threshold for 16-bit Brain Floating Point (BF16) format can be set to 1e. -2 The error threshold for 16-bit half-precision floating-point (FP16) format can be set to 1e. -3 .
[0133] At step 504, the computing device 100 outputs the target operator code in response to the maximum absolute error being less than the error threshold.
[0134] In some embodiments, in response to the maximum absolute error being less than an error threshold, the computing device 100 determines that the target operator code and the reference implementation code are consistent within the error threshold range, and then outputs the target operator code.
[0135] At step 506, the computing device 100 performs an error distribution analysis based on the maximum absolute error in response to the maximum absolute error being greater than or equal to an error threshold.
[0136] In some embodiments, a method for performing error distribution analysis based on the maximum absolute error in response to the maximum absolute error being greater than or equal to an error threshold includes, for example: the computing device 100 extracts sub-tensors corresponding to the regions responsible for each stream processor cluster (SPC) in the target processor from the reference implementation code according to the task allocation table (e.g., spc_map) of the target processor; compares the sub-tensors extracted from the reference implementation code with the local output tensors corresponding to each SPC region element by element; in response to the comparison result error of at least one SPC region exceeding the threshold, the SPC region is considered to be abnormal.
[0137] At step 508, the computing device 100 updates the structured specification document based on the error distribution analysis results.
[0138] In some embodiments, the method for updating the structured specification document based on the error distribution analysis results includes: the computing device 100 optimizing the corresponding prompts for operator implementation details in the structured specification document based on the error distribution analysis results. The operator implementation details include, for example, overview information, tensor information (e.g., tensor contract), startup configuration information, circular block information, shared memory information, synchronization method information, and pseudo-kernel information.
[0139] At step 510, the computing device 100 regenerates the target operator code based on the updated structured specification document in order to perform the correctness verification again.
[0140] In the above scheme, the output of the reference implementation code is compared element-by-element with the output of the target operator code using floating-point comparison. The maximum absolute error is calculated based on the comparison results, and it is determined whether the maximum absolute error is less than an error threshold. It should be understood that during automated operator generation, due to differences in hardware architecture or numerical precision deviations introduced by optimization strategies, the directly generated target operator code may contain errors exceeding the allowable range, rendering the calculation results unusable. However, the correctness verification method of this invention automatically performs precision verification after the target operator code is generated. When the error exceeds the threshold, the structured specification document is updated based on the error distribution analysis results. Then, the target operator code is regenerated based on the updated document, and correctness verification is performed again, forming an iterative closed loop of verification, analysis, updating, and regeneration. This mandatory correctness gating ensures the verification principle of "correctness first, optimization later." Therefore, this invention not only improves the automation level of operator generation but also ensures, from a process perspective, that the final output target operator code meets the accuracy requirements, achieving a balance between efficiency and correctness.
[0141] In some embodiments, the method 200 may further include, for example, a method 600 for automatically updating structured specification documents. The following will be combined with... Figure 3 and Figure 6 This invention describes a method 600 for automatically updating structured specification documents, according to embodiments of the present invention. It should be understood that method 600 can, for example, be used in... Figure 1a The described computing device 100 performs the operation. Method 600 may also include additional actions not shown and / or the actions shown may be omitted; the scope of the invention is not limited in this respect.
[0142] At step 602, the computing device 100 acquires hardware performance data in response to the target operator code passing the correctness verification.
[0143] The collected hardware performance data includes, for example, an overview, instruction distribution, stall breakdown, memory bandwidth, and L2 cache. The overview represents macroscopic indicators of the target processor's operating status, such as utilization; instruction distribution indicates the execution ratio of various instructions to analyze instruction mix; stall breakdown shows the causes and proportions of pipeline stalls; memory bandwidth measures the actual bandwidth utilization of various memory levels, including global memory and shared memory; and L2 cache data includes L2 cache hit rate, miss rate, and requested read / write volume to assess the impact of caching behavior on performance.
[0144] At step 604, the computing device 100 calculates the performance characterization data of the target operator code based on the hardware performance data.
[0145] A method for calculating performance characterization data of the target operator code based on the hardware performance data includes, for example: a computing device 100 obtains the actual execution cycle of the target operator code based on the hardware performance data; obtains the current theoretical optimal cycle and the ideal optimal cycle of the target operator code; and determines the performance characterization data based on the actual execution cycle, the current theoretical optimal cycle and the ideal optimal cycle.
[0146] Regarding the actual execution cycle of the target operator code, it indicates, for example, the execution cycle of the target operator code measured based on a hardware performance counter (PFC) during execution. The hardware performance counter is used to collect performance-related data of the target operator code during hardware execution.
[0147] The theoretically optimal cycle time of the target operator code under the current algorithm indicates, for example, the theoretically optimal execution cycle time that the target operator code can achieve under the current algorithm.
[0148] Regarding the ideal optimal period of the target operator code, it indicates, for example, the ideal theoretical optimal period that the target operator code can achieve under ideal theoretical conditions.
[0149] The performance characterization data includes, for example, implementation efficiency and algorithm efficiency. The implementation efficiency is determined based on the ratio of the theoretically optimal cycle time of the current algorithm to the actual execution cycle time, and the algorithm efficiency is determined based on the ratio of the ideal optimal cycle time to the theoretically optimal cycle time of the current algorithm.
[0150] The calculation method for the aforementioned efficiency is shown in formula (1) for example:
[0151] op_score=T_algo / T_actual(1)
[0152] Here, op_score represents the implementation efficiency, which is used to measure the implementation efficiency of the target processor's computing core (e.g., GPU Kernel); T_algo represents the theoretically optimal cycle time of the current algorithm; and T_actual represents the actual execution cycle time.
[0153] In some embodiments, when the operator performance score op_score is close to 1, it indicates that the actual performance of the kernel is close to the theoretical limit of the current algorithm on the current hardware; when the operator performance score op_score is low, it indicates that there are inefficient operations in the implementation that can be optimized, such as unnecessary memory accesses, synchronization overhead, and instruction idleness. By calculating the operator performance score, it can be used as a quantitative target for performance optimization; when the operator performance score meets the target, the pipeline completes; when the operator performance score does not meet the target, optimization suggestions are generated and an iterative loop is entered.
[0154] The calculation method for the efficiency of the algorithm is shown in formula (2) for example:
[0155] algo_overhead=T_ideal / T_algo(2)
[0156] Wherein, `algo_overhead` represents the algorithm efficiency; `T_ideal` represents the ideal optimal cycle time; and `T_algo` represents the theoretically optimal cycle time of the current algorithm. Therefore, the actual execution cycle time, the theoretically optimal cycle time of the current algorithm, and the ideal optimal cycle time can be compared to calculate the implementation efficiency and the algorithm efficiency, thereby quantifying the distance between the performance of the generated target operator code and its theoretical limit.
[0157] At step 606, the computing device 100 generates an automatic optimization strategy based on the performance characterization data in response to the performance characterization data being less than a target threshold.
[0158] In some embodiments, a method for generating an automatic optimization strategy based on the performance characterization data includes, for example, the computing device 100 mapping at least one type of performance characterization data (e.g., op_score, overall_score, or PFC bottleneck metric) from the performance characterization data to an optimization mapping decision table to generate an automatic optimization strategy. The automatic optimization strategy may include, for example, various optimization strategies such as memory access optimization, algorithm modification, and synchronization overhead elimination.
[0159] In some embodiments, the computing device 100 optimizes for one of the automatic optimization strategies in each round of optimization, thereby facilitating accurate attribution of performance changes in the optimized target operator code.
[0160] In some embodiments, the generated automatic optimization strategy includes, for example, a PFC analysis report. The PFC analysis report may include, for example, a PFC analysis report in Markdown format. In some embodiments, the PFC report may include, for example, performance metrics, bottleneck classification, op_score calculation, and optimization suggestions.
[0161] Table 3 presents PFC analysis reports from some embodiments of the present invention, including, for example, bottleneck classification, implementation efficiency score, overall efficiency score, execution unit stall breakdown (EUStall), memory bandwidth metrics, and optimization suggestions. The bottleneck classification `bottleneck_type` is, for example, an enumeration type used to represent the performance bottleneck type of the target operator code, including, for example, memory-bound, compute-bound, or balanced. `memory-bound` indicates that the performance bottleneck is mainly limited by memory access; `compute-bound` indicates that the performance bottleneck is mainly limited by computational power; and `balanced` indicates that the performance bottleneck is relatively balanced between memory access and computational power. The implementation efficiency score `op_score` is, for example, a floating-point type used to represent the implementation efficiency score of the target operator code, with a value range of, for example, [0.0, 1.0]. The overall efficiency score `overall_score` is, for example, a floating-point type used to represent the overall efficiency score of the target operator code, with a value range of, for example, [0.0, 1.0]. EU Stall decomposition (stall_breakdown) is, for example, a table type used to represent the decomposition results of execution unit pauses, including, for example, the percentage corresponding to each pause cause. Memory bandwidth metric (memory_bw) is, for example, a table type used to represent the memory bandwidth metric during the execution of the target operator code, with units such as Bytes / Cycle. Optimization suggestions (optimization_suggestions) is, for example, a list type used to represent optimization suggestions generated for the target operator code, with content such as strategy description text.
[0162] Table 3
[0163] At step 608, the computing device 100 updates the structured specification document based on the automatic optimization strategy.
[0164] In some embodiments, the method for updating the structured specification document based on the automatic optimization strategy includes one or more of the following: optimizing the corresponding prompts for operator implementation details in the structured specification document based on the automatic optimization strategy. The operator implementation details include, for example, overview information, tensor information (e.g., tensor contract), startup configuration information, circular block information, shared memory information, synchronization method information, and pseudo-kernel information.
[0165] At step 610, the computing device 100 regenerates the target operator code based on the updated structured specification document in order to perform the correctness verification again.
[0166] In some embodiments, in response to the updated target operator code passing the correctness verification, the target operator code is output; in response to the updated target operator code failing the correctness verification, the process returns to step 608 so that the target operator code can be regenerated.
[0167] In the above scheme, hardware performance data is collected in response to the target operator code passing correctness verification, and performance characterization data is calculated based on this data. Then, when the performance characterization data is less than the target threshold, an optimization strategy is automatically generated and the structured specification document is updated. Finally, the target operator code is regenerated based on the updated document to perform correctness verification again. It should be understood that in traditional operator code generation processes, a systematic verification method is lacking. Performance evaluation and optimization often rely on manual experience for repeated trial and error, with each iteration taking several hours, and it is difficult to guarantee the stability and reproducibility of the optimization results. However, the method of this invention achieves a fully automated performance evaluation and iterative optimization closed loop, reducing the time for each optimization iteration from several hours to minutes, significantly improving the efficiency and quality of operator code generation. Therefore, this invention effectively solves the problem of lacking systematic verification in the operator code generation process, ensuring continuous optimization of operator performance and the standardization of code generation.
[0168] In some embodiments, this disclosure can sequentially schedule each agent to perform interactive tasks based on the operator development requirements, and return the final product according to the interaction results. Therefore, the automated operator development process may, for example, include an interactive method 700 for generating target operator code. The following will be combined with... Figure 7 The present invention describes a method 700 for generating target operator code according to embodiments thereof. Method 700 may further include additional actions not shown and / or actions shown may be omitted; the scope of the invention is not limited in this respect.
[0169] like Figure 7 As shown, method 700 involves multiple interactive agents, including, for example, a user, a pipeline orchestrator, and agents in the first to fifth stages.
[0170] At step 702, the user sends an operator development request to the pipeline orchestrator.
[0171] In some embodiments, the user is, for example, the initiator of the pipeline, who submits operator development requests and ultimately receives the development artifacts.
[0172] At step 704, the pipeline orchestrator parses the operator development requirements and starts the pipeline.
[0173] In some embodiments, the pipeline orchestrator, for example, serves as the control center of the pipeline, for receiving user requests, sequentially dispatching tasks to agents at each stage, receiving returned products, performing gating decisions, and managing optimization iteration loops.
[0174] At step 706, the pipeline orchestrator sends a task to the agent in the first stage to trigger the acquisition of reference implementation code for the operator and an algorithm summary document associated with the operator.
[0175] At step 708, the agent in the first phase performs a multi-priority source search.
[0176] In some embodiments, the agent in the first phase is responsible for obtaining reference implementation code for an operator and an algorithm summary document associated with the operator, and independently performs a multi-priority source search and returns the reference implementation code and the algorithm summary document.
[0177] At step 710, the agent in the first stage returns the reference implementation code and algorithm summary document to the pipeline orchestrator.
[0178] At step 712, the pipeline orchestrator sends a task to the agent in the second stage to trigger the generation of a structured specification document of the operator on the target processor of the artificial intelligence chip.
[0179] In some embodiments, the agent in the second stage is used, for example, to generate a structured specification document of the operator on the target processor of the artificial intelligence chip, independently perform dimension classification and mapping scheme generation based on the received reference implementation code and algorithm summary document, and return the structured specification document.
[0180] At step 714, the agent in the second phase generates a structured specification document.
[0181] At step 716, the agent in the second phase returns a structured specification document to the pipeline orchestrator.
[0182] At step 718, the pipeline orchestrator sends a task to the third-stage agent to trigger the generation of target operator code for the operator to run on the target processor.
[0183] In step 720, the third-stage agent parses the structured specification document, generates the target operator code, and compiles and runs it.
[0184] In some embodiments, the third-stage agent is used, for example, to generate target operator code that the operator runs on the target processor, such as by performing a translation process based on a received structured specification document and returning target operator code executable by the target platform.
[0185] At step 722, the third-stage agent returns the target operator code to the pipeline orchestrator.
[0186] At step 724, the pipeline orchestrator triggers a task in the fourth phase to perform a correctness verification against the target operator code based on the reference implementation code.
[0187] In some embodiments, the fourth stage, for example, is used for correctness gating, whereby a comparison test is performed after receiving the compiled target operator code and reference implementation code, and the verification result is returned.
[0188] At step 726, a correctness verification is performed in the fourth phase.
[0189] At step 728, the verification results are returned to the pipeline orchestrator in the fourth stage.
[0190] In some embodiments, in response to a code correctness repair verification that fails the verification result, step 730 is performed for code correctness repair verification, which includes steps 730a to 730d.
[0191] Specifically, in step 730a, the pipeline orchestrator sends feedback error information to the agent in the third stage, requesting repair; in step 730b, the agent in the third stage returns the repaired code to the pipeline orchestrator; in step 730c, the pipeline orchestrator sends a retest instruction to the fourth stage; and in step 730d, the fourth stage returns the verification result to the pipeline orchestrator.
[0192] Step 730 (e.g., including steps 730a, 730b, 730c, and 730d) illustrates a gated rollback stage in one embodiment of the present invention. This stage includes, for example, the pipeline orchestrator dispatching a task to the agent in the fourth stage to perform correctness verification on the target operator code based on the reference implementation code. In response to a failed verification result (e.g., as indicated by label 730d), the pipeline orchestrator feeds back the error information to the agent in the third stage for repair, and resubmits the repaired target operator code to the fourth stage for verification. This process may be repeated multiple times until verification passes. In this way, a targeted repair and retesting mechanism is automatically triggered when a correctness defect is detected, allowing errors to be located and corrected immediately. This achieves strict control over code correctness, prevents defects from flowing downstream, significantly reduces later repair costs, and ensures the reliability of basic functions.
[0193] At step 732, the pipeline orchestrator sends a task to the agent in the fifth stage to trigger the acquisition of hardware performance data and the calculation of performance characterization data. In some embodiments, the agent in the fifth stage, for example, is used for performance analysis, receiving the compiled target operator code, performing data acquisition and analysis, and returning performance characterization data.
[0194] At step 734, the agent in the fifth stage performs hardware performance data acquisition and calculates performance characterization data.
[0195] At step 736, the agent in the fifth stage returns performance characterization data to the pipeline orchestrator.
[0196] In some embodiments, regarding the performance optimization iteration loop in response to performance characterization data being less than a target threshold, it proceeds to step 738, which is a performance optimization iteration loop that includes steps 738a to 738f and is executed in this loop.
[0197] Specifically, in step 738a, the pipeline orchestrator sends a trigger optimization task to the agent in the third stage; in step 738b, the agent in the third stage returns the optimized code to the pipeline orchestrator; in step 738c, the pipeline orchestrator sends a re-correction verification instruction to the fourth stage; in step 738d, the fourth stage returns a test pass result to the pipeline orchestrator; in step 738e, the pipeline orchestrator sends a re-performance analysis instruction to the agent in the fifth stage; and in step 738f, the agent in the fifth stage returns new performance characterization data to the pipeline orchestrator.
[0198] Step 738 illustrates the optimization iteration stage in one embodiment of the present invention. This includes, for example, the pipeline orchestrator dispatching a task to the agent in the fifth stage to collect hardware performance data and calculate performance characterization data. In response to the performance characterization data being less than a target threshold (e.g., as indicated by label 738f), the pipeline orchestrator passes optimization suggestions to the agent in the third stage for application optimization. After optimization, the process passes the correctness verification in the fourth stage again. Upon successful correctness verification, the agent re-enters the fifth stage to perform performance analysis and compare the old and new performance characterization data. This cyclical process continues until the performance target is met. In this way, while pursuing performance improvement, a mandatory correctness verification is introduced as a safety net, ensuring that each performance optimization does not disrupt existing functional logic and that the final output operator code meets both accuracy and performance targets.
[0199] Finally, at step 740, the pipeline orchestrator returns the final product to the user.
[0200] In some embodiments, the final deliverables returned to the user by the pipeline orchestrator may include, for example, optimized code, reference implementation code, structured specification documents, and performance characterization data.
[0201] In the above steps, steps 702, 704, 706, 712, 718, 724, and 732 illustrate the forward dispatch phase in some embodiments of the present invention. This includes, for example, a pipeline orchestrator parsing requirements to start the pipeline, and sequentially dispatching tasks to the first-stage agent to obtain the reference implementation code of the operator and the algorithm summary document related to the operator, and receiving the returned product; dispatching tasks to the second-stage agent to generate the structured specification document of the operator on the target processor of the artificial intelligence chip, and receiving the returned product; dispatching tasks to the third-stage agent to generate the target operator code running on the target processor, and receiving the returned product; dispatching tasks to the fourth-stage agent to perform correctness verification on the target operator code based on the reference implementation code, and receiving the verification result; and dispatching tasks to the fifth-stage agent to collect hardware performance data and calculate performance characterization data, and receiving the performance characterization data. Each stage only receives the output product of the previous stage and necessary context information, and does not directly communicate with other stages. In this way, the complex operator development process is decoupled into mutually isolated independent subtasks, avoiding state interference and communication coupling between stages. This achieves standardized scheduling with high cohesion and low coupling in the pipeline, ensuring the clarity of data flow and the scalability of the system architecture.
[0202] The various processes and procedures described above, such as methods 200, 400, 500, 600, and 700, can be executed at a computing device. This computing device may include, for example, at least one processor (at least one graphics processor and at least one central processing unit); and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor. In some embodiments, methods 200, 400, 500, 600, and 700 may be implemented as a computer software program or program product tangibly contained in a machine-readable medium. In some embodiments, part or all of the computer program may be loaded and / or installed on the computing device via read-only memory (ROM) and / or a communication unit. When the computer program is loaded into random-access memory (RAM) and executed by the GPU and CPU, one or more actions of methods 200, 400, 500, 600, and 700 described above can be performed.
[0203] This invention can be a method, apparatus, system, and / or computer program product. A computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of the invention. The computer-readable storage medium may be a tangible device capable of holding and storing instructions used by an instruction execution device. The computer-readable storage medium may be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof.
[0204] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to various computing / processing devices, or downloaded via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network, to an external computer or external storage device. Various aspects of the invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0205] These computer-readable program instructions can be provided to the central processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the central processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of a flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of a flowchart and / or block diagram.
[0206] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
[0207] It should be understood that the various forms of processes shown above can be used to reorder, add, or delete steps. For example, the steps loaded in this application can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this application can be achieved, and this is not limited herein.
[0208] The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors.
Claims
1. A method for generating target operator code, characterized in that, The method includes: Obtain the reference implementation code of the operator and the algorithm summary document related to the operator; Based on the reference implementation code and the algorithm summary document, a structured specification document of the operator on the target processor of the artificial intelligence chip is generated; Based on the structured specification document, generate the target operator code for the operator to run on the target processor; and Based on the reference implementation code, a correctness verification is performed on the target operator code, and the verified target operator code is output. Generating a structured specification document for the operator on the target processor of an artificial intelligence chip, based on the reference implementation code and the algorithm summary document, includes: performing structured parsing on the reference implementation code and the algorithm summary document to extract the algorithm feature information of the operator; and mapping each computational dimension of the operator to the hardware parameters of the target processor based on the algorithm feature information, in order to generate the structured specification document.
2. The method according to claim 1, characterized in that, The reference implementation code for the operator and the algorithm summary document related to the operator include: Based on the operator identifier of the operator, the corresponding reference implementation code and algorithm summary document are matched and obtained from the preset operator code library.
3. The method according to claim 2, characterized in that, Based on the operator identifier of the operator, the corresponding reference implementation code and algorithm summary document are matched and obtained from a preset operator code library, including: Based on a multi-priority source search strategy, the corresponding reference implementation code and algorithm summary document are matched in the preset operator code library; In response to the lack of a matching reference implementation code and algorithm summary document, a corresponding reference implementation code and algorithm summary document are generated based on the algorithm expression corresponding to the operator.
4. The method according to claim 1, characterized in that, Based on the algorithm feature information, mapping each computational dimension of the operator to the hardware parameters of the target processor includes: The algorithm's feature information is dimensionally classified to determine the corresponding dimensional category; Based on the corresponding dimension category, the parallel computing mode matching the operator is determined, and the mapping template corresponding to the parallel computing mode is obtained. The parallel computing mode indicates the way the operator runs in parallel on the artificial intelligence chip; and Based on the corresponding mapping template and the hardware parameters of the target processor, the structured specification document is generated. The hardware parameters of the target processor include at least the amount of shared memory, the capacity of shared memory, the number of registers, and the memory bandwidth.
5. The method according to claim 4, characterized in that, The dimension categories include the following: batch dimension, spatial parallel dimension, feature parallel dimension, reduction dimension, and broadcast dimension.
6. The method according to claim 1, characterized in that, According to the structured specification document, generating the target operator code for running the operator on the target processor includes: The structured specification document is parsed to extract the implementation decision information required for operation on the target processor; Based on the implementation decision information, the structured specification document is converted into API call information adapted to the target processor; and Based on the API call information, the code is compiled in the compiler to generate the target operator code.
7. The method according to claim 6, characterized in that, Parsing the structured specification document to extract implementation decision information required for operation on the target processor includes one or more of the following: Based on the overview information in the structured specification document, the execution mode and precision are determined, and the corresponding kernel function keywords are selected; Based on the tensor information in the structured specification document, determine each tensor information and select the corresponding tensor type; Based on the startup configuration information in the structured specification document, determine the grid dimension and number of threads, and set the kernel startup parameters; Based on the circular chunking information in the structured specification document, determine the chunk size constant and the number of burst transmissions, and decide whether to use template parameters or compile-time constants. Based on the shared memory information in the structured specification document, determine the size and number of bytes of the shared memory buffer, and write the shared memory declaration; Based on the synchronization method information in the structured specification document, the synchronization point location and synchronization range are determined, and the corresponding synchronization API of the target processor is selected.
8. The method according to claim 6, characterized in that, Based on the implementation decision information, converting the structured specification document into API call information adapted to the target processor includes: Primitive information describing computational logic is extracted from the pseudo-kernel information in the structured specification document. The pseudo-kernel information is a transformation mapping table used to convert the primitive information into API call information adapted to various target processors. The primitive information is converted into API call information adapted to the target processor.
9. The method according to claim 6, characterized in that, The compilation process based on the API call information to generate the target operator code includes: In response to a compilation failure in the compiler based on the API call information, automatic repair is performed according to the corresponding error information in the compiler, and the compilation is performed again.
10. The method according to claim 1, characterized in that, Based on the reference implementation code, the correctness of the target operator code is verified, and the verified target operator code is output, including: The output of the reference implementation code is compared element-wise with the output of the target operator code using floating-point comparison. Based on the comparison results, the maximum absolute error is calculated, and it is determined whether the maximum absolute error is less than an error threshold. In response to the maximum absolute error being less than the error threshold, the target operator code is output.
11. The method according to claim 10, characterized in that, Performing correctness verification on the target operator code based on the reference implementation code includes: In response to the maximum absolute error being greater than or equal to an error threshold, an error distribution analysis is performed based on the maximum absolute error; The structured specification document is updated based on the error distribution analysis results; and The target operator code is regenerated based on the updated structured specification document in order to perform correctness verification again.
12. The method according to claim 1, characterized in that, The method further includes: In response to the target operator code passing the correctness verification, hardware performance data is collected; Based on the hardware performance data, calculate the performance characterization data of the target operator code; In response to the performance characterization data being less than the target threshold, an automatic optimization strategy is generated based on the performance characterization data; The structured specification document is updated based on the automatic optimization strategy; and The target operator code is regenerated based on the updated structured specification document in order to perform correctness verification again.
13. The method according to claim 12, characterized in that, Based on the hardware performance data, the performance characterization data of the target operator code is calculated as follows: Based on the hardware performance data, the actual execution cycle of the target operator code is obtained; Obtain the current theoretical optimal period and ideal optimal period of the target operator code; and The performance characterization data is determined based on the actual execution cycle, the theoretically optimal cycle of the current algorithm, and the ideal optimal cycle.
14. The method according to any one of claims 1-13, characterized in that, The method can be executed by one or more intelligent agents.
15. A computing device, characterized in that, include: At least one processor; as well as A memory that is communicatively connected to the at least one processor; in The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-14.
16. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a machine, performs the method according to any one of claims 1-14.
17. A computer program product, characterized in that, Includes a computer program, which, when executed by a machine, performs the method according to any one of claims 1-14.