Neural network model inference code automatic generation method and device, and electronic equipment

CN122779201APending Publication Date: 2026-09-18BEIHANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610660537.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-13
Publication Date
2026-09-18

AI Technical Summary

Technical Problem

[0004]本申请实施例的目的是提供一种神经网络模型推理代码自动生成方法、装置及电子设备,以解决现有技术中手写算子开发成本高、跨架构迁移困难的技术问题

Benefits of technology

[0010] As can be seen from the technical solutions provided in the embodiments of this application above, the embodiments of this application obtain the source code of the neural network model through a code analysis agent, convert the source code of the neural network model into a static single-assignment intermediate representation (SSA IR), and a compilation optimization agent performs compilation optimization on the SSA IR. This compilation optimization is used to reduce the number of operators and the memory access of intermediate results without changing the computational semantics. The kernel generation agent generates multiple inference units based on the compiled and optimized SSA IR, and the process orchestration agent combines these multiple inference units into inference flow code, realizing the automatic generation of neural network model inference code. Compared with the hand-written operator mode, this greatly shortens the development cycle, significantly reduces development costs, and can match the rapid iteration of neural network models. Moreover, by using a multi-agent collaborative architecture and leveraging the code understanding and generation capabilities of a large language model to generate inference flow code, the cross-architecture migration capability is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122779201A_ABST
    Figure CN122779201A_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a neural network model inference code automatic generation method and device and electronic equipment. The method comprises: a code analysis agent obtaining a neural network model source code, and converting the neural network model source code into a static single assignment intermediate representation (SSA IR); a compilation optimization agent performing compilation optimization on the SSA IR, the compilation optimization being used to reduce the number of operators and the memory access of intermediate results without changing the calculation semantics; a kernel generation agent generating a plurality of inference units according to the SSA IR after the compilation optimization; and a flow arrangement agent combining the plurality of inference units into inference flow code.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence, and in particular to a method, apparatus and electronic device for automatically generating inference code for a neural network model. Background Technology

[0002] With the widespread application of deep learning models in fields such as robot control and autonomous driving, how to efficiently deploy various neural network models on resource-constrained edge devices has become an important research topic. Especially in recent years, compared to pure language models, multimodal neural network models, which include heterogeneous structures such as visual encoders and cross-modal projectors, often need to fuse multimodal information to generate control commands. This heterogeneous computing model presents edge-side inference with more severe challenges than traditional language models.

[0003] Existing inference optimization solutions primarily rely on handwritten operator development. Inference engines, such as TensorRT, provide rich libraries of basic operators, allowing developers to write high-performance custom operators for specific models. However, the handwritten operator development model is extremely costly. An experienced developer typically needs several weeks to design and tune all operators from scratch for a new, complex neural network model, making it difficult to keep pace with rapid model iteration. Furthermore, optimization experience accumulated through handwritten operators is difficult to directly transfer to new architectures, leading to challenges in cross-architecture migration and significant duplication of development. Summary of the Invention

[0004] The purpose of this application is to provide a method, apparatus, and electronic device for automatically generating inference code for neural network models, so as to solve the technical problems of high development cost of handwritten operators and difficulty in cross-architecture migration in the prior art.

[0005] In a first aspect, the embodiments of this application provide a method for automatically generating inference code for a neural network model, comprising: The code analysis agent obtains the source code of the neural network model and converts the source code of the neural network model into a static single-assignment intermediate representation (SSA IR). The compiler optimization agent performs compiler optimization on the SSA IR, which is used to reduce the number of operators and memory access of intermediate results without changing the computational semantics; The kernel generates an intelligent agent that generates multiple inference units based on the compiled and optimized SSA IR; The process orchestration agent combines the multiple inference units into inference process code.

[0006] Secondly, embodiments of this application provide an automatic neural network model inference code generation device, comprising: A code analysis agent is used to obtain the source code of a neural network model and convert the source code of the neural network model into a static single-assignment intermediate representation (SSAIR). The compilation optimization agent performs compilation optimization on the SSA IR, which is used to reduce the number of operators and memory access of intermediate results without changing the computational semantics. The kernel generates an intelligent agent, which generates multiple inference units based on the compiled and optimized SSA IR; The process orchestration agent combines the multiple inference units into inference process code.

[0007] Thirdly, embodiments of this application provide an electronic device, including a processor, a memory, and a computer program stored in the memory and executable on the processor. When the computer program is executed by the processor, it implements the steps of the automatic generation method for neural network model inference code provided in the above embodiments.

[0008] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the neural network model inference code automatic generation method provided in the above embodiments.

[0009] Fifthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the steps of the neural network model inference code automatic generation method provided in the above embodiments.

[0010] As can be seen from the technical solutions provided in the embodiments of this application above, the embodiments of this application obtain the source code of the neural network model through a code analysis agent, convert the source code of the neural network model into a static single-assignment intermediate representation (SSA IR), and a compilation optimization agent performs compilation optimization on the SSA IR. This compilation optimization is used to reduce the number of operators and the memory access of intermediate results without changing the computational semantics. The kernel generation agent generates multiple inference units based on the compiled and optimized SSA IR, and the process orchestration agent combines these multiple inference units into inference flow code, realizing the automatic generation of neural network model inference code. Compared with the hand-written operator mode, this greatly shortens the development cycle, significantly reduces development costs, and can match the rapid iteration of neural network models. Moreover, by using a multi-agent collaborative architecture and leveraging the code understanding and generation capabilities of a large language model to generate inference flow code, the cross-architecture migration capability is improved. Attached Figure Description

[0011] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 This application provides a schematic diagram of a system architecture for automatically generating inference code for a neural network model. Figure 2 A flowchart illustrating an automatic generation method for neural network model inference code provided in an embodiment of this application; Figure 3 A schematic diagram illustrating the process of generating and verifying the SSA IR for code analysis agents provided in this application embodiment; Figure 4 A schematic diagram illustrating the process of a compilation optimization agent performing compilation optimization on SSA IR, as provided in an embodiment of this application; Figure 5 A schematic diagram illustrating the closed-loop mechanism for the kernel generation agent to perform generation, verification, repair, and regeneration provided in the embodiments of this application; Figure 6 A schematic diagram of the structure of the automatic generation device for neural network model inference code provided in the embodiments of this application; Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0013] This application provides a method, apparatus, and electronic device for automatically generating inference code for neural network models.

[0014] To enable those skilled in the art to better understand the technical solutions in this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this application.

[0015] First, the terms used in this application will be explained.

[0016] 1. SSA IR (Static Single Assignment Intermediate Representation): A program representation in which each variable is assigned a value only once, which facilitates data flow analysis and optimization.

[0017] 2. ATen (A Tensor Library) Operators: Low-level tensor operation primitives in the PyTorch framework, serving as the foundational implementation for all high-level APIs.

[0018] 3. Operator fusion: Combine multiple consecutive operators into one operator to reduce memory accesses and improve performance.

[0019] 4. Triton: A domain-specific language for writing high-performance GPU kernels, providing a higher level of abstraction than CUDA, developed by OpenAI.

[0020] 5. CUDA Graph: A technology provided by NVIDIA that can pre-record a series of CUDA kernel calls and then replay the recorded call sequence during subsequent execution, thereby eliminating CPU-side scheduling overhead.

[0021] 6. FX module: A symbolic tracing tool built into the PyTorch framework, used to capture the computational graph structure of the model. It can record the operator call sequence and tensor shape information during the forward propagation process.

[0022] 7. HBM (High Bandwidth Memory): A high-performance stacked memory technology for GPUs, characterized by high bandwidth but relatively high access latency.

[0023] 8. HLO (High Level Optimizer): A high-level intermediate representation used by the TensorFlow XLA compiler, which preserves the semantics of relatively high-level operators.

[0024] The method and apparatus for automatically generating neural network model inference code provided in this application can be applied to electronic devices, which can be terminal devices or servers. The server can be a standalone server or a server cluster composed of multiple servers, and is not specifically limited. The terminal device can be a personal computer or a mobile terminal device such as a mobile phone or tablet, and is not specifically limited. The application scenarios of the above method and apparatus include scenarios where neural network inference code is automatically generated on terminal devices or servers. Through the above method and apparatus, the automatic generation of neural network model inference code can be achieved, which greatly shortens the development cycle and significantly reduces development costs compared to the handwritten operator mode, and can match the rapid iteration of neural network models. Moreover, by using a multi-agent collaborative architecture and leveraging the code understanding and generation capabilities of a large language model to generate inference flow code, the cross-architecture migration capability is improved.

[0025] Figure 1 This is a schematic diagram of a system architecture for automatically generating inference code for a neural network model, provided as an embodiment of this application. Figure 1 As shown, the system architecture includes: CodeAnalyzer (code analysis agent), SSAOptimizer (compilation optimization agent), KernelGenerator (kernel generation agent), FlowComposer (flow orchestration agent), and Validator (validation agent throughout the entire process).

[0026] The code analysis agent CodeAnalyzer acquires the source code of neural network models, such as PyTorch models, and converts it into a fully flattened SSA IR, ensuring that each computational step is explicit. The compilation optimization agent SSAOptimizer performs one or more stages of compilation optimization on the SSA IR, reducing the number of operators and memory access to intermediate results without altering the computational semantics. The kernel generation agent KernelGenerator generates multiple inference units based on the compiled and optimized SSA IR, such as converting each SSA instruction into a corresponding Triton kernel. The flow orchestration agent FlowComposer combines discrete inference units into a complete end-to-end inference flow code, such as Triton code, completing variable mapping and scheduling orchestration. The validation agent Validator is embedded in each of the above stages to verify the correctness of the outputs at each stage, ensuring that the generated code is computationally equivalent to the original PyTorch model.

[0027] like Figure 2 As shown in the figure, this application provides a method for automatically generating inference code for a neural network model, which may specifically include the following steps: Step S202: The code analysis agent obtains the source code of the neural network model and converts the source code of the neural network model into a static single-assignment intermediate representation (SSA IR).

[0028] In this embodiment of the application, the source code of the neural network model can be of various types, such as PyTorch model source code, etc., and there is no specific limitation.

[0029] In this embodiment, the code analysis agent serves as the entry point to the entire pipeline. Its core task is to convert the high-level abstract code of the PyTorch model into a structured intermediate representation called SSA. This conversion is crucial because only when the SSA IR is completely and accurately extracted can subsequent optimization and code generation have a reliable basis.

[0030] In this embodiment of the application, the code analysis agent can receive various types of input files, including but not limited to at least one of the following: 1) The source code of the forward method of a neural network model, such as the layers.py file in a PyTorch model.

[0031] 2) Computation graph log of the neural network model. This is a log file generated by pre-computation using PyTorch's FX module. It records runtime information such as tensor shape and data type during the forward propagation of the neural network model.

[0032] When receiving the two types of input files mentioned above, these two types of input files form the basis of code analysis. The forward method source code provides the static code structure, while the computation graph log supplements the dynamic tensor shape information.

[0033] In this embodiment, prompts can also be used during the generation of SSA IR. These prompts can include four key components: First, a role definition, clearly defining the large language model's responsibility as a PyTorch-to-SSA IR translation expert, emphasizing the need to faithfully reflect the original computational semantics and not omit any operators. Second, an SSA specification document, providing a complete list of aten operators and semantic descriptions to ensure consistent output format. This specification document adopts a modular structure, including design principles, YAML format definitions, detailed descriptions of over 80 operators, and complete SSA IR examples. Third, translation examples, using typical operators such as F.silu and F.linear as examples, demonstrating how to decompose high-level APIs into corresponding operation sequences. Fourth, constraints, emphasizing that no shape transformations or type conversions should be omitted, as these operations may play a crucial role in subsequent optimizations.

[0034] like Figure 3As shown in the embodiments of this application, after the code analysis agent generates the SSA IR, a verification agent can be used to perform static verification on the SSA IR. The standard result obtained by reasoning from the neural network model source code, such as Torch source code, can be compared with the IR result obtained by reasoning from the SSA IR based on the Torch API. The SSA IR that fails verification is returned to the code analysis agent for regeneration.

[0035] Step S204: The compiler optimizes the agent to perform compiler optimization on the SSA IR. The compiler optimization is used to reduce the number of operators and memory access of intermediate results without changing the computational semantics.

[0036] In this embodiment, the SSA IR structure can include four parts. The first part: inputs, defines the input parameters for the forward operation of the neural network model. Each input can include an SSA ID, name, shape, and data type. The second part: outputs, defines the output parameters of the neural network model. The third part: weights, defines the weight parameters of the neural network model, referenced by string names rather than SSA IDs, because the weight values ​​are determined when the neural network model is loaded and do not participate in the SSA IR computation graph. The fourth part: body, contains a sequence of operator instructions arranged in execution order, and is the core of the SSA IR. Each instruction can include results (a list of result SSA IDs), opcode (operator name), operands (a list of operands), and attributes (operator attributes).

[0037] Step S206: The kernel generates an agent that generates multiple inference units based on the compiled and optimized SSA IR.

[0038] In this embodiment, kernel-generated agents are the most technically challenging part of the entire pipeline, as they convert SSA instructions into executable Triton kernels. The core challenge at this stage lies in the fact that different operators have drastically different computational characteristics and performance requirements, thus necessitating the adoption of site-specific processing strategies.

[0039] In this embodiment, the kernel-generated agent can generate multiple inference units based on prompt words using a language model. The construction strategy for these prompt words can specifically include information from the following multiple dimensions: 1) Operator definition: This includes the operator's opcode name and SSA definition, which clarifies the type of operation that needs to be implemented.

[0040] 2) Input / output specifications: including shape and data type, to ensure that the generated kernel processes the correct tensor dimensions.

[0041] 3) Original PyTorch Implementation: Provides reference implementation code of this operator in PyTorch to help large language models understand the computational semantics and expected behavior of the operator.

[0042] 4) Hardware constraints: These specify implementation requirements such as data type configuration and parallelism parameters. For example, the accumulator uses fp32, and intermediate results use bf16; automatic tuning must be enabled using @triton.autotune; and BLOCK_SIZE follows naming conventions.

[0043] 5) Reference Implementation: For complex operators, a validated reference implementation can be provided for large language models to learn from. For example, for Flash Attention, the prompt word will additionally include a standard description of the Flash Attention algorithm and a reference code snippet.

[0044] This unified generation strategy simplifies the processing logic, and through differentiated prompt word design, it can ensure that various operators can obtain appropriate implementation schemes, thereby obtaining high-quality generation results.

[0045] Step S208: The process orchestration agent combines multiple inference units into inference process code.

[0046] In the embodiments of this application, each operator is processed independently and generates its own inference unit in the previous stage. However, these inference units are isolated from each other and require a process orchestration agent to connect them into an end-to-end inference link.

[0047] In this embodiment of the application, the step S202 above, which converts the neural network model source code into a static single-assignment intermediate representation (SSA IR), may specifically include the following steps: The code analysis agent converts the neural network model source code into a static single-assignment intermediate representation (SSA IR) according to the following principles: 1) Each SSA instruction corresponds to an ATen operator, and the operator granularity in the SSA IR is consistent with the kernel granularity; The above principle can be called the granularity consistency principle, which means that the mapping relationship from SSA IR to Kernel is clear and concise, requiring no additional decomposition steps. In contrast, frameworks such as TVM / Relay HLO retain high-level operators such as dense and softmax, which are then decomposed by the compiler at compile time. While user-friendly, this increases the complexity of intermediate transformations.

[0048] 2) All shape transformation and type conversion operations are represented by explicit commands; The above principles can be called the explicitness principle. Shape transformation operations such as `view`, `unsqueeze`, `permute`, `transpose`, and `reshape` change the dimensional layout of a tensor but not its data content. Type conversion operations such as `to`, `float`, and `bfloat16` change the data type of a tensor. The advantage of this explicit representation is that the optimizer can see the complete computation graph and will not lose optimization opportunities due to high-level abstraction. Taking operator fusion as an example, if shape transformation operations are hidden within high-level APIs, the optimizer will find it difficult to determine whether adjacent operators can be safely fused.

[0049] 3) Each variable is assigned a value once and referenced through its SSA ID, forming an acyclic computation dependency graph.

[0050] The above principles can be called the SSA form principles. This form makes computational dependencies clear, facilitates data flow analysis, and allows multiple uses of values ​​to be referenced by SSA IDs instead of copied, avoiding the risk of data inconsistency and also facilitating optimizations such as dead code elimination.

[0051] In this embodiment of the application, the goal of performing compilation optimization in step S204 is to reduce the number of operators and the memory access of intermediate results without changing the computational semantics, thereby improving inference efficiency.

[0052] like Figure 4 As shown, the compilation optimization agent in step S204 above performs compilation optimization on SSA IR, which may include one or more stages of compilation optimization, specifically including the following steps: The compiler-optimizing agent performs at least one of the following compiler optimizations on SSA IR: weight preprocessing, constant folding, common subexpression elimination, dead code elimination, and operator fusion.

[0053] In the case of performing compilation optimization in multiple stages, the input of each stage is the output of the previous stage.

[0054] In this embodiment of the application, the above-mentioned weight preprocessing may specifically include: advancing the transpose operation of the weights to the neural network model loading stage, and removing the operator corresponding to the transpose operation from the SSA IR.

[0055] In the implementation of a standard linear layer, weights are stored in the form [out_features, in_features], but during computation, they need to be transposed to [in_features, out_features] and multiplied by the input matrix. In the original SSA IR, this transpose operation exists as an independent operator before each matrix multiplication instruction. Before each matrix multiplication, the program performs a transpose operation and produces an intermediate result, and this overhead accumulates repeatedly during inference.

[0056] The core idea of ​​weight preprocessing is to bring forward transpose operations—which are executed only once but referenced multiple times—to the loading stage of the neural network model. The specific implementation consists of two steps. First, identify all transpose operations applied to the weights, corresponding to the `aten.t` operator, record their results, and remove them from the computation graph. Second, iterate through the remaining operators, replacing all references to the transpose result with the original weights, and marking these weights as preprocessed weights.

[0057] An example of the optimization effect is as follows: Before optimization, the shape of the weight `linear.weight` is [4096, 2048]. The computation graph contains an `aten.t` operator that transposes it to [2048, 4096]. Subsequent matrix multiplications use the transposed result. After optimization, the shape in the weight definition is directly [2048, 4096]. The `preprocessed` flag is set to true, the transpose operator in the computation graph is removed, and matrix multiplications directly use the weights. This optimization eliminates the runtime transpose overhead, resulting in a significant performance improvement for models with multiple linear layers.

[0058] In this embodiment of the application, the above-mentioned constant folding may specifically include: pre-compiling operators whose operands are all constants at compile time, storing the calculation results in the constant storage area, and removing operators whose operands are all constants from the SSA IR.

[0059] In the actual execution of model inference, not all computations depend on dynamic input. The operands of some operators are actually fixed values, such as the eps parameter in RMSnorm (typically 1e-5) and the statistical mean in normalization layers. These constants do not change after the neural network model is loaded, and repeatedly calculating the same results for each inference is clearly wasteful.

[0060] The constant folding algorithm described above can maintain a constant storage area, traverse each operator in the SSA IR, and determine whether all its operands have been registered in the storage area. If so, the calculation corresponding to the operator is executed at compile time, the calculation result is stored in the storage area, and the operator is removed from the output; otherwise, it is retained in the processed operator sequence.

[0061] The benefits of constant folding described above are reflected on multiple levels. Regarding runtime overhead, constant computation and its instruction scheduling are completely eliminated; regarding the generated code, the reduced number of instructions translates to smaller compiled output size and faster kernel generation speed. Furthermore, the folded constant results remain in the constant storage area, which can serve as input for subsequent optimization stages such as common subexpression elimination, further eliminating redundant computations.

[0062] In this embodiment of the application, the above-mentioned common subexpression elimination may specifically include: caching the operator name, operands and calculation result for each calculated subexpression, and reusing the corresponding calculation result in the cache when encountering an operator with the same operator name and operands as those in the cache during the traversal of the SSA IR.

[0063] In SSA IR flattening, the same sub-computations often appear repeatedly in different locations. For example, suppose an intermediate variable x is generated by operator A and subsequently referenced by operators B, C, and D. Without handling, each time B, C, or D references x, the operator logic in A will be repeatedly executed, resulting in significant redundant computation.

[0064] The core strategy of common subexpression elimination is to cache each evaluated subexpression, specifically caching the operator name, operands, and computation result. When traversing the SSA IR, the operator name opcode and all its operands are used as indices for matching within the cache. If the current operator has the exact same opcode and operands as an expression in the cache, the computation result of that expression is directly reused without re-execution. Furthermore, ID remapping can be performed on all results of the current operator; that is, the discrete IDs after common subexpression elimination are reordered and remapped to consecutive IDs for easy subsequent referencing.

[0065] Since each value in an SSA IR is defined only once, and the source of operands is explicitly indicated by the SSA ID, there is no ambiguity. Therefore, by caching subexpressions to eliminate common subexpressions, redundant calculations can be effectively reduced and inference efficiency improved.

[0066] In this embodiment of the application, the above-mentioned dead code elimination may specifically include: tracing back the referenced operators from the output of the neural network model through the SSA ID, and filtering out unreferenced operators in the SSA IR.

[0067] In actual deployment of neural network models, not all intermediate variables are used in subsequent calculations. Some tensors may be debugging artifacts left over from the development phase, some activation values ​​will never be hit under certain input conditions, and some conditional branch code has become a dead path in the neural network model. The existence of this code does not affect the output of the neural network model, but it still consumes valuable computing resources.

[0068] Dead code elimination starts from the output of the neural network model and traces backwards through all actually used SSA IDs, retaining only the operators corresponding to them in the computation graph. Since the definition and referencing relationships of each value in the SSA form are explicit, dependency analysis can be precise and complete. The dead code elimination algorithm traverses the SSA IR in reverse order starting from the output ID, collecting the SSA IDs appearing in all operands to obtain a set of referenced values; then it filters the SSA IR, retaining only the operators whose resulting SSA IDs appear in this set.

[0069] This optimization often triggers a chain reaction; eliminating a piece of dead code may cause some previously indirectly referenced weights or activation values ​​to lose all their subsequent references, thus becoming new dead code, which in turn triggers further elimination. This chain-like contraction is particularly pronounced in multi-branch networks and neural network models containing a large amount of conditional logic.

[0070] In this embodiment of the application, the above-mentioned operator fusion may specifically include: merging multiple adjacent operators that have direct data dependencies into a single operator, and completing the calculation of multiple adjacent operators in a register.

[0071] After SSA transformation and scheduling optimization, the program's data flow is clear enough, but a considerable number of operators are still scattered in the SSA IR in a fine-grained form. Taking RMSnorm as an example, its complete computation process involves multiple steps such as squaring, accumulation, square root, reciprocal, and normalization. In the standard implementation, each step is an independent kernel call. This means that even if the scheduler has arranged them in adjacent positions, the intermediate results between kernels still need to be written back to HBM and then read back for the next step. For these one-time temporary variables, HBM reads and writes are essentially redundant.

[0072] The operator fusion described above merges multiple adjacent operators with direct data dependencies into a single operator, allowing all computations to be completed in registers without falling back to HBM. The fully flattened design of SSA IR enables the optimizer to clearly identify fusionable operator patterns, while explicit shape transformations and type conversions ensure the safety of fusion.

[0073] In this embodiment of the application, the above step S206, kernel generating agent, generates multiple inference units based on the compiled and optimized SSA IR, which may specifically include the following steps a to d.

[0074] Step a: The kernel generates an intelligent agent that classifies the operators in the compiled and optimized SSA IR according to their computational complexity, and uses a large language model to convert each type of operator into a corresponding inference unit.

[0075] Step b: Based on the same input, compare the source code of the neural network model and the inference results of each inference unit to determine the inference unit that failed the verification.

[0076] Step c: Analyze the reasons for the errors in the inference units that failed to be validated and generate repair suggestions.

[0077] Step d: Regenerate the corresponding inference unit based on the repair suggestions.

[0078] In this embodiment of the application, step a above may specifically include the following steps: Step a1: The kernel generates an intelligent agent and classifies the operators in the compiled and optimized SSA IR into computationally intensive operators, non-computationally intensive operators, and complex operators based on the computational complexity of the operators.

[0079] Step a2: For computationally intensive operators, use a large language model to generate a dedicated optimized kernel as an inference unit.

[0080] Among them, computationally intensive operators include, but are not limited to, matrix multiplication, activation functions, normalization, attention mechanisms, etc. These operators are the key bottlenecks in inference performance and deserve a lot of effort to optimize. Therefore, we use large language models to generate dedicated optimized kernels for these operators.

[0081] Step a3: For non-computational operators, use the corresponding operators in the neural network model source code as inference units.

[0082] Non-computational operators include, but are not limited to, shape transformation and element-wise operations. These operators either have very low computational cost or already have highly optimized underlying implementations. Manual optimization yields limited benefits, and using PyTorch source code directly can save development costs.

[0083] Step a4: For complex operators, use a large language model based on a preset template to generate adapted code as an inference unit according to the characteristics of the target hardware.

[0084] Complex operators such as Flash Attention can reuse validated handwritten templates and have a large language model generate more suitable Triton code based on the target hardware characteristics. These operators have high algorithmic complexity, but mature implementations already exist. Direct reuse can reduce development risks and significantly improve the generation speed of end-to-end operators.

[0085] Since not all operators require or are worth investing a lot of effort in generating dedicated kernels, classifying operators in SSAIR and allocating resources reasonably can maximize automation and reduce human intervention while ensuring generation quality.

[0086] like Figure 5 As shown in this embodiment, the kernel generation agent can use a closed-loop mechanism of generation, verification, repair, and regeneration. Multiple generated inference units need to be verified before proceeding to subsequent processes. Verification refers to using a verification agent to verify the numerical correctness of the multiple inference units generated by the kernel generation agent. If the error exceeds a preset threshold, a repair process is triggered. Specifically, taking the generated inference unit as an example, under the same input, the kernel's output is compared with the output of the PyTorch source code to determine if their values ​​are consistent. The testing method is as follows: a set of random input tensors is constructed and calculated using both the kernel and the PyTorch source code. Each element of the two output tensors is compared, and the maximum absolute error and relative error are calculated. If neither the maximum absolute error nor the relative error exceeds its respective threshold, the verification is successful. For kernels that fail the test, the Validator enters the repair process. This process additionally starts a sub-agent, KernelFixer, responsible for analyzing the cause of the error and providing repair suggestions. Then, the KernelGenerator regenerates the kernel based on these suggestions. KernelFixer can use prompt word templates for error analysis rather than code generation. Specifically, it can include the following: error message (maximum error of test failure, input / output shape); original prompt words (operator specifications, hardware constraints); generated code (Triton source code); and repair requirements (analyze the cause of the error and provide specific repair suggestions).

[0087] In this embodiment of the application, step S208 may specifically include the following steps e to g.

[0088] Step e: The process orchestration agent performs variable mapping on multiple inference units and concatenates the multiple inference units into inference process code according to the topological order in SSA IR.

[0089] The variable mapping described above solves the problem of converting SSA IDs to valid variable names. Values ​​in the SSA IR are referenced by SSA IDs, but meaningful variable names are required in the generated Python code. The rules for the variable mapping are as follows: convert SSA IDs (e.g., %123) to descriptive variable names (e.g., temp_123), maintain an ID remapping table to handle ID aliases caused by common subexpression elimination, and remap discrete IDs to consecutive IDs.

[0090] The above-described sequence enables scheduling and orchestration. Because the SSA form guarantees an acyclic dependency graph, topological ordering can be performed efficiently. The scheduler generates an execution plan, ensuring that all operands for each operator have been computed in previous steps.

[0091] Step f: Encapsulate the operator sequences in SSA IR that do not have dynamic control flow into a CUDA graph and record the CUDA graph.

[0092] Step g: Add playback logic to the inference flow code. The playback logic is used to select the recorded playback to be executed when the operator sequence is called.

[0093] CUDA Graph is a technology provided by NVIDIA that can record a series of CUDA kernel calls and replay them directly during subsequent execution, thereby reducing CPU scheduling overhead.

[0094] In this embodiment, the validator can be used throughout the entire pipeline, responsible for verifying the correctness of the output results at different stages, ensuring that the final generated inference process code is semantically equivalent to the original PyTorch model.

[0095] In one implementation, the above method may further include at least one of the following steps x to z.

[0096] Step x: Verify the agent performs static verification on the SSA IR, and revert the SSA IR that fails verification to the code analysis agent for regeneration.

[0097] The static verification of the SSA IR by the aforementioned verification agent includes format verification, which can specifically include three items: 1) Verify the validity of the opcode name to ensure that all operator names are in the predefined set of aten operators; 2) Verify the validity of operand references to ensure that all operand references to SSA IDs have been defined in the preceding text and that there are no forward reference errors; 3) Verify the self-consistency of shape inference by verifying whether the output shape of each operator is consistent with the expectation through the shape propagation rules of the computation graph.

[0098] SSA IRs that fail verification are returned with detailed error information, which the code analysis agent can then use a large language model to regenerate based on. This verification step ensures that the SSA IRs obtained in subsequent stages are complete and correctly formatted.

[0099] Step y: Verify the correctness of the numerical values ​​of multiple inference units by the intelligent agent. If the error exceeds the preset threshold, trigger the repair process.

[0100] The method for verifying the correctness of the aforementioned numerical values ​​may include: constructing multiple sets of random input tensors, covering different shape boundary conditions and data types, calculating them using both the inference unit and PyTorch source code, comparing each element of the two outputs, and calculating the maximum absolute error and relative error. If the maximum absolute error exceeds a preset threshold or the relative error exceeds a preset threshold (e.g., the relative error exceeds 1e-3), the inference unit is deemed to have failed to generate, triggering a repair process.

[0101] Step z: Verify the equivalence of the intelligent agent with the inference process code. Run the neural network model source code and the inference process code with the same input respectively, and verify the consistency of the output shape and data type, the cosine similarity of the output tensor, and the relative error of each element. If all verifications pass, output the inference process code.

[0102] If the output shape of the neural network model source code is consistent with the output shape of the inference process code, and the data type of the neural network model source code is consistent with the data type of the inference process code, then the first step of verification is successful.

[0103] If the cosine similarity between the output tensor of the neural network model source code and the output tensor of the inference process code is within the allowable range, then the second verification step is successful.

[0104] If the relative error of each element in the neural network model source code is within the allowable range as well as the relative error of each element in the inference process code, then the third step of the verification is successful.

[0105] Wherein, the foregoing equivalence verification may run the PyTorch source code and the generated inference flow code respectively using the same input data, and compare the output tensors of the two. In consideration of the minor differences that may be caused by the non-associativity of floating-point operations, the end-to-end verification adopts a hierarchical error evaluation strategy, which specifically includes the following steps: first, compare whether the output shapes and data types of the two are consistent; second, calculate the cosine similarity of the output tensors to ensure consistent directions; finally, calculate the element-wise relative error to ensure that it is within an acceptable numerical accuracy range. If the end-to-end verification is passed, the entire pipeline outputs the final inference code; if the verification fails, the problematic link is located according to the error analysis result, and the corresponding stage is re-executed after targeted repair.

[0106] The foregoing method provided in the embodiments of the present application obtains the source code of a neural network model through a code analysis agent, converts the neural network model source code into static single assignment intermediate representation SSA IR, a compilation optimization agent performs compilation optimization on the SSA IR, wherein the compilation optimization is used to reduce the number of operators and the video memory access to intermediate results without changing the computational semantics, a kernel generation agent generates a plurality of inference units according to the compiled and optimized SSA IR, and a flow orchestration agent combines the plurality of inference units into inference flow code, which realizes automatic generation of neural network model inference code. Compared with the hand-written operator mode, the present invention greatly shortens the development cycle, substantially reduces development costs, and can match the rapid iteration of neural network models. Moreover, by adopting a multi-agent collaborative architecture, the code understanding and generation capability of large language models is utilized to generate inference flow code, which improves the cross-architecture migration capability.

[0107] In addition, by decomposing the complex code conversion task into a plurality of sub-tasks with clear responsibilities, which are respectively processed by specialized agents, end-to-end automatic generation from PyTorch source code to high-performance Triton inference code is realized, the development cycle of neural network model inference code is shortened from several weeks to several hours, and the dependence on expert knowledge of GPU programming is substantially reduced.

[0108] The above is the automatic generation method for neural network model inference code provided by the embodiments of the present application. Based on the same idea, the embodiments of the present application also provide an automatic generation apparatus for neural network model inference code, as Figure 6 shown, the automatic generation apparatus for neural network model inference code comprises: a code analysis agent 601, configured to acquire neural network model source code, and convert the neural network model source code into static single assignment intermediate representation SSA IR.

[0109] a compilation optimization agent 602, configured to perform compilation optimization on the SSA IR, wherein the compilation optimization is used to reduce the number of operators and the video memory access to intermediate results without changing computational semantics.

[0110] The kernel generates agent 603, which generates multiple inference units based on the compiled and optimized SSA IR.

[0111] The process orchestration agent 604 combines multiple inference units into inference process code.

[0112] In this embodiment of the application, the code analysis agent 601 is specifically used to: convert the neural network model source code into a static single-assignment intermediate representation (SSAIR) according to the following principles: 1) Each SSA instruction corresponds to an ATen operator, and the operator granularity in the SSA IR is consistent with the kernel granularity; 2) All shape transformation and type conversion operations are represented by explicit commands; 3) Each variable is assigned a value once and referenced through its SSA ID, forming an acyclic computation dependency graph.

[0113] In this embodiment of the application, the compiler optimization agent 602 is specifically used to perform at least one of the following compiler optimizations on the SSA IR: weight preprocessing, constant folding, common subexpression elimination, dead code elimination, and operator fusion.

[0114] In this embodiment of the application, the weight preprocessing includes: advancing the transpose operation of the weights to the neural network model loading stage, and removing the operator corresponding to the transpose operation from the SSA IR.

[0115] In this embodiment of the application, constant folding includes: pre-compiling operators whose operands are all constants at compile time, storing the calculation results in the constant storage area, and removing operators whose operands are all constants from the SSA IR.

[0116] In this embodiment of the application, common subexpression elimination includes: caching the operator name, operands, and calculation result for each calculated subexpression; and reusing the corresponding calculation result in the cache when encountering an operator with the same operator name and operands as those in the cache during the traversal of the SSA IR.

[0117] In this embodiment of the application, dead code elimination includes: tracing back the referenced operators from the output of the neural network model through the SSA ID, and filtering out unreferenced operators in the SSA IR.

[0118] In this embodiment of the application, operator fusion includes: merging multiple adjacent operators that have direct data dependencies into a single operator, and performing calculations on multiple adjacent operators in a register.

[0119] In this embodiment of the application, the kernel-generated intelligent agent 603 is specifically used for: Based on the computational complexity of the operators, the operators in the compiled and optimized SSA IR are classified, and the large language model is used to convert each type of operator into a corresponding inference unit. Based on the same input, compare the source code of the neural network model and the inference results of each inference unit to determine the inference unit that failed the verification. Analyze the reasons for the failure of the inference unit and generate repair suggestions; The corresponding inference units are regenerated based on the repair recommendations.

[0120] In this embodiment of the application, the kernel-generated intelligent agent 603 is specifically used for: Based on the computational complexity of the operators, the operators in the compiled and optimized SSA IR are divided into computationally intensive operators, non-computational operators, and complex operators; For computationally intensive operators, a dedicated optimized kernel is generated using a large language model as the inference unit; For non-computational operators, the corresponding operators in the neural network model source code are used as inference units; For complex operators, a large language model is used to generate adapted code as an inference unit based on the target hardware characteristics, using a preset template.

[0121] In this embodiment of the application, the process orchestration agent 604 is specifically used for: Variable mapping is performed on multiple inference units, and the multiple inference units are chained together into inference flow code according to the topological order in SSA IR; The operator sequences that do not have dynamic control flow in SSA IR are encapsulated into CUDA graphs, and the CUDA graphs are recorded. Add playback logic to the inference flow code. The playback logic is used to select to execute the recorded playback when calling the operator sequence.

[0122] In this embodiment of the application, the above-described apparatus may further include a verification agent for performing at least one of the following: Perform static validation on the SSA IR, and revert any SSA IRs that fail validation to the code analysis agent for regeneration; Numerical correctness verification is performed on multiple inference units, and a repair process is triggered if the error exceeds a preset threshold. The inference process code is verified for equivalence. The neural network model source code and the inference process code are run separately using the same input. The consistency of output shape and data type, cosine similarity of output tensors, and relative error per element are verified respectively. If all verifications pass, the inference process code is output.

[0123] The apparatus provided in this application embodiment can execute the methods in the above method embodiments. For detailed process, please refer to the description in the method embodiments, which will not be repeated here.

[0124] The apparatus provided in this application obtains the source code of a neural network model through a code analysis agent, converts the source code into a static single-assignment intermediate representation (SSA IR), and a compilation optimization agent performs compilation optimization on the SSA IR. This compilation optimization reduces the number of operators and memory access to intermediate results without changing the computational semantics. A kernel generation agent generates multiple inference units based on the compiled and optimized SSA IR, and a process orchestration agent combines these multiple inference units into inference flow code. This achieves automatic generation of neural network model inference code, significantly shortening the development cycle and greatly reducing development costs compared to the hand-written operator mode, and is able to match the rapid iteration of neural network models. Moreover, by using a multi-agent collaborative architecture and leveraging the code understanding and generation capabilities of a large language model to generate inference flow code, the cross-architecture migration capability is improved.

[0125] Figure 7 This is a schematic diagram of the hardware structure of an electronic device to implement the various embodiments of this application. The electronic device 700 includes, but is not limited to: a radio frequency unit 701, a network module 702, an audio output unit 703, an input unit 704, a sensor 705, a display unit 706, a user input unit 707, an interface unit 708, a memory 709, a processor 710, and a power supply 711, etc. Those skilled in the art will understand that... Figure 7 The electronic device structures shown are not intended to limit the electronic device. An electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements. In the embodiments of this application, the electronic device includes, but is not limited to, mobile phones, tablets, laptops, PDAs, in-vehicle terminals, wearable devices, and pedometers.

[0126] The processor 710 is configured to: use a code analysis agent to obtain the source code of a neural network model; convert the source code of the neural network model into a static single-assignment intermediate representation (SSAIR); use a compilation optimization agent to perform compilation optimization on the SSAIR, wherein the compilation optimization is used to reduce the number of operators and the memory access of intermediate results without changing the computational semantics; use a kernel generation agent to generate multiple inference units based on the compiled and optimized SSAIR; and use a process orchestration agent to combine the multiple inference units into inference flow code.

[0127] This application provides an electronic device that obtains the source code of a neural network model through a code analysis agent, converts the source code into a static single-assignment intermediate representation (SSA IR), and a compilation optimization agent performs compilation optimization on the SSA IR. This optimization reduces the number of operators and memory access to intermediate results without changing the computational semantics. A kernel generation agent generates multiple inference units based on the compiled and optimized SSA IR, and a process orchestration agent combines these multiple inference units into inference flow code. This achieves automatic generation of neural network model inference code, significantly shortening the development cycle and reducing development costs compared to hand-written operator mode, and can match the rapid iteration of neural network models. Moreover, the use of a multi-agent collaborative architecture leverages the code understanding and generation capabilities of a large language model to generate inference flow code, improving cross-architecture migration capabilities.

[0128] It should be understood that, in this embodiment, the radio frequency unit 701 can be used for receiving and transmitting signals during information transmission or calls. Specifically, it receives downlink data from the base station and processes it with the processor 710; additionally, it transmits uplink data to the base station. Typically, the radio frequency unit 701 includes, but is not limited to, an antenna, at least one amplifier, a transceiver, a coupler, a low-noise amplifier, a duplexer, etc. Furthermore, the radio frequency unit 701 can also communicate with networks and other electronic devices via a wireless communication system.

[0129] Electronic devices provide users with wireless broadband internet access through network module 702, such as helping users send and receive emails, browse web pages, and access streaming media.

[0130] The audio output unit 703 can convert audio data received by the radio frequency unit 701 or the network module 702 or stored in the memory 709 into audio signals and output them as sound. Furthermore, the audio output unit 703 can also provide audio output related to specific functions performed by the electronic device 700 (e.g., call signal reception sound, message reception sound, etc.). The audio output unit 703 includes a speaker, a buzzer, and a receiver, etc.

[0131] Input unit 704 is used to receive audio or video signals. Input unit 704 may include a graphics processing unit (GPU) 7041 and a microphone 7042. The GPU 7041 processes image data of still images or videos acquired by an image capture device (such as a camera) in video capture mode or image capture mode. The processed image frames can be displayed on display unit 706. The image frames processed by GPU 7041 can be stored in memory 709 (or other storage media) or transmitted via radio frequency unit 701 or network module 702. Microphone 7042 can receive sound and process such sound into audio data. The processed audio data can be converted into a format that can be transmitted to a mobile communication base station via radio frequency unit 701 in telephone call mode.

[0132] The electronic device 700 also includes at least one sensor 705, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor includes an ambient light sensor and a proximity sensor. The ambient light sensor can adjust the brightness of the display panel 7061 according to the ambient light level, and the proximity sensor can turn off the display panel 7061 and / or backlight when the electronic device 700 is moved to the ear. As a type of motion sensor, an accelerometer sensor can detect the magnitude of acceleration in various directions (generally three axes). When stationary, it can detect the magnitude and direction of gravity and can be used to identify the posture of the electronic device (such as landscape / portrait switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc. The sensor 705 may also include a fingerprint sensor, pressure sensor, iris sensor, molecular sensor, gyroscope, barometer, hygrometer, thermometer, infrared sensor, etc., which will not be described in detail here.

[0133] The display unit 706 is used to display information input by the user or information provided to the user. The display unit 706 may include a display panel 7061, which may be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), or the like.

[0134] User input unit 707 can be used to receive input numerical or character information, and to generate key signal inputs related to user settings and function control of electronic devices. Specifically, user input unit 707 includes a touch panel 7071 and other input devices 7072. Touch panel 7071, also known as a touch screen, can collect touch operations performed by the user on or near it (such as operations performed by the user using a finger, stylus, or any suitable object or accessory on or near touch panel 7071). Touch panel 7071 may include two parts: a touch detection device and a touch controller. The touch detection device detects the user's touch position and the signal generated by the touch operation, and transmits the signal to the touch controller; the touch controller receives touch information from the touch detection device, converts it into touch point coordinates, and sends it to the processor 710, which receives and executes commands from the processor 710. In addition, touch panel 7071 can be implemented using various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to touch panel 7071, user input unit 707 may also include other input devices 7072. Specifically, other input devices 7072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, joysticks, etc., which will not be described in detail here.

[0135] Furthermore, the touch panel 7071 can cover the display panel 7061. When the touch panel 7071 detects a touch operation on or near it, it transmits the information to the processor 710 to determine the type of touch event. Subsequently, the processor 710 provides corresponding visual output on the display panel 7061 based on the type of touch event. Although in Figure 7 In this embodiment, the touch panel 7071 and the display panel 7061 are two independent components to realize the input and output functions of the electronic device. However, in some embodiments, the touch panel 7071 and the display panel 7061 can be integrated to realize the input and output functions of the electronic device. The specific implementation is not limited here.

[0136] Interface unit 708 serves as an interface for connecting external devices to electronic device 700. For example, external devices may include a wired or wireless headphone port, an external power supply (or battery charger) port, a wired or wireless data port, a memory card port, a port for connecting a device with an identification module, an audio input / output (I / O) port, a video I / O port, a headphone port, and so on. Interface unit 708 can be used to receive input from external devices (e.g., data, power, etc.) and transmit the received input to one or more components within electronic device 700, or it can be used to transmit data between electronic device 700 and external devices.

[0137] The memory 709 can be used to store software programs and various data. The memory 709 may primarily include a program storage area and a data storage area. The program storage area may store the operating system, applications required for at least one function (such as sound playback, image playback, etc.), etc.; the data storage area may store data created based on the use of the mobile phone (such as audio data, phonebook, etc.). Furthermore, the memory 709 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.

[0138] The processor 710 is the control center of the electronic device. It connects various parts of the electronic device via various interfaces and lines. By running or executing software programs and / or modules stored in the memory 709, and by calling data stored in the memory 709, it performs various functions and processes data, thereby providing overall monitoring of the electronic device. The processor 710 may include one or more processing units; preferably, the processor 710 may integrate an application processor and a modem processor. The application processor mainly handles the operating system, user interface, and applications, while the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 710.

[0139] The electronic device 700 may also include a power supply 711 (such as a battery) that supplies power to various components. Preferably, the power supply 711 is logically connected to the processor 710 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system.

[0140] Preferably, this application embodiment also provides an electronic device, including a processor 710, a memory 709, and a computer program stored in the memory 709 and executable on the processor 710. When the computer program is executed by the processor 710, it implements the various processes of the above method embodiments and can achieve the same technical effects. To avoid repetition, it will not be described again here.

[0141] This application also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the various processes of the above method embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0142] This application also provides a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the various processes of the above-described method embodiments and achieves the same technical effects. To avoid repetition, it will not be described again here. The computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0143] The computer-readable storage medium provided in this application embodiment obtains the source code of a neural network model through a code analysis agent, converts the source code into a static single-assignment intermediate representation (SSA IR), and a compilation optimization agent performs compilation optimization on the SSA IR. This compilation optimization reduces the number of operators and the memory access of intermediate results without changing the computational semantics. A kernel generation agent generates multiple inference units based on the compiled and optimized SSA IR, and a process orchestration agent combines these multiple inference units into inference flow code. This achieves automatic generation of neural network model inference code, which greatly shortens the development cycle and significantly reduces development costs compared to the hand-written operator mode, and can match the rapid iteration of neural network models. Moreover, by using a multi-agent collaborative architecture and leveraging the code understanding and generation capabilities of a large language model to generate inference flow code, the cross-architecture migration capability is improved.

[0144] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0145] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0146] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0147] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0148] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0149] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0150] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information by any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0151] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0152] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A method for automatically generating inference code for a neural network model, characterized in that, The method includes: The code analysis agent obtains the source code of the neural network model and converts the source code of the neural network model into a static single-assignment intermediate representation (SSA IR). The compiler optimization agent performs compiler optimization on the SSA IR, which is used to reduce the number of operators and memory access of intermediate results without changing the computational semantics; The kernel generates an intelligent agent that generates multiple inference units based on the compiled and optimized SSA IR; The process orchestration agent combines the multiple inference units into inference process code.

2. The method according to claim 1, characterized in that, The step of converting the neural network model source code into a static single-assignment intermediate representation (SSAIR) includes: The code analysis agent converts the neural network model source code into a static single-assignment intermediate representation (SSAIR) according to the following principles: Each SSA instruction corresponds to an ATen operator, and the granularity of the operators in the SSA IR is consistent with the kernel granularity. All shape transformation and type conversion operations are represented by explicit instructions; Each variable is assigned a value once and referenced via its SSA ID, forming an acyclic computational dependency graph.

3. The method according to claim 1, characterized in that, The compilation optimization agent performs compilation optimization on the SSA IR, including: The compilation optimization agent performs at least one of the following compilation optimizations on the SSA IR: Weight preprocessing, constant folding, common subexpression elimination, dead code elimination, and operator fusion.

4. The method according to claim 3, characterized in that, The weight preprocessing includes: performing the weight transpose operation in advance to the neural network model loading stage, and removing the operator corresponding to the transpose operation from the SSA IR; The constant folding includes: pre-compiling operators whose operands are all constants at compile time, storing the calculation results in the constant storage area, and removing the operators whose operands are all constants from the SSA IR; The elimination of common subexpressions includes: caching the operator name, operands, and calculation result for each calculated subexpression; and reusing the corresponding calculation result in the cache when encountering an operator with the same operator name and operands as those in the cache during the traversal of the SSA IR. The dead code elimination includes: tracing back the referenced operators from the output of the neural network model through the SSA ID, and filtering out unreferenced operators in the SSA IR; The operator fusion includes: merging multiple adjacent operators that have direct data dependencies into a single operator, and performing the calculation of the multiple adjacent operators in a register.

5. The method according to claim 1, characterized in that, The kernel-generated agent generates multiple inference units based on the compiled and optimized SSA IR, including: The kernel-generated intelligent agent classifies the operators in the compiled and optimized SSA IR according to the computational complexity of the operators, and uses a large language model to convert each type of operator into a corresponding inference unit. Based on the same input, the source code of the neural network model and the inference results of each inference unit are compared to determine the inference unit that failed the verification. Analyze the error causes of the failed inference units and generate repair suggestions; The corresponding inference unit is regenerated based on the repair recommendations.

6. The method according to claim 5, characterized in that, The kernel-generated agent classifies the operators in the compiled and optimized SSA IR according to their computational complexity, and uses a large language model to convert each type of operator into a corresponding inference unit, including: The kernel-generated intelligent agent classifies the operators in the compiled and optimized SSA IR into computationally intensive operators, non-computationally intensive operators, and complex operators based on the computational complexity of the operators. For the computationally intensive operator, a dedicated optimized kernel is generated using a large language model as the inference unit; For the non-computational operators, the corresponding operators in the source code of the neural network model are used as inference units; For the complex operator, a large language model is used to generate adapted code as an inference unit based on the target hardware characteristics, using a preset template.

7. The method according to claim 1, characterized in that, The process orchestration agent combines the multiple inference units into inference process code, including: The process orchestration agent performs variable mapping on the multiple inference units and strings the multiple inference units together into inference process code according to the topological order in the SSA IR; The operator sequences that do not have dynamic control flow in the SSA IR are encapsulated into a CUDA graph, and the CUDA graph is recorded. Add playback logic to the inference process code. The playback logic is used to select to execute the recorded playback when the operator sequence is called.

8. The method according to claim 1, characterized in that, The method further includes at least one of the following: The verification agent performs static verification on the SSA IR, and returns the SSA IR that fails verification to the code analysis agent for regeneration; The verification agent verifies the numerical correctness of the multiple inference units, and if the error exceeds a preset threshold, a repair process is triggered. The verification agent performs equivalence verification on the inference process code. Using the same input, it runs the neural network model source code and the inference process code respectively, and verifies the consistency of output shape and data type, cosine similarity of output tensors, and relative error per element. If all verifications pass, the inference process code is output.

9. An automatic generation device for neural network model inference code, characterized in that, The device includes: A code analysis agent is used to obtain the source code of a neural network model and convert the source code of the neural network model into a static single-assignment intermediate representation (SSAIR). The compilation optimization agent performs compilation optimization on the SSA IR, which is used to reduce the number of operators and memory access of intermediate results without changing the computational semantics. The kernel generates an intelligent agent, which generates multiple inference units based on the compiled and optimized SSA IR; The process orchestration agent combines the multiple inference units into inference process code.

10. An electronic device, characterized in that, The method includes a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the method for automatically generating inference code for a neural network model as described in any one of claims 1 to 8.