An RWKV-7 network transplantation and optimization method, system, terminal and medium for Ascend NPU

CN122433803BActive Publication Date: 2026-09-08GUANGDONG LAB OF ARTIFICIAL INTELLIGENCE & DIGITAL ECONOMY (SZ)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610874558.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-17
Publication Date
2026-09-08
Estimated Expiration
2046-06-17

AI Technical Summary

Technical Problem

[0006]本申请的主要目的在于提供一种面向昇腾NPU的RWKV-7网络移植与优化方法、系统、终端及介质,旨在解决现有技术中RWKV网络的实现基于某厂商的GPU及其CUDA并行计算平台,且RWKV核心算子的计算特性与昇腾的硬件特性不匹配,导致计算资源利用率较低的问题

Benefits of technology

[0017] Beneficial Effects: This application provides a method, system, terminal, and medium for porting and optimizing RWKV-7 networks for Ascend NPUs. After analyzing and identifying performance bottlenecks, this application develops dedicated hardware operators to replace these bottlenecks. It employs a progressive combination of network-level and operator-level optimizations, and by comparing the performance after each step, the benefits of each optimization action can be clearly quantified, facilitating analysis and localization. Based on the Ascend AI chip, CANN heterogeneous computing architecture, and MindSpore deep learning framework, this application enables the complete inference application of the RWKV-7 language model on the target region's AI platform, thereby improving the utilization rate of computing resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122433803B_ABST
    Figure CN122433803B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of artificial intelligence, and discloses an RWKV-7 network transplantation and optimization method, system, terminal and medium for Ascend NPU. The method comprises the following steps: constructing an original RWKV-7 model on a set computing platform; performing performance analysis and bottleneck positioning on the original RWKV-7 model by using a performance analysis tool of the set computing platform, outputting bottleneck root cause conclusions and an optimization scheme of the original RWKV-7 model; performing hierarchical optimization on the original RWKV-7 model according to the bottleneck root cause conclusions and the optimization scheme, obtaining an optimized high-performance RWKV-7 model; and performing integration and end-to-end verification according to the original RWKV-7 model, the bottleneck root cause conclusions, the optimization scheme and the high-performance RWKV-7 model, obtaining a target optimized RWKV-7 model for executing a natural language processing inference task on the Ascend NPU. The application can improve the utilization rate of computing resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a method, system, terminal and medium for porting and optimizing RWKV-7 network for Ascend NPU. Background Technology

[0002] With the continuous growth in the parameter scale of deep learning models, model inference efficiency in generative artificial intelligence and natural language processing has become a key technical challenge. Although the traditional Transformer architecture has achieved great success in the field of large language models (LLM), its self-attention mechanism has a time complexity of O(n²), which is computationally expensive in long sequence inference scenarios, severely restricting the deployment efficiency and application scope of the model.

[0003] Linear attention, a novel attention computation method, reduces time complexity from O(n²) to O(n) by changing the order of attention calculations, achieving linear inference complexity similar to RNNs while maintaining the parallel training capability of Transformers. The RWKV (Receptance Weighted Key Value) network is an innovative architecture designed based on the principle of linear attention. It combines the advantages of recurrent neural networks (RNNs) and Transformers, supporting parallel training while maintaining constant computational complexity and memory usage during the inference phase.

[0004] Currently, the official implementation of the RWKV network is mainly based on the PyTorch deep learning framework and the Triton programming language, optimized for a specific vendor's GPU and CUDA platform. Existing solutions suffer from the following problems: First, strong hardware platform dependence: The existing implementation is highly dependent on a specific vendor's GPU and CUDA ecosystem, making it difficult to directly migrate to the target region's computing power platform; Second, lack of NPU adaptation solutions for the target region: Ascend NPU adopts the Da Vinci architecture and heterogeneous computing paradigm, which differs significantly from the GPU architecture, making existing code unable to run directly; Third, low operator implementation efficiency: The computational characteristics of the RWKV core operator (Time-Mixing) are incompatible with the hardware characteristics of the Ascend AI Core, resulting in low utilization of computing resources; Fourth, unoptimized memory management: It has not been optimized for the Ascend NPU's multi-level storage architecture (HBM2e, L2 Cache, Unified Buffer), resulting in high data transfer overhead.

[0005] Therefore, existing technologies still need to be improved and developed. Summary of the Invention

[0006] The main purpose of this application is to provide a method, system, terminal and medium for porting and optimizing RWKV-7 networks for Ascend NPU, aiming to solve the problem that the implementation of RWKV networks in the prior art is based on a certain manufacturer's GPU and its CUDA parallel computing platform, and the computing characteristics of RWKV core operators are not compatible with the hardware characteristics of Ascend, resulting in low utilization of computing resources.

[0007] The first aspect of this application provides a method for porting and optimizing RWKV-7 networks for Ascend NPUs, the method comprising the following steps: The original RWKV-7 model was built on the specified computing power platform; The original performance logs of the original RWKV-7 model are obtained by setting the performance analysis tool of the computing power platform. The computation graph structure is analyzed based on the original performance logs to obtain a list of inefficient nodes. The operator-level time consumption statistics are performed based on the original performance logs to obtain a hot operator ranking table. The list of inefficient nodes and the hot operator ranking table are cross-attribution analysis to obtain the bottleneck root cause conclusion. The priority ranking of optimization tasks is determined based on the bottleneck root cause conclusion. Based on the bottleneck root cause conclusions and the optimization task list, the original RWKV-7 model is optimized by core operators to obtain high-performance custom operators. Based on the high-performance custom operators and the original RWKV-7 model, memory access optimization is performed to obtain a memory optimization strategy. Based on the computation graph of the original RWKV-7 model and the high-performance custom operators, graph execution optimization is performed to obtain a target optimization computation graph. Based on the high-performance custom operators, the memory optimization strategy, and the target optimization computation graph, a high-performance RWKV-7 model is obtained. The original RWKV-7 model, the bottleneck root cause conclusions, the optimization task list, and the high-performance RWKV-7 model are integrated and validated end-to-end to obtain the target optimization RWKV-7 model. A natural language processing inference task is obtained, and the natural language processing inference task is input into the target optimization RWKV-7 model. After the target optimization RWKV-7 model executes the natural language processing inference task on the Ascend NPU, it outputs the task execution result.

[0008] Optionally, in one embodiment of this application, the step of constructing the original RWKV-7 model on a set computing power platform specifically includes: The MindSpore model is obtained by reconstructing the network structure using the MindSpore deep learning framework in the computing platform. Obtain the checkpoint file that matches the MindSpore model and the operator implementation fragment corresponding to the MindSpore model; Functional verification was performed based on the MindSpore model, the checkpoint file, and the operator implementation fragment to obtain the original RWKV-7 model.

[0009] Optionally, in one embodiment of this application, the high-performance custom operator is the Ascend C WKV7 operator; The process of optimizing the original RWKV-7 model based on the bottleneck root cause conclusions and the optimization task list to obtain high-performance custom operators specifically includes: Based on the bottleneck root cause conclusions and the optimization task list, the original RWKV-7 model is decomposed into state update formulas, and an instruction sequence blueprint is output. Design a Tiling strategy based on the instruction sequence blueprint and output a Tiling parameter table; Based on the instruction sequence blueprint and the Tiling parameter table, a pipelined parallel design is performed to generate the Ascend CWKV7 operator.

[0010] Optionally, in one embodiment of this application, the memory optimization strategy includes a memory reuse mapping table and double-buffered scheduling logic; The memory optimization strategy obtained by optimizing memory access based on the high-performance custom operator and the original RWKV-7 model specifically includes: Based on the high-performance custom operator and the original RWKV-7 model, local tensor pre-allocation is performed, and a pre-allocated memory pool is output. Based on the computation graph of the original RWKV-7 model and the pre-allocated memory pool, memory reuse is performed, a memory reuse mapping table is output, and the memory reuse mapping table is integrated into the MindSpore memory manager; Based on the operator data access mode of the Ascend C WKV7 operator, output the data layout conversion code; Based on the Tiling parameter and the pipeline design of the Ascend C WKV7 operator, a double-buffered scheduling logic is output and integrated into the Ascend C WKV7 operator.

[0011] Optionally, in one embodiment of this application, the step of obtaining the target optimization computation graph by performing graph execution optimization based on the computation graph of the original RWKV-7 model and the high-performance custom operator specifically includes: Based on the computation graph of the original RWKV-7 model and the high-performance custom operator, operator fusion is performed to output a fused computation graph; Redundancy is eliminated from the fused computation graph, and a simplified computation graph is output. Based on the simplified computation graph and the static configuration of the original RWKV-7 model, constant folding is performed, and a constant folding computation graph is output. Custom operator integration is performed based on the constant folding computation graph and the high-performance custom operator to output the target optimization computation graph.

[0012] Optionally, in one embodiment of this application, the computing platform includes Ascend NPU, CANN heterogeneous computing architecture, and MindSpore deep learning framework.

[0013] Optionally, in one embodiment of this application, the step of integrating and end-to-end verifying the original RWKV-7 model, the bottleneck root cause conclusions, the optimization task list, and the high-performance RWKV-7 model to obtain the target optimization RWKV-7 model specifically includes: The high-performance custom operator, the memory optimization strategy, and the target optimization computation graph in the high-performance RWKV-7 model are integrated into the MindSpore computation graph to obtain the integrated model. The integrated model is compared with the original RWKV-7 model for correctness verification. Based on the bottleneck root cause conclusion and the optimization task list, different batch sizes and sequence lengths are set to verify the performance of the integrated model. The integrated model that has passed the correctness verification and the performance verification is used as the target to optimize the RWKV-7 model.

[0014] A second aspect of this application also provides a system for porting and optimizing an RWKV-7 network for Ascend NPU, wherein the system is used to implement the method for porting and optimizing an RWKV-7 network for Ascend NPU as described in any of the above solutions; the system includes: The model building module is used to build the original RWKV-7 model on a specified computing power platform. The performance analysis and bottleneck location module is used to obtain the original performance logs of the original RWKV-7 model by setting the performance analysis tools of the computing power platform, perform computation graph structure analysis based on the original performance logs to obtain a list of inefficient nodes, perform operator-level time consumption statistics based on the original performance logs to obtain a hot operator ranking table, perform cross-attribution analysis on the list of inefficient nodes and the hot operator ranking table to obtain the bottleneck root cause conclusion, and determine the priority ranking of optimization tasks based on the bottleneck root cause conclusion. The hierarchical optimization and performance acceleration module is used to optimize the core operators of the original RWKV-7 model based on the bottleneck root cause conclusions and the optimization task list to obtain high-performance custom operators; to optimize memory access based on the high-performance custom operators and the original RWKV-7 model to obtain a memory optimization strategy; to optimize graph execution based on the computation graph of the original RWKV-7 model and the high-performance custom operators to obtain a target optimization computation graph; and to obtain a high-performance RWKV-7 model based on the high-performance custom operators, the memory optimization strategy, and the target optimization computation graph. The graph execution optimization module is used to integrate and perform end-to-end verification based on the original RWKV-7 model, the bottleneck root cause conclusion, the optimization task list and the high-performance RWKV-7 model to obtain the target optimized RWKV-7 model. The model execution task module is used to acquire a natural language processing inference task, input the natural language processing inference task into the target optimization RWKV-7 model, and after the target optimization RWKV-7 model executes the natural language processing inference task on the Ascend NPU, it outputs the task execution result.

[0015] A third aspect of this application also provides a terminal, wherein the terminal includes: a memory, a processor, and an RWKV-7 network porting and optimization program for Ascend NPU stored in the memory and executable on the processor, wherein when the RWKV-7 network porting and optimization program for Ascend NPU is executed by the processor, it implements the steps of the RWKV-7 network porting and optimization method for Ascend NPU as described above.

[0016] A fourth aspect of this application also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a program for porting and optimizing an RWKV-7 network for an Ascend NPU, and when the program is executed by a processor, it implements the steps of the method for porting and optimizing an RWKV-7 network for an Ascend NPU as described above.

[0017] Beneficial Effects: This application provides a method, system, terminal, and medium for porting and optimizing RWKV-7 networks for Ascend NPUs. After analyzing and identifying performance bottlenecks, this application develops dedicated hardware operators to replace these bottlenecks. It employs a progressive combination of network-level and operator-level optimizations, and by comparing the performance after each step, the benefits of each optimization action can be clearly quantified, facilitating analysis and localization. Based on the Ascend AI chip, CANN heterogeneous computing architecture, and MindSpore deep learning framework, this application enables the complete inference application of the RWKV-7 language model on the target region's AI platform, thereby improving the utilization rate of computing resources. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a schematic diagram of the overall architecture of the RWKV-7 network in this application on the Ascend NPU; Figure 2 This is a flowchart of a preferred embodiment of the RWKV-7 network porting and optimization method for Ascend NPU in this application; Figure 3 This is a schematic diagram of the RWKV7 architecture of this application; Figure 4 This is a computation flow graph of the Time-Mixing module in a preferred embodiment of the RWKV-7 network porting and optimization method for Ascend NPU in this application; Figure 5 This is a schematic diagram of the Ascend 910B storage architecture and data flow in a preferred embodiment of the RWKV-7 network porting and optimization method for Ascend NPU in this application; Figure 6 The figure below is a schematic diagram of the fusion optimization scheme in a preferred embodiment of the RWKV-7 network porting and optimization method for Ascend NPU in this application; Figure 7 This is a performance comparison chart of the rwkv7 model before and after optimization in a preferred embodiment of the RWKV-7 network porting and optimization method for Ascend NPU in this application. Figure 8 This is a structural diagram of a preferred embodiment of the RWKV-7 network porting and optimization system for Ascend NPU according to this application; Figure 9 This is a structural diagram of a preferred embodiment of the terminal of this application.

[0020] Explanation of reference numerals in the attached figures: 100. Model building module; 200. Performance analysis and bottleneck identification module; 300. Hierarchical optimization and performance acceleration module; 400. Graph execution optimization module; 500. Model execution task module. Detailed Implementation

[0021] To make the objectives, technical solutions, and effects of this application clearer and more explicit, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. The described embodiments are only possible technical implementations of this application and not all possible implementations. Based on the embodiments in this application, those skilled in the art can obtain other embodiments without creative effort, and these embodiments are also within the protection scope of this application.

[0022] First, let's introduce the terms used in the embodiments of this application: AI Core, Artificial Intelligence Core (AI computing core). AIC, AI Cube, Matrix Computing Unit (Cube Core); AIV, AI Vector, Vector Computation Unit (Vector Core). Ascend C, Ascend C (programming language name), the Ascend C programming language; Attention, the mechanism of attention; BF16, Brain Floating Point 16-bit, is a 16-bit data type. CANN, Compute Architecture for Neural Networks (Ascend Computing Architecture); Channel-Mixing, Channel-Mixing module; CUDA, Compute Unified Device Architecture; Decay, Decay, attenuation (coefficient); GM stands for Global Memory (HBM memory). GPU, Graphics Processing Unit; HBM, High Bandwidth Memory; HBM2e, High Bandwidth Memory 2e, is a second-generation enhanced high-bandwidth memory. INT88-bit Integer, an 8-bit integer data type; Key, Key (K in the attention mechanism); L1 Buffe, Level 1 Buffe, is a first-level buffer. L2 Cache, Level 2 Cache; Linear Attention; Linear attention. LLM, Large Language Model; Mamba, Mamba (model name), Mamba architecture (a linear attention model); MindSpore, MindSpore (deep learning framework), Ascend deep learning framework; MindSpore Profiler, a MindSpore performance analysis tool; Naive WKV7 Kernel; ND / NZ, N-Dimensional / NZ format. ND format (general multidimensional format) / NZ format (Ascend-specific data layout); NPU, Neural Processing Unit (AI chip). PyTorch, the PyTorch deep learning framework; Receptance (gating signal in RWKV); RetNet, Retention Network, RetNet network (a linear attention architecture). RNN, Recurrent Neural Network; RWKV, Receptance Weighted Key Value (RWKV network). Softmax, a flexible maximum value function; SPMD, Single Program Multiple Data, is a parallel mode for single-program multiple data. Time-Mixing, a time-mixing module; Tiling, data chunking (chunking strategy); Transformer, Transformer, Transformer architecture; Triton (a programming language) is used for writing GPU cores. UB, Unified Buffer (on-chip cache); Value (V in the attention mechanism); WKV, Weighted Key Value (RWKV core calculation mode).

[0023] The core idea of ​​linear attention is to linearize the softmax operation in attention computation using a kernel function. Traditional attention computation is as follows: Attention(Q,K,V) = softmax(QK V; Linear attention uses a feature mapping function φ(·) to map Q and K to a new feature space, so that the attention calculation can be rewritten as: Attention(Q,K,V) = φ(Q)(φ(K) V); Due to the associative law of matrix multiplication, φ(K) can be calculated first. V (time complexity O(n)) is then multiplied by φ(Q), thereby reducing the overall complexity from O(n²) to O(n). The RWKV network further introduces mechanisms such as receptance, key, value, and decay, forming a unique Time-Mixing computation model.

[0024] This application presents a method for porting and optimizing the RWKV-7 network for the Ascend AI platform. Based on the Huawei Ascend AI chip, CANN heterogeneous computing architecture, and MindSpore deep learning framework, it enables the complete inference application of the RWKV-7 language model on the target region's AI platform. The method utilizes MindSpore for model migration and framework adaptation to achieve complete porting of the RWKV-7 network structure; it employs a custom operator design based on the Ascend C Time-Mixing module, deeply optimizing the Cube and Vector computing units of the Ascend AI Core; it incorporates a multi-level storage architecture-aware memory optimization strategy, including local tensor pre-allocation and memory reuse mechanisms; and it streamlines and optimizes graph execution paths to reduce data copying overhead between the framework layer and the operator layer.

[0025] It should be noted that this application focuses on the RWKV7 language model inference task, and is divided into three parts: computational graph reproduction, network optimization, and operator design. The computational graph reproduction uses the dynamic graph mode of the MindSpore deep learning framework. By examining existing RWKV-7 inference networks implemented using PyTorch and Triton frameworks, equivalent interfaces are used to complete the parsing of RWKV7-World series model weight files, the forward inference network, and the initialization of network weights. For computational steps using CUDA operators in existing implementations, the computational graph reproduction will transcribe them into matrix operation forms suitable for graph operations, avoiding the negative impact of unreasonable design on the reliability of performance evaluation data. After the computational graph reproduction, the MindSpore framework can independently complete the inference function. After the computational graph reproduction is completed, the model's inference function is verified. The expected result is that the text prediction task can be completed normally, that is, it can compute and infer subsequent sequences and output them for the input sequence. Network optimization is carried out after the functionality of the computational graph is verified. Based on the MindSpore Profler results and the computational graph structure, an overall analysis is performed to locate performance bottlenecks and customize a network optimization scheme. The network optimization scheme follows the principle of ensuring functionality, that is, the adjustment of the network cannot change the computation results of any process. Otherwise, the network will not be able to obtain reasonable final inference results, which is mainly manifested in the destruction of the semantic coherence of the output sequence.

[0026] The operator design adopts the PyTorch implementation of the Naive WKV7 Kernel proposed by RWKV-7 for the temporal mixing layer. It also references existing CUDA implementations and fully considers the characteristics of Ascend hardware, making certain improvements to allocate the Tiling strategy as rationally as possible and fully utilize the task pipeline mechanism. The operator design undergoes both accuracy and functional verification. Accuracy verification uses the same tensor and performs calculations using both the custom operator and the implementation based on the same principles of the PyTorch framework. The results are then compared to ensure the difference is within an acceptable range. Functional verification ensures that the operator's registration and integration into the network do not affect the network's final inference results.

[0027] For different network optimization schemes, this paper tests the relative optimization effect of each stack on top of the previous one, according to the stacking order of the optimization schemes. Finally, the relative optimization effect is tested by stacking a custom operator. Ultimately, for the complete design, different batch sizes and sequence lengths are set to verify both the overall performance of the complete design on the Huawei platform and the correctness of the overall design, i.e., correctly demonstrating the constant speed characteristic of the RWKV network inference process.

[0028] Following the above line of thought, see Figure 1The overall structure of the solution is divided into a model layer, a framework layer, an operator layer, a memory and scheduling layer, and a hardware execution layer. It adopts a three-level collaborative optimization architecture of framework-operator-hardware, with the overall architecture divided into five layers: Model layer (RWKV-7 Model): including core modules such as word vector embedding, Time-Mixing, Channel-Mixing, and word lookup; Framework layer (MindSpore Graph / Runtime): responsible for computation graph construction, operator scheduling, and memory management; Operator layer (AI Core Kernel): custom WKV7 operators implemented based on Ascend C; Memory and scheduling layer (UB / L1 / L2): managing multi-level storage such as Unified Buffer, L1 Buffer, and L2 Cache; Hardware execution layer (Ascend NPU): the AI ​​Core execution unit of the Ascend 910B chip.

[0029] The technical solutions of this application will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0030] The preferred embodiment of this application describes the RWKV-7 network porting and optimization method for Ascend NPU, such as... Figure 2 As shown, the method for porting and optimizing the RWKV-7 network for Ascend NPU includes the following steps: In step S10, the original RWKV-7 model is constructed on the set computing power platform.

[0031] In one possible implementation, step S10 specifically includes: reconstructing the network structure by setting the MindSpore deep learning framework in the computing platform to obtain the MindSpore model; obtaining the checkpoint file matching the MindSpore model and the operator implementation fragment corresponding to the MindSpore model; and performing functional verification based on the MindSpore model, the checkpoint file, and the operator implementation fragment to obtain the original RWKV-7 model.

[0032] This application first completes the migration of the RWKV-7 network from PyTorch to MindSpore. For example... Figure 3 As shown, the migration process includes: Network structure reconstruction: Reimplementing the RWKV-7 layer structure using MindSpore's nn.Module interface, maintaining the same computational logic as the original model; Weight format conversion: Converting the PyTorch format model weight file to a MindSpore-compatible format to ensure correct parameter loading; Data flow alignment: Unifying the input and output data formats and dimension arrangement to ensure semantic consistency with the original implementation. Figure 3 In the diagram, x represents the main input feature vector of the current time step or the current layer, and x0 represents the input feature tensor of the 0th time step. This indicates intermediate features after time mixing or channel mixing (the subscript t indicates the time step). This represents the output feature tensor after passing through the time-mixing or channel-mixing module at time step 0, and wkv represents the weighted key-value state (i.e., wkv) calculated by the Time-Mixing core. t wkv0 represents the weighted key-value state at time step 0.

[0033] Specifically, on the Ascend NPU and MindSpore framework, a fully functional but unoptimized RWKV-7 model was built from scratch. This included: Step S11 involves network structure reconstruction. Each layer of RWKV-7 (Embedding, Time-Mixing, Channel-Mixing, output layer, etc.) is reimplemented using MindSpore's nn.Module API, producing an empty, structurally complete MindSpore model instance. This model has the exact same layer structure, parameter shapes, and connections as the original PyTorch model, but all parameters are not yet loaded (weights are random values ​​or empty), and some complex operators (such as the original WKV7 operation in Time-Mixing) are not yet implemented. A mapping table of layer names and parameter shapes is provided for the weight transformation step S12. Simultaneously, the location of operator call points and input / output specifications are provided for the operator softening step S13.

[0034] Step S12: Weight format conversion. Read the PyTorch pre-trained weight file (.pth or .bin). Based on the layer name mapping relationship determined in step S11, reorganize the parameters and save them as a MindSpore checkpoint file (.ckpt). Output: A weight file (.ckpt) that perfectly matches the MindSpore model structure, which can be directly loaded by MindSpore.

[0035] Step S13, operator softening, identifies operators in the original PyTorch / Triton implementation that cannot run directly on the Ascend NPU (mainly complex WKV7 operations within Time-Mixing, certain custom CUDA cores, etc.), decomposes these operators into a combination of basic operations natively supported by MindSpore, and produces a functionally equivalent but low-performance operator implementation fragment, usually existing in the form of MindSpore's nn.Cell or ops.Composite, used to replace the custom operators in the original code. This implementation is inserted into the network structure reconstructed in step S11, so that the entire computation graph is entirely composed of operators supported by Ascend.

[0036] Step S14: Perform functional verification by integrating the reconstructed model instance, the converted weight file, and the softened operator implementation. Load the model and weights during MindSpore runtime, perform a complete forward inference, and check if the output is reasonable. If successful, output a functionally correct but poorly performing RWKV-7 model; if it fails, generate an error report.

[0037] In step S20, the original performance logs of the original RWKV-7 model are obtained by setting the performance analysis tool of the computing platform. The computational graph structure is analyzed based on the original performance logs to obtain a list of inefficient nodes. The operator-level time consumption statistics are performed based on the original performance logs to obtain a hotspot operator ranking table. The list of inefficient nodes and the hotspot operator ranking table are cross-attribution analysis to obtain the bottleneck root cause conclusion. The priority ranking of optimization tasks is determined based on the bottleneck root cause conclusion.

[0038] The goal of this step is to use performance analysis tools, based on a functionally correct model, to accurately identify the key factors that restrict inference speed, and to provide a quantitative basis for subsequent optimization.

[0039] Specifically, this includes: Step S21, performance data collection, running a functional but poorly performing RWKV-7 model, and simultaneously launching the MindSpore Profiler tool to generate raw performance logs (including the execution time of each operator, memory allocation / copy events, AI Core utilization, data transfer timestamps, etc.).

[0040] Step S22, computation graph structure analysis: Based on the computation graph structure information in the original performance logs, analyze the connections, dependencies, parallelism, and data flow paths between operators from the perspective of the overall computation graph constructed by MindSpore. The focus is primarily on structural inefficiencies, such as unnecessary reshape / transpose sequences, small operator groups that could be merged but haven't, and excessive cross-core synchronization points. The output is a list of computation graph redundancies, i.e., inefficient nodes.

[0041] Step S23: Operator-level time consumption statistics. Extract the execution time (including scheduling overhead, computation time, and data transfer time) of each operator from the original performance log, sort them from largest to smallest time consumption, and identify the hot operators (Top 5 or Top 10). At the same time, analyze the number of operator calls, average time consumption, maximum and minimum time consumption, etc., and output a hot operator ranking table.

[0042] Step S24: Bottleneck comprehensive location. Perform cross-attribution analysis on the list of inefficient nodes and the hotspot operator sorting table to determine the root cause of the bottleneck and output a clear conclusion on the root cause of the bottleneck.

[0043] Step S25: Output optimization plan. Based on the identified bottleneck root causes, formulate a list of specific and executable optimization plans, and prioritize them according to expected benefits and implementation difficulty to generate a priority-ranked optimization task list.

[0044] In step S30, the original RWKV-7 model is optimized using core operators based on the bottleneck root cause conclusion and the optimization task list to obtain a high-performance custom operator. Memory access optimization is performed based on the high-performance custom operator and the original RWKV-7 model to obtain a memory optimization strategy. Graph execution optimization is performed based on the computation graph of the original RWKV-7 model and the high-performance custom operator to obtain a target optimization computation graph. Finally, a high-performance RWKV-7 model is obtained based on the high-performance custom operator, the memory optimization strategy, and the target optimization computation graph.

[0045] The computing platform includes Ascend NPU, CANN heterogeneous computing architecture, and MindSpore deep learning framework.

[0046] The goal of this step, tiered optimization and performance acceleration, is to divide optimization into three levels based on bottleneck analysis results, and implement optimization measures in an orderly manner to significantly improve the inference performance of the RWKV-7 model on the Ascend NPU. The input includes a functionally correct but poorly performing RWKV-7 model (including network structure, weights, and softened operators), bottleneck root cause conclusions, and a priority list of optimization schemes. The final output includes three levels of optimization modules (custom operators, memory optimization strategies, and graph optimization patches), and an optimized model that can be integrated (not yet finalized).

[0047] In one possible implementation, the high-performance custom operator is the Ascend C WKV7 operator; the specific implementation of obtaining the high-performance custom operator is as follows: based on the bottleneck root cause conclusion and the optimization task list, the original RWKV-7 model is split into state update formulas, and an instruction sequence blueprint is output; based on the instruction sequence blueprint, a Tiling strategy is designed, and a Tiling parameter table is output; based on the instruction sequence blueprint and the Tiling parameter table, a pipeline parallel design is performed to generate the Ascend C WKV7 operator.

[0048] In the design of the custom Time-Mixing operator, Time-Mixing is a core module of the RWKV-7 network, and its computation process involves state updates and attention weighting. This application designs a dedicated WKV7 custom operator based on the Ascend C programming language, such as... Figure 4 As shown, it includes: (1) Optimization of state update formula: The original Time-Mixing state update formula of RWKV-7 is: wkv t =wkv t-1 ⊙ w t +wkv t-1 +v t Among them, wkv t For the weighted key-value state at the current time step, wkv t-1 This is the weighted key-value state from the previous time step. w t The decay weights for each time step. As the attention gating factor, v t For value vectors, Let be the key vector, and ⊙ represent element-wise multiplication. Indicates transpose. This represents the transpose of the key vector at the current time step. Figure 4 middle, o represents the transpose of the receive vector at the current time step. tThe final output vector at the current time step is represented. Based on the characteristics of the Ascend Vector computing unit, this application decomposes the above calculation into a sequence of vector operations that can be executed in parallel, and uses AICore's SPMD (Single Program Multiple Data) execution model to achieve multi-core parallelism. (2) Tiling strategy design: Based on the Unified Buffer capacity (192KB) of Ascend 910B, the optimal data block strategy is designed. The input tensor is divided according to the batch and head dimensions to ensure that the amount of data processed by each computing core does not exceed the on-chip cache capacity and reduce the number of off-chip memory accesses. (3) Pipeline parallel optimization: Utilizing the parallel execution capabilities of the matrix computing unit (AIC) and vector computing unit (AIV) of Ascend NPU, the state update calculation and matrix multiplication operation are pipelined to hide the data transfer delay.

[0049] Further, in step S311, the state update formula is decomposed (algorithm decomposition). Based on the bottleneck conclusion, the original RWKV-7 Time-Mixing state update mathematical formula (involving matrix dot product, outer product, addition, etc.) is decomposed into a set of basic matrix operation sequences that Ascend AI Core can efficiently execute in parallel. This produces a parallelizable instruction sequence blueprint.

[0050] Step S312, Tiling strategy design (data partitioning): Based on the instruction sequence blueprint and the UnifiedBuffer capacity (192KB) of the Ascend 910B, design a data partitioning scheme to divide the input tensors (e.g., batch × seq) into blocks. len ×hidden size The data is divided into multiple "tiles" based on the batch and head dimensions, where batch is the batch size and seq is the number of tiles. len The sequence length is hidden. size To hide the layer dimension, ensure that the data block processed by each computing core can be completely loaded into the UB, avoid frequent access to the off-chip HBM, and generate the Tiling parameter table.

[0051] Step S313, pipelined parallel design (time overlap): Based on the instruction sequence blueprint and Tiling parameters, three pipelines are designed using independent Cube units, Vector units, and DMA units on the Ascend NPU, producing a pipeline scheduling scheme. The final output is a high-performance custom operator, which directly replaces the inefficient software implementation in step S10.

[0052] In one possible implementation, the memory optimization strategy includes a memory reuse mapping table and double-buffered scheduling logic. The specific implementation of the memory optimization strategy is as follows: Local tensor pre-allocation is performed based on the high-performance custom operator and the original RWKV-7 model, outputting a pre-allocated memory pool; memory reuse is performed based on the computation graph of the original RWKV-7 model and the pre-allocated memory pool, outputting a memory reuse mapping table, and integrating the memory reuse mapping table into MindSpore's memory manager; data layout conversion code is output based on the operator data access mode of the Ascend C WKV7 operator; and double-buffered scheduling logic is output based on the Tiling parameter and the pipeline design of the Ascend C WKV7 operator, and integrated into the Ascend C WKV7 operator.

[0053] During the memory optimization strategy, the Ascend 910B chip has a complex multi-level storage architecture, such as... Figure 5 As shown, the AI ​​Core includes a system control module, a storage conversion unit, a matrix calculation unit, a vector calculation unit, a scalar calculation unit, and a bus. The system control module further includes a bus interface unit, an instruction cache, a scalar instruction processing queue, an instruction dispatch module, a matrix operation queue, a vector operation queue, a storage conversion queue, and an event synchronization module. The bus interface unit receives external instructions and data via the bus; the instruction cache caches instructions to be executed; the scalar instruction processing queue temporarily stores scalar instructions; the instruction dispatch module distributes instructions of various types to their corresponding operation queues; the matrix operation queue, the vector operation queue, and the storage conversion queue respectively cache matrix operation instructions, vector operation instructions, and storage conversion instructions to be executed; the event synchronization module enables synchronization between different execution units. The storage conversion unit performs data format conversion, including decompression and transpose operations. It connects an input buffer and an output buffer. The input buffer, managed by an input buffer controller, temporarily stores data to be calculated, and the output buffer temporarily stores the calculated results. The matrix calculation unit performs matrix multiplication and addition operations; the vector calculation unit performs vector operations; and the scalar calculation unit performs scalar operations. The matrix calculation unit is connected to an accumulator to accumulate parts and results of matrix operations. The data precision module adapts to calculations with different precisions (INT8 / BF16, etc.). The dedicated and general-purpose registers store intermediate variables and configuration parameters during the calculation process. All parts interact via a bus to achieve efficient parallel computing for the AI ​​Core.

[0054] The AI ​​Core employs a pipelined parallel architecture. The system control module is responsible for acquiring, decoding, and issuing instructions, distributing matrix operations, vector operations, and storage conversion instructions to their respective operation queues. The storage conversion unit decompresses and transposes the input data, converting it to the Ascend NPU-friendly NZ format before storing it in the input buffer. The matrix calculation unit and vector calculation unit execute computation tasks in parallel, outputting the results through the output buffer after completion. The accumulator handles partial matrix multiplication and addition operations and accumulation. The data precision module supports multiple precisions such as INT8 / BF16 to adapt to the needs of different inference scenarios. All parts interact via a bus, achieving efficient parallel computing.

[0055] Further, in step S321, local tensor pre-allocation is performed by inputting the model structure (knowing the shape and lifetime of all intermediate tensors) and custom operators, and producing a pre-allocated memory pool.

[0056] Step S322, memory reuse (lifecycle analysis): Analyze the lifecycle of each tensor in the computation graph, and allow tensors with non-overlapping lifecycles to share the same memory region to reduce peak memory usage and reduce the number of memory allocations. That is, input the computation graph (or a new graph from the first-level optimization) and the memory pool layout, and produce a memory reuse mapping table.

[0057] Step S323 involves data layout optimization, converting the arrangement of tensors in memory from the default ND format to the more Ascend NPU-friendly NZ format. This allows the AI ​​Core to access data continuously, improving memory bandwidth utilization. Specifically, it involves inputting the operator's data access mode, generating data layout conversion code, and calculating the modified operator memory access offset.

[0058] Step S324, double buffering technique, uses two Unified Buffers inside the operator, which alternately serve as the current computation buffer and the data prefetch buffer for the next round, further superimposing data handling and computation. That is, inputting Tiling parameters, pipeline design, and data layout, and producing double buffering scheduling logic (embedded in the Ascend C operator code).

[0059] In one possible implementation, the specific implementation of obtaining the target optimization computation graph is as follows: operator fusion is performed based on the computation graph of the original RWKV-7 model and the high-performance custom operator to output a fused computation graph; redundancy elimination is performed on the fused computation graph to output a simplified computation graph; constant folding is performed based on the simplified computation graph and the static configuration of the original RWKV-7 model to output a constant folded computation graph; and custom operator integration is performed based on the constant folded computation graph and the high-performance custom operator to output the target optimization computation graph.

[0060] During graph execution optimization, the computational graph is deeply optimized using MindSpore's graph compilation capabilities. (See [link to relevant documentation]). Figure 6 (a) and Figure 6 In (b), E represents Element The wise operation, where B stands for Broadcast operation, has the following main optimization points: (1) Operator fusion: merging multiple small operators into a single composite operator to reduce kernel startup overhead; (2) Redundancy elimination: eliminating redundant broadcast, transformation and other operations in the computation graph to simplify the execution path; (3) Constant folding: pre-compiling constant expressions during compilation to reduce runtime computation; (4) Custom operator integration: seamlessly integrating the WKV7 operator implemented by Ascend C into the MindSpore computation graph to achieve end-to-end optimization.

[0061] Further, in step S331, operator fusion is performed. Based on the computation graph and custom operators, multiple consecutive small operators are merged into a composite operator, reducing kernel startup overhead and the need to write back intermediate results. This produces a fused computation graph (with fewer nodes and the appearance of composite operators).

[0062] Step S332: Redundancy elimination. Based on the fused computation graph, perform static analysis on the computation graph, delete all nodes that do not have a practical impact, and produce a simplified computation graph (with a further reduction in the number of nodes).

[0063] Step S333: Constant folding. Based on the simplified computation graph and the static configuration of the model, identify all operators whose inputs are known at compile time, calculate the results in advance, replace them with constant nodes, and produce a computation graph with constant folding (containing constant nodes to reduce runtime computation).

[0064] Step S334: Custom operator integration. Based on the optimized computation graph and the Ascend C operator, the Ascend C WKV7 custom operator generated in the first level is registered as a native MindSpore operator in the computation graph, replacing the original part composed of the softening subgraph (a group of small operators). The final optimized computation graph is produced (containing custom operator nodes, replacing the original multiple small operators).

[0065] In step S40, the original RWKV-7 model, the bottleneck root cause conclusion, the tuning scheme, and the high-performance RWKV-7 model are integrated and verified end-to-end to obtain the target optimized RWKV-7 model. In one possible implementation, the high-performance custom operator, the memory optimization strategy, and the target optimization computation graph in the high-performance RWKV-7 model are integrated into the MindSpore computation graph to obtain an integrated model; the integrated model is compared with the original RWKV-7 model for correctness verification, and different batch sizes and sequence lengths are set according to the bottleneck root cause conclusion and the optimization task list to verify the performance of the integrated model; the integrated model that passes the correctness verification and the performance verification is taken as the target optimized RWKV-7 model.

[0066] Understandably, the integrated model is compared with the original RWKV-7 model for correctness verification: the same test text is input into the integrated model and the original RWKV-7 model, and the output results of the two are compared to confirm that the semantic coherence of the output results is consistent and the numerical error is within an acceptable range; according to the bottleneck root cause conclusion and the optimization task list, different batch sizes and sequence lengths are set to verify the performance of the integrated model, and inference latency, throughput and memory usage are measured to confirm that the performance indicators meet the expected improvement.

[0067] Specifically, the optimized memory strategy and graph execution scheme, especially the newly developed Ascend C WKV7 operator, were seamlessly integrated into the MindSpore computation graph. Correctness verification was performed to confirm that, after this series of complex optimizations, the model's final output was completely consistent with the original output. Performance verification was conducted, testing the final inference speed, throughput, and latency at different batch sizes and sequence lengths, demonstrating a significant improvement in overall performance.

[0068] In step S50, a natural language processing inference task is obtained, and the natural language processing inference task is input into the target optimization RWKV-7 model. After the target optimization RWKV-7 model executes the natural language processing inference task on the Ascend NPU, it outputs the task execution result.

[0069] The technical solution of this invention brings about significant performance improvements, such as... Figure 7 As shown, the specific performance benefits include: WKV7 custom operators provide multiple sub-level speedups compared to the native Python implementation; increased character fragment throughput; reduced average inference latency; and improved overall inference performance.

[0070] Understandably, the method in this application can be extended to larger-scale RWKV models (such as RWKV-7-3B, RWKV-7-7B, etc.) by adjusting the Tiling strategy and parallelism to adapt to a larger model parameter scale. Building upon single-card optimization, multi-card data parallelism and model parallelism are further implemented, supporting inference with larger batch sizes. Combining the INT8 / BF16 computing power of the Ascend NPU, quantized inference of the RWKV-7 model is achieved, further improving inference efficiency. The optimization method can be extended to Ascend adaptation of other linear attention architectures (such as RetNet, Mamba, etc.). Based on inference optimization, training tasks of the RWKV-7 model on the Ascend platform are further supported.

[0071] This application provides a complete technical path for deploying advanced AI models such as RWKV on target region computing platforms, which helps promote the construction of the AI ​​chip ecosystem and technological self-reliance in the target region. The optimization method fully considers the hardware characteristics of different Ascend chip models and has good platform migration capabilities.

[0072] Next, referring to the accompanying drawings, a system for porting and optimizing RWKV-7 networks for Ascend NPU, as proposed in the embodiments of this application, is described to implement the method for porting and optimizing RWKV-7 networks for Ascend NPU as described in any of the above schemes.

[0073] Figure 8 This is a structural diagram of the RWKV-7 network porting and optimization system for Ascend NPU according to an embodiment of this application.

[0074] like Figure 8 As shown, the RWKV-7 network porting and optimization system for Ascend NPU includes: a model building module 100, a performance analysis and bottleneck location module 200, a hierarchical optimization and performance acceleration module 300, a graph execution optimization module 400, and a model execution task module 500.

[0075] Specifically, the model building module 100 is used to build the original RWKV-7 model on a set computing power platform; The performance analysis and bottleneck location module 200 is used to obtain the original performance logs of the original RWKV-7 model running by setting the performance analysis tool of the computing power platform, perform computation graph structure analysis based on the original performance logs to obtain a list of inefficient nodes, perform operator-level time consumption statistics based on the original performance logs to obtain a hot operator ranking table, perform cross-attribution analysis on the list of inefficient nodes and the hot operator ranking table to obtain a bottleneck root cause conclusion, and determine a priority-ranked optimization task list based on the bottleneck root cause conclusion. The hierarchical optimization and performance acceleration module 300 is used to optimize the core operators of the original RWKV-7 model according to the bottleneck root cause conclusion and the optimization task list to obtain a high-performance custom operator; to optimize memory access according to the high-performance custom operator and the original RWKV-7 model to obtain a memory optimization strategy; to optimize graph execution according to the computation graph of the original RWKV-7 model and the high-performance custom operator to obtain a target optimization computation graph; and to obtain a high-performance RWKV-7 model according to the high-performance custom operator, the memory optimization strategy and the target optimization computation graph. The graph execution optimization module 400 is used to integrate and verify the original RWKV-7 model, the bottleneck root cause conclusion, the tuning scheme and the high-performance RWKV-7 model to obtain the target optimized RWKV-7 model. The model execution task module 500 is used to acquire a natural language processing inference task, input the natural language processing inference task into the target optimization RWKV-7 model, and after the target optimization RWKV-7 model executes the natural language processing inference task on the Ascend NPU, it outputs the task execution result.

[0076] Figure 9 A structural diagram of a terminal provided in an embodiment of this application. The terminal may include: The memory 501, the processor 502, and the computer program stored on the memory 501 and capable of running on the processor 502.

[0077] When the processor 502 executes the program, it implements the RWKV-7 network porting and optimization method for Ascend NPU provided in the above embodiments.

[0078] Furthermore, the terminal also includes: Communication interface 503 is used for communication between memory 501 and processor 502.

[0079] The memory 501 is used to store computer programs that can run on the processor 502.

[0080] Memory 501 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.

[0081] If the memory 501, processor 502, and communication interface 503 are implemented independently, then the communication interface 503, memory 501, and processor 502 can be interconnected via a bus to complete communication between them. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EIS) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, Figure 9 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0082] Optionally, in a specific implementation, if the memory 501, processor 502, and communication interface 503 are integrated on a single chip, then the memory 501, processor 502, and communication interface 503 can communicate with each other through an internal interface.

[0083] Processor 502 may be a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application.

[0084] This embodiment also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for porting and optimizing RWKV-7 networks for Ascend NPU.

[0085] One embodiment of this application provides a computer program product, including a computer program that, when executed by a processor, implements the features described in this application. Figure 1 The corresponding embodiments provide RWKV-7 network porting and optimization methods for Ascend NPU.

[0086] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0087] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "N" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0088] Any process or method described in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or N executable instructions for implementing custom logic functions or processes, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of this application pertain.

[0089] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable storage medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable storage medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable storage media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable storage medium could be paper or other suitable media on which the program can be printed, since the program can be obtained electronically by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.

[0090] It should be understood that the various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0091] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.

[0092] Furthermore, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0093] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of this application.

[0094] It should be understood that the application of this application is not limited to the examples above. Those skilled in the art can make improvements or modifications based on the above description, and all such improvements and modifications should fall within the protection scope of the appended claims.

[0095] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.

Claims

1. A RWKV-7 network transplantation and optimization method for Ascend NPU, characterized in that, The method for porting and optimizing the RWKV-7 network for Ascend NPU includes: The original RWKV-7 model was built on the specified computing power platform; The original performance logs of the original RWKV-7 model are obtained by setting the performance analysis tool of the computing power platform. The computation graph structure is analyzed based on the original performance logs to obtain a list of inefficient nodes. The operator-level time consumption statistics are performed based on the original performance logs to obtain a hot operator ranking table. The list of inefficient nodes and the hot operator ranking table are cross-attribution analysis to obtain the bottleneck root cause conclusion. The priority ranking of optimization tasks is determined based on the bottleneck root cause conclusion. Based on the bottleneck root cause conclusions and the optimization task list, the original RWKV-7 model is optimized to obtain a high-performance custom operator, including: based on the bottleneck root cause conclusions and the optimization task list, the original RWKV-7 model is decomposed into a state update formula and an instruction sequence blueprint is output, including: the state update formula of the original RWKV-7 model is decomposed into independent vector multiplication and addition and outer product operation units, and rearranged into a parallel-issue instruction sequence according to data dependencies; The Tiling strategy is designed based on the instruction sequence blueprint, and a Tiling parameter table is output; the pipeline parallel design is performed based on the instruction sequence blueprint and the Tiling parameter table, and an Ascend C WKV7 operator is generated; wherein, the high-performance custom operator is an Ascend C WKV7 operator; Memory optimization strategies are derived by optimizing memory access based on the high-performance custom operator and the original RWKV-7 model. These strategies include: pre-allocating local tensors based on the high-performance custom operator and the original RWKV-7 model, and outputting a pre-allocated memory pool; reusing memory based on the computation graph of the original RWKV-7 model and the pre-allocated memory pool, outputting a memory reuse mapping table, and integrating the memory reuse mapping table into MindSpore's memory manager; outputting data layout conversion code based on the operator data access mode of the Ascend C WKV7 operator; and outputting double-buffered scheduling logic based on the Tiling parameter and the pipeline design of the Ascend C WKV7 operator, and integrating the double-buffered scheduling logic into the Ascend C WKV7 operator. The memory optimization strategy includes the memory reuse mapping table and the double-buffered scheduling logic. The target optimization computation graph is obtained by performing graph execution optimization based on the computation graph of the original RWKV-7 model and the high-performance custom operator. The high-performance RWKV-7 model is obtained based on the high-performance custom operator, the memory optimization strategy and the target optimization computation graph. The original RWKV-7 model, the bottleneck root cause conclusions, the optimization task list, and the high-performance RWKV-7 model are integrated and validated end-to-end to obtain the target optimization RWKV-7 model. A natural language processing inference task is obtained, and the natural language processing inference task is input into the target optimization RWKV-7 model. After the target optimization RWKV-7 model executes the natural language processing inference task on the Ascend NPU, it outputs the task execution result.

2. The NPU-oriented RWKV-7 network transplantation and optimization method according to claim 1, characterized in that, The construction of the original RWKV-7 model on the designated computing power platform specifically includes: The MindSpore model is obtained by reconstructing the network structure using the MindSpore deep learning framework in the computing platform. Obtain the checkpoint file that matches the MindSpore model and the operator implementation fragment corresponding to the MindSpore model; Functional verification was performed based on the MindSpore model, the checkpoint file, and the operator implementation fragment to obtain the original RWKV-7 model.

3. The NPU-oriented RWKV-7 network transplantation and optimization method according to claim 1, characterized in that, The step of obtaining the target optimization computation graph by performing graph execution optimization based on the computation graph of the original RWKV-7 model and the high-performance custom operator specifically includes: Based on the computation graph of the original RWKV-7 model and the high-performance custom operator, operator fusion is performed to output a fused computation graph; Redundancy is eliminated from the fused computation graph, and a simplified computation graph is output. Based on the simplified computation graph and the static configuration of the original RWKV-7 model, constant folding is performed, and a constant folding computation graph is output. Custom operator integration is performed based on the constant folding computation graph and the high-performance custom operator to output the target optimization computation graph.

4. The method for porting and optimizing RWKV-7 networks for Ascend NPU according to claim 1, characterized in that, The computing platform includes Ascend NPU, CANN heterogeneous computing architecture, and MindSpore deep learning framework.

5. The method for porting and optimizing RWKV-7 networks for Ascend NPU according to claim 1, characterized in that, The process of integrating and end-to-end verifying the original RWKV-7 model, the bottleneck root cause conclusions, the optimization task list, and the high-performance RWKV-7 model to obtain the target optimization RWKV-7 model specifically includes: The high-performance custom operator, the memory optimization strategy, and the target optimization computation graph in the high-performance RWKV-7 model are integrated into the MindSpore computation graph to obtain the integrated model. The integrated model is compared with the original RWKV-7 model for correctness verification. Based on the bottleneck root cause conclusion and the optimization task list, different batch sizes and sequence lengths are set to verify the performance of the integrated model. The integrated model that has passed the correctness verification and the performance verification is used as the target to optimize the RWKV-7 model.

6. A system for porting and optimizing RWKV-7 networks for Ascend NPU, characterized in that, The RWKV-7 network porting and optimization system for Ascend NPU is used to implement the RWKV-7 network porting and optimization method for Ascend NPU as described in any one of claims 1-5. The RWKV-7 network porting and optimization system for Ascend NPU includes: The model building module is used to build the original RWKV-7 model on a specified computing power platform. The performance analysis and bottleneck location module is used to obtain the original performance logs of the original RWKV-7 model by setting the performance analysis tools of the computing power platform, perform computation graph structure analysis based on the original performance logs to obtain a list of inefficient nodes, perform operator-level time consumption statistics based on the original performance logs to obtain a hot operator ranking table, perform cross-attribution analysis on the list of inefficient nodes and the hot operator ranking table to obtain the bottleneck root cause conclusion, and determine the priority ranking of optimization tasks based on the bottleneck root cause conclusion. The hierarchical optimization and performance acceleration module is used to optimize the core operators of the original RWKV-7 model based on the bottleneck root cause conclusions and the optimization task list to obtain high-performance custom operators; to optimize memory access based on the high-performance custom operators and the original RWKV-7 model to obtain a memory optimization strategy; to optimize graph execution based on the computation graph of the original RWKV-7 model and the high-performance custom operators to obtain a target optimization computation graph; and to obtain a high-performance RWKV-7 model based on the high-performance custom operators, the memory optimization strategy, and the target optimization computation graph. The graph execution optimization module is used to integrate and perform end-to-end verification based on the original RWKV-7 model, the bottleneck root cause conclusion, the optimization task list and the high-performance RWKV-7 model to obtain the target optimized RWKV-7 model. The model execution task module is used to acquire a natural language processing inference task, input the natural language processing inference task into the target optimization RWKV-7 model, and after the target optimization RWKV-7 model executes the natural language processing inference task on the Ascend NPU, it outputs the task execution result.

7. A terminal, characterized in that, The terminal includes: a memory, a processor, and an RWKV-7 network porting and optimization program for Ascend NPU stored in the memory and executable on the processor. When the RWKV-7 network porting and optimization program for Ascend NPU is executed by the processor, it implements the steps of the RWKV-7 network porting and optimization method for Ascend NPU as described in any one of claims 1-5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program for porting and optimizing an RWKV-7 network for Ascend NPU. When the program is executed by a processor, it implements the steps of the method for porting and optimizing an RWKV-7 network for Ascend NPU as described in any one of claims 1-5.