A Software Performance Optimization Method and System under a Domestic Chip Architecture
Through the software performance optimization method under the domestic chip architecture, the problem of limited computing resources of power AI applications at the edge is solved, the computing efficiency and flexibility are improved, and the real-time and reliability requirements of the power grid system are met.
Patent Information
- Application Number
- CN202510505194.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-22
AI Technical Summary
In modern power grid systems, power AI applications face rising data volume, improved real-time requirements and reliability requirements in complex environments, especially when edge computing resources are limited.
The software performance optimization method under the domestic chip architecture is adopted, including task segmentation, independent service model, Kubernetes containerized management, Istio service routing, Kirin chip acceleration library, hybrid accuracy calculation, PaddlePaddle framework optimization and TensorRT model optimization.
It improves the computing efficiency and flexibility of power AI applications, realizes efficient and high-quality computing in resource-constrained environments, and supports real-time and reliability of edge nodes.
Smart Images

Figure CN120029741B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of software optimization, and particularly to a method and system for optimizing software performance under a domestic chip architecture. Background Art
[0002] In modern power grid systems, with the popularization of Internet of Things devices and the application of smart grid technologies, a large amount of real-time data has been generated. These data play an important role in power grid monitoring, prediction, and fault detection. However, the traditional centralized computing mode often fails to meet the requirements of power grid real-time performance and high reliability. Therefore, more and more application scenarios are beginning to adopt edge computing and artificial intelligence technologies to address these challenges.
[0003] With the development of smart grids and the popularization of Internet of Things devices, power AI applications are facing multiple challenges such as a sharp increase in data volume, higher real-time requirements, and reliability requirements in complex environments. Especially in the case of limited resources in edge computing, power AI applications face a series of challenges during operation, including limitations in computing power, storage, energy consumption, and communication. Summary of the Invention
[0004] In order to solve the above problems, the purpose of the present invention is to provide a method and system for optimizing software performance under a domestic chip architecture, effectively improving the operation efficiency and flexibility of power AI applications.
[0005] To achieve the above purpose, the present invention adopts the following technical solutions:
[0006] A method for optimizing software performance under a domestic chip architecture, comprising the following steps
[0007] S1: Subdivide the tasks of the power AI application and clarify the KPIs for improving application performance, including reducing latency and improving accuracy;
[0008] S2: According to the task subdivision, decompose the AI application into independent service models to improve scalability, and use Kubernetes for container management to achieve automatic scaling and fault recovery; and utilize Istio to achieve refined service routing and traffic management;
[0009] S3: Adopt the domestic chip Kunpeng that supports artificial intelligence acceleration, design a dedicated acceleration library using custom instruction sets, and configure mixed-precision computing and high bandwidth according to application requirements;
[0010] S4: Use the domestic framework of PaddlePaddle to optimize the independent service model to adapt to the domestic chip Kunpeng;
[0011] S5: Deploy the optimized independent service model at the edge nodes of the power grid, and use TensorRT to optimize the model to improve local processing capabilities.
[0012] Furthermore, according to the task breakdown, decompose the AI application into independent service models to improve scalability, and use Kubernetes for containerized management to achieve automatic scaling and fault recovery. Specifically:
[0013] Design independent services for each functional module. The independent services communicate through APIs and are independently deployed and scaled;
[0014] Use Docker to encapsulate each independent service into a container to ensure environmental consistency and independence. Create a container image containing the application and its dependencies and upload it to the container image repository;
[0015] Deploy a Kubernetes cluster to manage containerized applications. Use configuration files to define Pod, ReplicaSet, and Deployment resources. Utilize the automatic scaling function of Kubernetes to dynamically adjust the number of Pods according to CPU / memory usage to achieve automatic scaling;
[0016] Set up health checks and restart policies to ensure that failed services can be automatically detected and recovered. Configure persistent volumes and persistent volume claims to ensure data persistence after Pod restart.
[0017] Furthermore, utilize Istio to achieve fine-grained service routing and traffic management, specifically as follows:
[0018] Set up Istio in the Kubernetes cluster. By injecting the sidecar proxy into each Pod, intercept and route communications;
[0019] Configure the control plane components of Istio to manage the network traffic between services within the cluster;
[0020] Utilize Istio's VirtualService and DestinationRule to define the routing rules between services to achieve traffic routing, load balancing, and fault recovery;
[0021] Implement canary releases and blue-green deployments, allowing for gradual updates and rollbacks without affecting the overall system;
[0022] Use Istio's monitoring and observability components to obtain service call chains and performance metrics to help identify performance bottlenecks.
[0023] Furthermore, S3 is specifically as follows:
[0024] S31: Develop a dedicated acceleration library using the instruction set provided by the Kirin chip, and design library functions for accelerating the execution of AI models. The library functions are optimized for the NPU of the chip using low-level programming interfaces;
[0025] S32: Use mixed precision calculation of FP16 and FP32 during model training to maintain computational efficiency and result accuracy; and during the model inference stage, use INT8 quantization to reduce memory usage and accelerate model execution;
[0026] S33: Through data preloading and asynchronous data stream DM technology, hide memory access latency, and combine with the high bandwidth characteristics of Kirin to optimize the data inflow and outflow channels to ensure the synchronization of data transmission and processing;
[0027] S34: Implement intelligent memory allocation to ensure that critical data resides in the cache and reduce the memory round-trip of high-frequency data; allocate buffers for large-scale matrix operations in the input and output layers.
[0028] Further, S32 is specifically as follows:
[0029] During the training process, maintain the 32-bit floating-point precision of the model parameters to accumulate accurate gradient updates; before each batch calculation, convert the FP32 model parameters to FP16 for fast calculation on the Tensor Core unit of the GPU;
[0030] Use FP16 to perform matrix operations and convolution operations in the forward and backward propagations.
[0031] When calculating gradients in the backward propagation, use FP16 for the operations and then convert to FP32 to accumulate the gradients;
[0032] Use a dynamic loss scaling strategy to avoid numerical underflow problems that may be introduced by FP16;
[0033] During training or after training quantization, convert the FP32 model parameters to the INT8 format, collect the maximum and minimum values of activations and weights through the calibration dataset for calculating the quantization scaling factor, and use the quantized INT8 for matrix multiplication and convolution operations.
[0034] Further, use a dynamic loss scaling strategy to avoid numerical underflow problems that may be introduced by FP16, specifically as follows:
[0035] At the start of training, select a scaling factor S greater than the preset value to scale the loss to increase the numerical stability of the calculation;
[0036] Calculate the original loss L in the forward propagation;
[0037] Scale the original loss L by a scaling factor S:
[0038] L scaled = L × S;
[0039] During backpropagation, calculate the gradient according to the scaled loss L scaled Calculate the gradient;
[0040] Due to the scaled loss, the gradient is also amplified accordingly;
[0041] The calculated gradient g scaled Is amplified and restored to the original proportional gradient g by unscaling. Divide the gradient g scaled By S:
[0042]
[0043] Perform a gradient check on the calculation. If there are non - numbers NaN or infinite values Inf in the gradient, reduce the scaling factor S proportionally, discard the current update cycle and restore to the previous stable model state; If there are no abnormalities in the gradient calculation: Select to increase the scaling factor S proportionally to improve the operation accuracy.
[0044] Furthermore, S34 is specifically:
[0045] Analyze and identify the data with an access frequency greater than the threshold in the model calculation as high - frequency data; such as intermediate tensors;
[0046] Divide the memory into multiple levels, from the fastest and smallest - capacity L1 cache to the large - capacity main memory;
[0047] Save the high - frequency data in the L1 or L2 cache and utilize the high - speed cache characteristics of Kirin to ensure the access speed;
[0048] Migrate the intermediate results and other large - scale memory usage scenarios to the main memory; Pre - allocate dedicated buffers for the input / output layer and key computing tasks to ensure that large - scale matrix operations do not cause bottlenecks due to improper memory allocation; Implement a memory pool to dynamically allocate and release memory, avoid frequent malloc / free operations, and reduce memory fragmentation;
[0049] Adjust the data layout method according to the data access pattern to maximize the cache hit rate.
[0050] Furthermore, S4 is specifically:
[0051] Use the built - in model conversion tool in PaddlePaddle to convert the model into a form suitable for the Kirin architecture to make full use of the hardware's computing power;
[0052] Enable the AMP function of PaddlePaddle and use mixed-precision computing (FP16 + FP32) in training and inference to improve performance and reduce memory usage;
[0053] For operations enabled on Kirin chips, analyze and customize the implementation of operators, and optimize key operators using the instruction sets and parallel acceleration features provided by the chips;
[0054] Use the Operator Layer support of PaddlePaddle to perform operator-level performance tuning according to actual needs;
[0055] Use PaddleSlim for model pruning to reduce unimportant neurons or channels, maintaining model effectiveness while reducing the computational burden;
[0056] Use the ParallelExecutor of PaddlePaddle to perform model inference in parallel, leveraging NPU and DSP acceleration units;
[0057] Implement layer fusion in the model execution pipeline to reduce memory interaction and cache switching, use pipelining execution technology to improve data flow efficiency, and adapt to the characteristics of the chip's hardware execution units.
[0058] Furthermore, implement layer fusion in the model execution pipeline to reduce memory interaction and cache switching as follows:
[0059] Utilize the static mode of PaddlePaddle to customize operators or use supported built-in operators to attempt the fusion process; identify connectable computational parts in the computational graph optimization process and perform corresponding operation fusion;
[0060] Analyze the execution graph of the model, divide the model into multiple stages, and execute each stage in parallel on different computing resources;
[0061] Adopt the producer-consumer model to ensure that the data flow of each stage can be synchronously transmitted;
[0062] In the pipeline, submit each execution stage in the form of an asynchronous task, and achieve the parallelism of each stage through the Executor of PaddlePaddle;
[0063] Bind the computing resources to the types of hardware execution units of the Kirin chip to ensure that each pipeline segment matches the best resources, and use a circular buffer to ensure synchronous and delay-free data transmission between execution stages.
[0064] A software performance optimization system under a domestic chip architecture, including a processor, a memory, and a computer program stored on the memory. When the processor executes the computer program, it specifically executes the steps in the above-mentioned software performance optimization method under a domestic chip architecture.
[0065] The present invention has the following beneficial effects:
[0066] 1. Through the in-depth combination of software and hardware, from the framework to the underlying chips, and from application decomposition to deployment management, the present invention improves the computing efficiency and flexibility of power AI applications, enabling the applications to maintain high-efficiency and high-quality computing in the power grid edge environment where resources are limited and high reliability and real-time computing are required.
[0067] 2. The present invention decomposes power AI applications into independent service models, improves the flexibility and scalability of the system, and through the containerized management of Kubernetes, realizes dynamic resource scheduling, promotes automatic expansion and fault recovery, enables different models to be independently developed, deployed, and expanded for easy maintenance and optimization, and finely manages the traffic and security between services through Service Mesh.
[0068] 3. By customizing instruction sets and dedicated acceleration libraries, the present invention makes full use of the hardware acceleration characteristics of Kirin chips, implements mixed-precision computing and high-bandwidth memory strategies to maximize performance, effectively improves the model computing efficiency, shortens the inference time, reduces computing energy consumption, and thus better supports the efficient operation of each microservice on domestic hardware; combined with the optimization ability of PaddlePaddle on domestic chip Kirin, efficient model compression and performance optimization can be achieved, enabling deep learning models to operate efficiently on resource-constrained terminal devices. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] Figure 1 It is a flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0070] The following further describes the present invention in detail with reference to the drawings and specific embodiments:
[0071] Reference Figure 1 , in this embodiment, a software performance optimization method under a domestic chip architecture includes the following steps
[0072] S1: Subdivide the tasks of power AI applications (for AI applications in the power system, subdivide them into load prediction models, fault detection models, energy consumption optimization models, etc.), and clarify the KPIs for application performance improvement, including reducing latency and improving accuracy.
[0073] In this embodiment, load forecasting: Forecast the future power demand of the power grid to optimize power generation and distribution. KPI: Accuracy: The goal is to reduce the mean squared error (MSE) below a certain threshold; Response latency: The update latency of the forecast result is less than 5 seconds.
[0074] Fault detection: Real-time detection and identification of anomalies and faults in the power grid. KPI: Identification accuracy: Higher than 98%; Response time: Less than 1 second to ensure a quick response.
[0075] Energy consumption optimization: Optimize the power consumption of the power system to improve the overall energy efficiency. KPI: Energy efficiency improvement: The target for system energy efficiency improvement is 5%; Cost savings: The target for reducing operating costs is 2%.
[0076] S2: According to the task breakdown, decompose the AI application into independent service models to improve scalability, and use Kubernetes for containerized management to achieve automatic scaling and fault recovery; and utilize Istio to achieve fine-grained service routing and traffic management;
[0077] S3: Adopt the domestic chip Kirin that supports artificial intelligence acceleration, design a dedicated acceleration library using custom instruction sets, and configure mixed-precision computing and high bandwidth according to application requirements;
[0078] S4: Use the domestic framework of PaddlePaddle to optimize the independent service model to adapt to the domestic chip Kirin;
[0079] S5: Deploy the optimized independent service model at the power grid edge nodes, use TensorRT to optimize the model to improve local processing capabilities;
[0080] In this embodiment, use the Python API of TensorRT or the command-line tool (trtexec) to convert the model into a TensorRT engine. Specify optimization parameters such as batch size, maximum workspace size, and target precision. Save the TensorRT engine for quick loading and startup on edge devices. Develop an inference service interface at the edge nodes to expose the model inference service through methods such as Rest API or RPC. Combine with specific power grid application scenarios to integrate sensor data collection and prediction inference. Utilize the parallel capabilities of edge devices to set parallel execution plans to maximize the use of the asynchronous execution characteristics of TensorRT. Optimize the preprocessing and postprocessing steps of the input data and use them as part of the pipeline to reduce the overall latency. Using TensorRT to optimize the model can significantly improve the execution efficiency of the model at the power grid edge nodes, ensuring fast and reliable inference results under resource-constrained conditions. This helps with real-time monitoring, analysis, and response in a smart grid environment to improve the overall efficiency and reliability of the power grid.
[0081] Preferably, use Prometheus to build a real-time monitoring system to monitor system performance and model inference time; implement A / B testing and user feedback collection mechanisms, and use data-driven methods to improve model algorithms and system bottlenecks.
[0082] In this embodiment, according to the task breakdown, the AI application is decomposed into independent service models to improve scalability, and Kubernetes is used for containerized management to achieve automatic scaling and fault recovery. Specifically:
[0083] Design independent services for each functional module (such as load prediction, fault detection, energy consumption optimization). The independent services communicate through APIs and are independently deployed and scaled;
[0084] Use Docker to encapsulate each independent service into a container to ensure environmental consistency and independence. Create a container image containing the application and its dependencies and upload it to the container image repository;
[0085] Deploy a Kubernetes cluster to manage containerized applications. Use configuration files (YAML) to define Pod, ReplicaSet, and Deployment resources. Utilize the automatic scaling function of Kubernetes, such as Horizontal Pod Autoscaler (HPA), to dynamically adjust the number of Pods according to CPU / memory usage to achieve automatic scaling;
[0086] Set up health checks and restart policies to ensure that failed services can be automatically detected and recovered. Configure Persistent Volume and Persistent Volume Claim to ensure data persistence after Pod restart.
[0087] In this embodiment, use Istio to achieve fine-grained service routing and traffic management, specifically as follows:
[0088] Set up Istio in the Kubernetes cluster. By injecting sidecar proxies (Envoy) into each Pod, communication interception and routing are achieved;
[0089] Configure the control plane components of Istio (such as pilot, mixer) to manage network traffic between services in the cluster;
[0090] Utilize Istio's VirtualService and DestinationRule to define routing rules between services to achieve traffic routing, load balancing, and fault recovery;
[0091] Implement gray release and blue-green deployment, allowing for gradual updates and rollbacks without affecting the overall system;
[0092] Use Istio's monitoring and observability components (such as Kiali, Jaeger) to obtain service call chains and performance metrics to help identify performance bottlenecks.
[0093] In this embodiment, S3 is specifically as follows:
[0094] S31: Develop a dedicated acceleration library using the instruction set provided by the Kirin chip, especially optimized for common operations in deep learning (such as matrix multiplication, convolution operations); and design library functions for accelerating the execution of AI models, and optimize the performance of the library functions for the NPU of the chip using low-level programming interfaces;
[0095] S32: Utilize mixed-precision computing of FP16 and FP32 in model training to maintain computing efficiency and result accuracy; and in the model inference stage, use INT8 quantization to reduce memory usage and accelerate model execution;
[0096] S33: Through data preloading and asynchronous data flow DM technology, hide memory access latency, and combine with the high-bandwidth characteristics of Kirin to optimize the data inflow and outflow channels to ensure the synchronization of data transmission and processing;
[0097] S34: Implement intelligent memory allocation to ensure that critical data (such as intermediate tensors) resides in the cache to reduce memory round-trips for high-frequency data; allocate buffers for large-scale matrix operations in the input and output layers.
[0098] In this embodiment, S32 is specifically as follows:
[0099] During the training process, maintain the 32-bit floating-point precision of the model parameters to accumulate accurate gradient updates; before each batch calculation, convert the FP32 model parameters to FP16 for fast calculation on the Tensor Core unit of the GPU;
[0100] Use FP16 to perform matrix operations and convolution operations in forward and backward propagations.
[0101] When calculating gradients in the backward propagation, use FP16 for the operation and then convert to FP32 to accumulate the gradients;
[0102] Use a dynamic loss scaling strategy to avoid numerical underflow problems that may be introduced by FP16;
[0103] During training or post-training quantization, convert FP32 model parameters to INT8 format, collect the maximum and minimum values of activations and weights through a calibration dataset for calculating quantization scaling factors, and perform matrix multiplication and convolution operations using the quantized INT8, running on hardware such as the accelerator of Kirin chips that supports this calculation.
[0104] In this embodiment, a dynamic loss scaling strategy is used to avoid the numerical underflow problem that FP16 may introduce, as follows:
[0105] When starting training, select a scaling factor S greater than a preset value (such as 1024 or higher) to scale the loss to increase the numerical stability of the calculation;
[0106] Calculate the original loss L during forward propagation;
[0107] Scale the original loss L by the scaling factor S for calculation:
[0108] L scaled =L×S;
[0109] During backpropagation, calculate the gradient according to the scaled loss L scaled Calculate the gradient;
[0110] Due to the scaled loss, the gradient is also amplified accordingly;
[0111] The calculated gradient g scaled Is amplified and restored to the original proportion gradient g by inverse scaling. Divide the gradient g scaled By S:
[0112]
[0113] Perform gradient checking on the calculation. If there are non-numeric NaN or infinite values Inf in the gradient, reduce the scaling factor S proportionally (for example, reduce it to half of the previous value), discard the current update cycle and restore to the previous stable model state; if there is no abnormality in the gradient calculation: select to increase the scaling factor S proportionally to improve the operation accuracy (such as after a certain number of steps).
[0114] In this embodiment, S34 is specifically:
[0115] Analyze and identify data with an access frequency greater than a threshold in model calculations as high-frequency data; such as intermediate tensors;
[0116] Divide the memory into multiple levels, from the fastest and smallest-capacity L1 cache to the large-capacity main memory;
[0117] Save the high-frequency data in the L1 or L2 cache and utilize the cache characteristics of Kirin to ensure the access speed;
[0118] Migrate intermediate results and other large-scale memory usage scenarios to the main memory; pre-allocate dedicated buffers for the input / output layer and critical computing tasks to ensure that large-scale matrix operations do not cause bottlenecks due to improper memory allocation; implement a Memory Pool to dynamically allocate and release memory, avoid frequent malloc / free operations, and reduce memory fragmentation;
[0119] Adjust the data layout according to the data access pattern (such as row-major or column-major) to maximize the cache hit rate.
[0120] In this embodiment, S4 is specifically as follows:
[0121] Use the built-in model conversion tool in PaddlePaddle to convert the model into a form suitable for the Kunpeng architecture to fully utilize the computing power of the hardware;
[0122] Enable the AMP function of PaddlePaddle and use mixed-precision computing (FP16 + FP32) in training and inference to improve performance and reduce memory usage;
[0123] For operations enabled on the Kunpeng chip, analyze and customize the implementation of operators, and optimize key operators using the instruction set and parallel acceleration features provided by the chip;
[0124] Use the Operator Layer support of PaddlePaddle to perform operator-level performance tuning according to actual needs;
[0125] Use PaddleSlim for model pruning to reduce unimportant neurons or channels, maintain the effectiveness of the model while reducing the computational burden;
[0126] Use the ParallelExecutor of PaddlePaddle to perform model inference in parallel and utilize the NPU and DSP acceleration units;
[0127] Implement layer fusion in the model execution pipeline to reduce memory interaction and cache switching, use pipeline execution technology to improve data flow efficiency, and adapt to the characteristics of the chip's hardware execution units.
[0128] Preferably, implement layer fusion in the model execution pipeline to reduce memory interaction and cache switching, as follows:
[0129] Utilize the static mode of PaddlePaddle to customize operators or use supported built-in operators to attempt the fusion process; identify connectable computational parts in the computational graph optimization process and perform corresponding operation fusion;
[0130] Analyze the execution graph of the model, divide the model into multiple stages, and each stage is executed in parallel on different computing resources (such as CPUs and NPUs);
[0131] Adopt the producer-consumer model to ensure that the data flow of each stage can be synchronously transmitted; (Each stage can be either a producer or a consumer: Producer: Responsible for computing and providing data to the next stage. Consumer: Receives data from the previous stage for processing)
[0132] In the pipeline, submit each execution stage in the form of an asynchronous task, and achieve the parallelism of each segment through the Executor of PaddlePaddle;
[0133] Bind the computing resources to the hardware execution unit types of Kirin chips to ensure that each pipeline segment matches the best resources, and use a circular buffer to ensure the synchronization and delay-free transmission of data between execution stages.
[0134] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can be implemented in the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can be implemented in the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0135] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the processes and / or blocks in the flowchart and / or block diagram can also be implemented. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate for implementing in the process Figure 1 a process or multiple processes and / or blocks Figure 1 a block or multiple blocks the device for the specified functions.
[0136] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device implements in the process Figure 1 a process or multiple processes and / or blocks Figure 1 a block or multiple blocks the specified functions.
[0137] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are executed on the computer or other programmable apparatus to produce a computer-implemented process, thereby the instructions executed on the computer or other programmable apparatus provide steps for realizing the functions specified in one process or a plurality of processes and / or one block or a plurality of blocks. Figure 1 one process or a plurality of processes and / or Figure 1 blocks.
[0138] As mentioned above, it is only the preferred embodiment of the present invention, and is not intended to limit the present invention to other forms. Any person skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes. However, any simple modification, equivalent change and modification made to the above embodiments based on the technical essence of the present invention without departing from the technical solution of the present invention still fall within the protection scope of the technical solution of the present invention.
Claims
1. A software performance optimization method under a domestic chip architecture, characterized in that It includes the following steps: S1: Subdivide the tasks of the power AI application and clarify the KPIs for improving application performance, including reducing latency and improving accuracy; S2: According to the task subdivision, decompose the AI application into independent service models to improve scalability, and use Kubernetes for containerized management to achieve automatic scaling and fault recovery; and utilize Istio to achieve fine-grained service routing and traffic management; S3: Adopt the domestic Kunpeng chip that supports artificial intelligence acceleration, design a dedicated acceleration library using custom instruction sets, and configure mixed-precision computing and high bandwidth according to application requirements; S4: Use the domestic framework of PaddlePaddle to optimize the independent service model to adapt to the Kunpeng chip; S5: Deploy the optimized independent service model at the grid edge node, use TensorRT to optimize the model, and improve the local processing ability; The specific content of S3 is as follows: S31: Develop a dedicated acceleration library using the instruction set provided by the Kunpeng chip, and design library functions for accelerating the execution of AI models. The library functions target the NPU of the chip and are optimized using low-level programming interfaces; S32: Utilize FP16 and FP32 mixed-precision computing during model training to maintain computing efficiency and result accuracy; and during the model inference stage, use INT8 quantization to reduce memory usage and accelerate model execution; S33: Through data preloading and asynchronous data stream DM technology, hide the memory access latency, and combine with the high-bandwidth characteristics of Kunpeng to optimize the data inflow and outflow channels to ensure the synchronization of data transmission and processing; S34: Implement intelligent memory allocation to ensure that critical data resides in the cache and reduce the memory round-trip of high-frequency data; allocate buffers for large-scale matrix operations of the input and output layers.
2. The software performance optimization method under a domestic chip architecture according to claim 1, characterized in that The specific content of decomposing the AI application into independent service models according to the task subdivision to improve scalability and using Kubernetes for containerized management to achieve automatic scaling and fault recovery is as follows: Design independent services for each functional module. The independent services communicate through APIs, and are independently deployed and scaled; Use Docker to encapsulate each independent service into a container to ensure environmental consistency and independence, create a container image containing the application and its dependencies, and upload it to the container image repository; Deploy a Kubernetes cluster to manage containerized applications, use configuration files to define Pod, ReplicaSet, and Deployment resources, and utilize the automatic scaling function of Kubernetes to dynamically adjust the number of Pods according to CPU / memory usage to achieve automatic scaling; Set up health checks and restart policies to ensure that faulty services can be automatically detected and recovered, and configure persistent volumes and persistent volume claims to ensure the persistence of data after the Pod restarts.
3. The software performance optimization method under a domestic chip architecture according to claim 2, characterized in that, The specific content of utilizing Istio to achieve fine-grained service routing and traffic management is as follows: Set up Istio in the Kubernetes cluster, and intercept and route communications by injecting sidecar proxies into each Pod; Configure the control plane components of Istio to manage the network traffic between services within the cluster; Utilize Istio's VirtualService and DestinationRule to define the routing rules between services, achieving traffic routing, load balancing, and fault recovery; Implement canary releases and blue-green deployments, allowing for gradual updates and rollbacks without affecting the overall system; Use Istio's monitoring and observability components to obtain service call chains and performance metrics to help identify performance bottlenecks.
4. A software performance optimization method under a domestic chip architecture according to claim 1, characterized in that The S32 specifically refers to: During the training process, maintain the 32-bit floating-point precision of the model parameters to accumulate accurate gradient updates; before each batch calculation, convert the FP32 model parameters to FP16 for fast calculation on the Tensor Core units of the GPU; Perform matrix operations and convolution operations in the forward and backward propagations using FP16; When calculating gradients in the backward propagation, perform operations using FP16 and then convert to FP32 to accumulate gradients; Use a dynamic loss scaling strategy to avoid numerical underflow problems that FP16 may introduce; During training or post-training quantization, convert the FP32 model parameters to INT8 format, collect the maximum and minimum values of activations and weights through a calibration dataset for calculating the quantization scaling factor, and use the quantized INT8 for matrix multiplication and convolution operations.
5. A method for optimizing software performance under a domestic chip architecture according to claim 4, characterized in that, The use of a dynamic loss scaling strategy to avoid numerical underflow problems that FP16 may introduce is specifically as follows: At the start of training, select a scaling factor S greater than a preset value to scale the loss to increase the numerical stability of the calculation; Calculate the original loss L in the forward propagation; Scale the original loss L by the scaling factor S for calculation: L scaled = L × S; In backpropagation, the gradient is calculated based on the scaled loss L scaled Calculate the gradient; Due to the scaled loss, the gradients are also magnified accordingly; The calculated gradient g scaled is magnified, and the original-scale gradient g is restored by anti-scaling. The gradient g scaled is divided by S: ; Perform a gradient check on the calculation. If there are non-numeric NaN or infinite values Inf in the gradients, Reduce the scaling factor S proportionally, discard the current update cycle, and restore to the previous stable model state; if there are no abnormalities in the gradient calculation: select to increase the scaling factor S proportionally to improve the operation accuracy.
6. The software performance optimization method under a domestic chip architecture according to claim 1, wherein The S34 specifically refers to: Analyze and identify data with an access frequency greater than a threshold in the model calculation as high-frequency data; Divide the memory into multiple levels, from the fastest and smallest-capacity L1 cache to the large-capacity main memory; Keep the high-frequency data in the L1 or L2 cache and utilize the high-speed cache characteristics of Kirin to ensure the access speed; Migrate the intermediate results and other large-scale memory usage scenarios to the main memory; pre-allocate dedicated buffers for the input / output layers and critical computing tasks to ensure that large-scale matrix operations do not cause bottlenecks due to improper memory allocation; implement a memory pool to dynamically allocate and release memory, avoid frequent malloc / free operations, and reduce memory fragmentation; Adjust the data layout method according to the data access pattern to maximize the cache hit rate.
7. A software performance optimization method under a domestic chip architecture according to claim 1, wherein, The S4 specifically refers to: Use the built-in model conversion tool in PaddlePaddle to convert the model into a form suitable for the Kirin architecture to make full use of the computing power of the hardware; Enable the AMP function of PaddlePaddle and use mixed-precision computing in training and inference to improve performance and reduce memory usage; For operations enabled on Kirin chips, analyze and customize the implementation of operators, and optimize key operators by leveraging the instruction sets and parallel acceleration features provided by the chips; Use the Operator Layer support of PaddlePaddle to perform operator-level performance tuning according to actual needs; Use PaddleSlim for model pruning to reduce unimportant neurons or channels, maintaining model effectiveness while reducing the computational burden; Use the ParallelExecutor of PaddlePaddle to perform model inference in parallel, leveraging NPU and DSP acceleration units; Implement layer fusion in the model execution pipeline to reduce memory interaction and cache switching, use pipelining execution technology to improve data flow efficiency, and adapt to the characteristics of the chip's hardware execution units.
8. A software performance optimization method under a domestic chip architecture according to claim 7, characterized in that, The implementation of layer fusion in the model execution pipeline to reduce memory interaction and cache switching is as follows: Utilize the static mode of PaddlePaddle to customize operators or use supported built-in operators to attempt the fusion process; identify connectable computational parts in the computational graph optimization process and perform corresponding operation fusion; Analyze the execution graph of the model, divide the model into multiple stages, and execute each stage in parallel on different computing resources; Adopt the producer-consumer mode to ensure synchronous data flow transmission between stages; In the pipeline, submit each execution stage in the form of an asynchronous task and achieve parallelism for each stage through the Executor of PaddlePaddle; Bind the computing resources to the types of the chip's hardware execution units of Kirin to ensure that each pipeline segment matches the optimal resources, and use a circular buffer to ensure synchronous and delay-free data transmission between execution stages.
9. A software performance optimization system under a domestic chip architecture, characterized in that, It includes a processor, a memory, and a computer program stored on the memory. When the processor executes the computer program, it specifically executes the steps in a software performance optimization method under a domestic chip architecture as described in any one of claims 1-8.
Citation Information
Patent Citations
Adaptive reasoning acceleration method and device, computer equipment and storage medium
CN114997401A
Domestic deep learning algorithm and operation equipment thereof
CN118397432A
Business application high-reliability optimization method and system based on micro-service architecture
CN119211013A