Software performance optimization method and system under domestic chip architecture

By performing task segmentation and model optimization on the domestic chip architecture Kirin, combined with Kubernetes and Istio management, the problem of edge computing resources in the power grid system is solved, and efficient computing and flexibility of power AI applications is achieved.

CN120029741AActive Publication Date: 2025-05-23FUJIAN YIRONG INFORMATION TECH
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510505194.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-05-23
Estimated Expiration
2045-04-22

AI Technical Summary

Technical Problem

In power grid systems, traditional centralized computing mode cannot meet the requirements of real-time and high reliability, especially when edge computing resources are limited, power AI applications face limitations in computing power, storage, energy consumption and communications.

Method used

It adopts the domestic chip architecture Kirin, which combines the PaddlePaddle framework for model optimization and deployment through task segmentation, independent service model, Kubernetes containerized management, Istio service routing and traffic management, customized instruction set acceleration library, hybrid precision computing and high bandwidth memory strategies.

Benefits of technology

It effectively improves the computing efficiency and flexibility of power AI applications, ensures efficient and high-quality computing in resource-constrained power grid edge environments, realizes automatic expansion and failure recovery, and reduces computing energy consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029741A_ABST
    Figure CN120029741A_ABST
Patent Text Reader

Abstract

The invention relates to a software performance optimization method and system under a domestic chip architecture, and the method comprises the following steps: S1, carrying out the task subdivision of an electric power AI application, and determining the KPI of the application performance improvement, including the delay reduction and the precision improvement; s2, decomposing the AI application into independent service models according to task subdivision, improving the expandability, and performing containerization management by using Kubernetes to realize automatic expansion and fault recovery; the Istio is utilized to realize refined service routing and flow management; s3, a domestic chip Kylin supporting artificial intelligence acceleration is adopted, a special acceleration library is designed through a user-defined instruction set, and mixing precision calculation and high bandwidth are configured according to application requirements; s4, using a domestic framework of PadlePaddle to optimize the independent service model so as to adapt to a domestic chip kylin; and S5, deploying the optimized independent service model at the edge node of the power grid, and optimizing the model by using TensorRT to improve the local processing capability. According to the method, the operation efficiency and flexibility of power AI application are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of software optimization, and in particular to a method and system for optimizing software performance under a domestic chip architecture. Background Art

[0002] In modern power grid systems, with the popularization of IoT devices and the application of smart grid technology, a large amount of real-time data is generated. This data plays an important role in monitoring, predicting and fault detection of power grids. However, traditional centralized computing models often cannot meet the requirements of real-time and high reliability of power grids. Therefore, more and more application scenarios are beginning to adopt edge computing and artificial intelligence technologies to solve these challenges.

[0003] With the development of smart grids and the popularization of IoT devices, power AI applications are facing multiple challenges such as a surge in data volume, increased real-time requirements, and reliability requirements in complex environments. Especially in the case of resource constraints in edge computing, power AI applications face a series of challenges in computing, including limitations in computing power, storage, energy consumption, and communication. Summary of the invention

[0004] In order to solve the above problems, the purpose of the present invention is to provide a software performance optimization method and system under a domestic chip architecture, so as to effectively improve the computing efficiency and flexibility of power AI applications.

[0005] To achieve the above object, the present invention adopts the following technical solutions: A software performance optimization method under a domestic chip architecture includes the following steps S1: Segment the tasks of power AI applications and identify the KPIs for improving application performance, including reducing latency and improving accuracy. S2: Decompose AI applications into independent service models based on task segmentation to improve scalability, and use Kubernetes for containerized management to achieve automatic expansion and fault recovery; and use Istio to achieve refined service routing and traffic management; S3: Adopt the domestic Kirin chip that supports AI acceleration, design a dedicated acceleration library using a custom instruction set, and configure mixed-precision computing and high bandwidth according to application requirements; S4: Use PaddlePaddle's domestic framework to optimize the independent service model to adapt to the domestic chip Kirin; S5: Deploy the optimized independent service model at the edge nodes of the power grid and use TensorRT to optimize the model to improve local processing capabilities.

[0006] Furthermore, AI applications are decomposed into independent service models based on task segmentation to improve scalability, and containerized management using Kubernetes is used to achieve automatic expansion and fault recovery. Specifically: Design independent services for each functional module. Independent services communicate through APIs and are independently deployed and expanded. Use Docker to encapsulate each independent service as a container to ensure environmental consistency and independence, create a container image containing the application and its dependencies, and upload it to the container image repository; Deploy Kubernetes clusters to manage containerized applications, use configuration files to define Pod, ReplicaSet, and Deployment resources, and use Kubernetes's auto-scaling feature to dynamically adjust the number of Pods based on CPU / memory usage to achieve automatic expansion; Set up health checks and restart policies to ensure that failed services can be automatically detected and recovered, and configure persistent volumes and persistent volume declarations to ensure that data is persistent after a Pod restart.

[0007] Furthermore, Istio is used to implement refined service routing and traffic management, as follows: Set up Istio in the Kubernetes cluster and inject the sidecar proxy into each Pod to intercept and route communications. Configure Istio's control plane components to manage network traffic between services in the cluster; Use Istio's VirtualService and DestinationRule to define routing rules between services to implement traffic routing, load balancing, and fault recovery; Implement grayscale releases and blue-green deployments, allowing for gradual updates and rollbacks without affecting the overall system; Use Istio's monitoring and observability components to obtain service call links and performance metrics to help identify performance bottlenecks.

[0008] Furthermore, S3 is specifically: S31: Develop a dedicated acceleration library using the instruction set provided by the Kirin chip, and design library functions for accelerating the execution of AI models. The library functions are targeted at the NPU of the chip and use low-level programming interfaces for performance optimization. S32: Use FP16 and FP32 mixed precision calculations in model training to maintain computational efficiency and result accuracy; and use INT8 quantization in the model inference phase to reduce memory usage and accelerate model execution; S33: Through data preloading and asynchronous data flow DM technology, memory access latency is hidden. Combined with the high bandwidth characteristics of Kylin, data inflow and outflow channels are optimized to ensure synchronization of data transmission and processing. S34: Implement intelligent memory allocation to ensure that critical data resides in the cache and reduce memory round trips for high-frequency data; allocate buffers for large-scale matrix operations in the input and output layers.

[0009] Furthermore, S32 is specifically: During training, the 32-bit floating point precision of model parameters is maintained to accumulate accurate gradient updates; before each batch calculation, the FP32 model parameters are converted to FP16 for fast calculation on the GPU's Tensor Core unit; Matrix operations and convolution operations in forward and backward propagation are performed using FP16.

[0010] When calculating the gradient in back propagation, FP16 is used for calculation and then converted to FP32 for accumulating gradients; Use a dynamic loss scaling strategy to avoid numerical underflow issues that may be introduced by FP16; During or after training, quantize the model to convert FP32 parameters to INT8 format, collect the maximum and minimum values ​​of activations and weights through the calibration dataset for calculation of quantization scaling factors, and use quantized INT8 for matrix multiplication and convolution operations.

[0011] Furthermore, a dynamic loss scaling strategy is used to avoid the numerical underflow problem that FP16 may introduce, as follows: When starting training, choose a scaling factor S that is larger than the preset value to scale the loss to increase the numerical stability of the calculation; Calculate the original loss L in the forward propagation; The original loss L is scaled by the scaling factor S: L scaled =L×S; In the backward propagation, according to the scaled loss L scaled Calculate the gradient; Due to the scaling loss, the gradients are also amplified accordingly; The calculated gradient g scaled is amplified and restored to its original scale by inverse scaling. scaled Divide by S:

[0012] Perform a gradient check on the calculation. If there is a non-number NaN or infinite value Inf in the gradient, reduce the scaling factor S proportionally, discard the current update cycle and restore to the last stable model state; if there is no abnormality in the gradient calculation: choose to increase the scaling factor S proportionally to improve the calculation accuracy.

[0013] Furthermore, S34 is specifically: Analyze and identify data with access frequency greater than a threshold in model calculations as high-frequency data, such as intermediate tensors; Divide memory into multiple levels, from the fastest and smallest L1 cache to the large main memory; Store high-frequency data in L1 or L2 cache, and use the high-speed cache feature of Kylin to ensure access speed; Migrate intermediate results and other large-scale memory usage scenarios to main memory; pre-allocate dedicated buffers for input and output layers and key computing tasks to ensure that large-scale matrix operations do not cause bottlenecks due to improper memory allocation; implement memory pools, dynamically allocate and release memory, avoid frequent malloc / free operations, and reduce memory fragmentation; Adjust the data layout according to the data access pattern to maximize the cache hit rate.

[0014] Furthermore, S4 is specifically: Use PaddlePaddle's built-in model conversion tool to convert the model to a format suitable for the Kirin architecture in order to fully utilize the computing power of the hardware; Enable PaddlePaddle's AMP feature to use mixed precision computing (FP16 + FP32) in training and inference to improve performance and reduce memory usage; Analyze and customize the implementation of operators for the operations enabled on the Kirin chip, and optimize key operators using the instruction set and parallel acceleration features provided by the chip; Use PaddlePaddle's Operator Layer support to implement operator-level performance tuning based on actual needs; Use PaddleSlim to prune the model, reduce unimportant neurons or channels, maintain model effectiveness and reduce computational burden; Use PaddlePaddle's ParallelExecutor to perform model reasoning in parallel, taking advantage of NPU and DSP acceleration units; Implement layer fusion in the model execution link, reduce memory interaction and cache switching, use pipeline execution technology, improve data flow efficiency, and adapt to the hardware execution unit characteristics of the chip.

[0015] Furthermore, layer fusion is implemented in the model execution link to reduce memory interaction and cache switching as follows: Use PaddlePaddle's static mode to customize operators or use supported built-in operators to try the fusion process; identify the connectable operation parts in the calculation graph optimization process and perform the corresponding operation fusion; Analyze the execution graph of the model and divide the model into multiple stages. Each stage is executed in parallel on different computing resources. The producer-consumer model is adopted to ensure that the data flow at each stage can be transmitted synchronously; In the pipeline, each execution stage is submitted as an asynchronous task, and the parallelism of each stage is achieved through PaddlePaddle's Executor; Computing resources are bound to the hardware execution unit type of the Kirin chip to ensure that each pipeline stage matches the best resources, and a ring buffer is used to ensure that data is synchronized and passed without delay between execution stages.

[0016] A software performance optimization system under a domestic chip architecture includes a processor, a memory, and a computer program stored in the memory. When the processor executes the computer program, it specifically performs the steps in the software performance optimization method under a domestic chip architecture as described above.

[0017] The present invention has the following beneficial effects: 1. The present invention improves the computing efficiency and flexibility of power AI applications through the in-depth combination of software and hardware, from framework to underlying chip, from application decomposition to deployment management, so that applications can maintain efficient and high-quality computing in the edge environment of the power grid where resources are limited and high reliability and real-time computing are required; 2. The present invention decomposes power AI applications into independent service models, improves the flexibility and scalability of the system, realizes dynamic resource scheduling through Kubernetes container management, promotes automatic expansion and fault recovery, enables independent development, deployment, and expansion of different models, and facilitates maintenance and optimization. It also manages traffic and security between services in a refined manner through Service Mesh; 3. The present invention makes full use of the hardware acceleration characteristics of the Kirin chip through a custom instruction set and a dedicated acceleration library, implements mixed-precision computing and high-bandwidth memory strategies to maximize performance, effectively improves model computing efficiency, shortens inference time, and reduces computing energy consumption, thereby better supporting the efficient operation of various microservices on domestic hardware; and combined with PaddlePaddle's optimization capabilities on the domestic Kirin chip, efficient model compression and performance optimization can be achieved, enabling deep learning models to operate efficiently on terminal devices with limited resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 The figure is a flow chart of the method of the present invention. DETAILED DESCRIPTION

[0019] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments: refer to Figure 1 In this embodiment, a method for optimizing software performance under a domestic chip architecture includes the following steps: S1: Perform task segmentation for power AI applications (for power system AI applications, segmentation into load forecasting models, fault detection models, energy consumption optimization models, etc.), and define the KPIs for application performance improvement, including reducing latency and improving accuracy; In this embodiment, load forecasting: predict the future power demand of the power grid and optimize power generation and distribution. KPI: Accuracy: the goal is to reduce the mean square error (MSE) below a certain threshold; Response delay: the update delay of the prediction result is less than 5 seconds.

[0020] Fault detection: Real-time detection and identification of anomalies and faults in the power grid. KPI: Identification accuracy: higher than 98%; Response time: less than 1 second, ensuring rapid response.

[0021] Energy consumption optimization: optimize the power consumption of the power system and improve overall energy efficiency. KPI: Energy efficiency improvement: the system energy efficiency improvement target is 5%; cost savings: the operating cost reduction target is 2%.

[0022] S2: Decompose AI applications into independent service models based on task segmentation to improve scalability, and use Kubernetes for containerized management to achieve automatic expansion and fault recovery; and use Istio to achieve refined service routing and traffic management; S3: Adopt the domestic Kirin chip that supports AI acceleration, design a dedicated acceleration library using a custom instruction set, and configure mixed-precision computing and high bandwidth according to application requirements; S4: Use PaddlePaddle's domestic framework to optimize the independent service model to adapt to the domestic chip Kirin; S5: Deploy the optimized independent service model at the edge node of the power grid and use TensorRT to optimize the model to improve local processing capabilities; In this embodiment, the Python API or command line tool (trtexec) of TensorRT is used to convert the model to the TensorRT engine. Specify optimization parameters such as batch size, maximum workspace size, and target accuracy. Save the TensorRT engine for fast loading and startup on edge devices. Develop an inference service interface on the edge node and expose the model inference service through methods such as Rest API or RPC. Integrate sensor data acquisition and predictive reasoning in combination with specific power grid application scenarios. Utilize the parallel capabilities of edge devices and set up parallel execution plans to maximize the use of the asynchronous execution features of TensorRT. Optimize the preprocessing and postprocessing steps of input data and use them as part of the pipeline to reduce overall latency. Using TensorRT to optimize the model can greatly improve the execution efficiency of the model at the edge node of the power grid and ensure fast and reliable inference results under resource-constrained conditions. This helps to perform real-time monitoring, analysis, and response in a smart grid environment to improve the efficiency and reliability of the overall power grid.

[0023] Preferably, use Prometheus to build a real-time monitoring system to monitor system performance and model inference time; implement A / B testing and user feedback collection mechanisms, and use data-driven improvements to model algorithms and system bottlenecks.

[0024] In this embodiment, AI applications are decomposed into independent service models according to task segmentation to improve scalability, and Kubernetes is used for containerized management to achieve automatic expansion and fault recovery. Specifically: Design independent services for each functional module (such as load forecasting, fault detection, and energy consumption optimization). Independent services communicate through APIs and are independently deployed and expanded. Use Docker to encapsulate each independent service as a container to ensure environmental consistency and independence, create a container image containing the application and its dependencies, and upload it to the container image repository; Deploy Kubernetes clusters to manage containerized applications, use configuration files (YAML) to define Pod, ReplicaSet, and Deployment resources, and use Kubernetes' automatic scaling features, such as Horizontal PodAutoscaler (HPA), to dynamically adjust the number of Pods based on CPU / memory usage to achieve automatic expansion; Set up health checks and restart policies to ensure that failed services can be automatically detected and recovered, and configure persistent volumes and persistent volume claims to ensure that data is persistent after a Pod restart.

[0025] In this embodiment, Istio is used to implement refined service routing and traffic management, as follows: Set up Istio in the Kubernetes cluster and inject the sidecar proxy (Envoy) into each Pod to intercept and route communications. Configure Istio's control plane components (such as pilot and mixer) to manage network traffic between services in the cluster; Use Istio's VirtualService and DestinationRule to define routing rules between services to implement traffic routing, load balancing, and fault recovery; Implement grayscale releases and blue-green deployments, allowing for gradual updates and rollbacks without affecting the overall system; Use Istio's monitoring and observability components (such as Kiali and Jaeger) to obtain service call links and performance indicators to help identify performance bottlenecks.

[0026] In this embodiment, S3 is specifically: S31: Develop a dedicated acceleration library using the instruction set provided by the Kirin chip, especially optimizing common deep learning operations (such as matrix multiplication and convolution operations); and design library functions for accelerating the execution of AI models. The library functions use low-level programming interfaces for performance optimization for the chip's NPU. S32: Use FP16 and FP32 mixed precision calculations in model training to maintain computational efficiency and result accuracy; and use INT8 quantization in the model inference phase to reduce memory usage and accelerate model execution; S33: Through data preloading and asynchronous data flow DM technology, memory access latency is hidden. Combined with the high bandwidth characteristics of Kylin, data inflow and outflow channels are optimized to ensure synchronization of data transmission and processing. S34: Implement smart memory allocation to ensure that critical data (such as intermediate tensors) reside in cache and reduce memory round trips for high-frequency data; allocate buffers for large-scale matrix operations input to Output layers.

[0027] In this embodiment, S32 is specifically: During training, the 32-bit floating point precision of model parameters is maintained to accumulate accurate gradient updates; before each batch calculation, the FP32 model parameters are converted to FP16 for fast calculation on the Tensor Core unit of the GPU; Matrix operations and convolution operations in forward and backward propagation are performed using FP16.

[0028] When calculating the gradient in back propagation, FP16 is used for calculation and then converted to FP32 for accumulating gradients; Use a dynamic loss scaling strategy to avoid numerical underflow issues that may be introduced by FP16; During or after training, quantize the model to convert FP32 parameters to INT8 format. Use the calibration dataset to collect the maximum and minimum values ​​of activations and weights for calculating the quantization scaling factor. Use the quantized INT8 for matrix multiplication and convolution operations, and run on hardware that supports this calculation, such as the accelerator of the Kirin chip.

[0029] In this embodiment, a dynamic loss scaling strategy is used to avoid the numerical underflow problem that may be introduced by FP16, as follows: When starting training, choose a scaling factor S larger than the preset value (e.g. 1024 or higher) to scale the loss to increase numerical stability of the calculation; Calculate the original loss L in the forward propagation; The original loss L is scaled by the scaling factor S: L scaled =L×S; In the backward propagation, according to the scaled loss L scaled Calculate the gradient; Due to the scaling loss, the gradients are also amplified accordingly; The calculated gradient g scaled is amplified and restored to its original scale by inverse scaling. scaled Divide by S:

[0030] Perform a gradient check on the calculation. If there is a non-number NaN or infinite value Inf in the gradient, reduce the scaling factor S proportionally (for example, reduce it to half of the previous value), discard the current update cycle and restore to the last stable model state; if there is no abnormality in the gradient calculation: choose to proportionally increase the scaling factor S to improve the calculation accuracy (for example, after each certain number of steps).

[0031] In this embodiment, S34 is specifically: Analyze and identify data with access frequency greater than a threshold in model calculations as high-frequency data, such as intermediate tensors; Divide memory into multiple levels, from the fastest and smallest L1 cache to the large main memory; Store high-frequency data in L1 or L2 cache, and use the high-speed cache feature of Kylin to ensure access speed; Migrate intermediate results and other large-scale memory usage scenarios to main memory; pre-allocate dedicated buffers for input and output layers and key computing tasks to ensure that large-scale matrix operations do not cause bottlenecks due to improper memory allocation; implement memory pools to dynamically allocate and release memory, avoid frequent malloc / free operations, and reduce memory fragmentation; Adjust the layout of data based on data access patterns (such as row-major or column-major) to maximize cache hit rates.

[0032] In this embodiment, S4 is specifically: Use PaddlePaddle's built-in model conversion tool to convert the model to a format suitable for the Kirin architecture in order to fully utilize the computing power of the hardware; Enable PaddlePaddle's AMP feature to use mixed precision computing (FP16 + FP32) in training and inference to improve performance and reduce memory usage; Analyze and customize the implementation of operators for the operations enabled on the Kirin chip, and optimize key operators using the instruction set and parallel acceleration features provided by the chip; Use PaddlePaddle's Operator Layer support to implement operator-level performance tuning based on actual needs; Use PaddleSlim to prune the model, reduce unimportant neurons or channels, maintain model effectiveness and reduce computational burden; Use PaddlePaddle's ParallelExecutor to perform model reasoning in parallel, taking advantage of NPU and DSP acceleration units; Implement layer fusion in the model execution link, reduce memory interaction and cache switching, use pipeline execution technology, improve data flow efficiency, and adapt to the hardware execution unit characteristics of the chip.

[0033] Preferably, layer fusion is implemented in the model execution link to reduce memory interaction and cache switching, as follows: Use PaddlePaddle's static mode to customize operators or use supported built-in operators to try the fusion process; identify the connectable operation parts in the calculation graph optimization process and perform the corresponding operation fusion; Analyze the execution graph of the model and divide the model into multiple stages. Each stage is executed in parallel on different computing resources (such as CPU and NPU). The producer-consumer model is adopted to ensure that the data flow of each stage can be transmitted synchronously; (each stage can be both a producer and a consumer: Producer: responsible for calculation and providing data to the next stage. Consumer: receives data from the previous stage for processing) In the pipeline, each execution stage is submitted as an asynchronous task, and the parallelism of each stage is achieved through PaddlePaddle's Executor; Computing resources are bound to the hardware execution unit type of the Kirin chip to ensure that each pipeline stage matches the best resources, and a ring buffer is used to ensure that data is synchronized and passed without delay between execution stages.

[0034] It will be appreciated by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0035] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0036] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0037] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0038] The above is only a preferred embodiment of the present invention, and does not limit the present invention in other forms. Any technician familiar with the profession may use the above disclosed technical content to change or modify it into an equivalent embodiment with equivalent changes. However, any simple modification, equivalent change and modification made to the above embodiment according to the technical essence of the present invention without departing from the technical solution of the present invention still belongs to the protection scope of the technical solution of the present invention.

Claims

1. A software performance optimization method under a domestic chip architecture, characterized in that: The following steps are included S1: Segment the tasks of power AI applications and identify the KPIs for improving application performance, including reducing latency and improving accuracy. S2: Decompose AI applications into independent service models based on task segmentation to improve scalability, and use Kubernetes for containerized management to achieve automatic expansion and fault recovery; and use Istio to achieve refined service routing and traffic management; S3: Adopt the domestic Kirin chip that supports AI acceleration, design a dedicated acceleration library using a custom instruction set, and configure mixed-precision computing and high bandwidth according to application requirements; S4: Use PaddlePaddle's domestic framework to optimize the independent service model to adapt to the domestic chip Kirin; S5: Deploy the optimized independent service model at the edge nodes of the power grid and use TensorRT to optimize the model to improve local processing capabilities.

2. According to claim 1, the software performance optimization method under the domestic chip architecture is characterized in that: According to the task segmentation, AI applications are decomposed into independent service models to improve scalability, and Kubernetes is used for container management to achieve automatic expansion and fault recovery. Specifically: Design independent services for each functional module. Independent services communicate through APIs and are independently deployed and expanded. Use Docker to encapsulate each independent service as a container to ensure environmental consistency and independence, create a container image containing the application and its dependencies, and upload it to the container image repository; Deploy Kubernetes clusters to manage containerized applications, use configuration files to define Pod, ReplicaSet, and Deployment resources, and use Kubernetes's auto-scaling feature to dynamically adjust the number of Pods based on CPU / memory usage to achieve automatic expansion; Set up health checks and restart policies to ensure that failed services can be automatically detected and recovered, and configure persistent volumes and persistent volume declarations to ensure that data is persistent after a Pod restart.

3. According to claim 2, the software performance optimization method under the domestic chip architecture is characterized in that: The use of Istio to achieve refined service routing and traffic management is as follows: Set up Istio in the Kubernetes cluster and inject the sidecar proxy into each Pod to intercept and route communications. Configure Istio's control plane components to manage network traffic between services in the cluster; Use Istio's VirtualService and DestinationRule to define routing rules between services to implement traffic routing, load balancing, and fault recovery; Implement grayscale releases and blue-green deployments, allowing for gradual updates and rollbacks without affecting the overall system; Use Istio's monitoring and observability components to obtain service call links and performance metrics to help identify performance bottlenecks.

4. According to claim 1, the software performance optimization method under the domestic chip architecture is characterized in that: The S3 is specifically: S31: Develop a dedicated acceleration library using the instruction set provided by the Kirin chip, and design library functions for accelerating the execution of AI models. The library functions are targeted at the NPU of the chip and use low-level programming interfaces for performance optimization. S32: Use FP16 and FP32 mixed precision calculations in model training to maintain computational efficiency and result accuracy; and use INT8 quantization in the model inference phase to reduce memory usage and accelerate model execution; S33: Through data preloading and asynchronous data flow DM technology, memory access latency is hidden. Combined with the high bandwidth characteristics of Kylin, data inflow and outflow channels are optimized to ensure synchronization of data transmission and processing. S34: Implement intelligent memory allocation to ensure that critical data resides in the cache and reduce memory round trips for high-frequency data; allocate buffers for large-scale matrix operations in the input and output layers.

5. According to claim 4, the method for optimizing software performance under a domestic chip architecture is characterized in that: The S32 is specifically: During training, the 32-bit floating point precision of model parameters is maintained to accumulate accurate gradient updates; before each batch calculation, the FP32 model parameters are converted to FP16 for fast calculation on the Tensor Core unit of the GPU; Use FP16 to perform matrix operations and convolution operations in forward and backward propagation; When calculating the gradient in back propagation, FP16 is used for calculation and then converted to FP32 for accumulating gradients; Use a dynamic loss scaling strategy to avoid numerical underflow issues that may be introduced by FP16; During or after training, quantize the model to convert FP32 parameters to INT8 format, collect the maximum and minimum values ​​of activations and weights through the calibration dataset for calculation of quantization scaling factors, and use quantized INT8 for matrix multiplication and convolution operations.

6. The method for optimizing software performance under a domestic chip architecture according to claim 5, characterized in that: The dynamic loss scaling strategy is used to avoid the numerical underflow problem that may be introduced by FP16, as follows: When starting training, choose a scaling factor S that is larger than the preset value to scale the loss to increase the numerical stability of the calculation; Calculate the original loss L in the forward propagation; The original loss L is scaled by the scaling factor S: L scaled =L×S; In the backward propagation, according to the scaled loss L scaled Calculate the gradient; Due to the scaling loss, the gradients are also amplified accordingly; The calculated gradient g scaled is amplified and restored to its original scale by inverse scaling. scaled Divide by S: ; Perform a gradient check on the calculation. If there is a non-number NaN or infinite value Inf in the gradient, reduce the scaling factor S proportionally, discard the current update cycle and restore to the last stable model state; if there is no abnormality in the gradient calculation: choose to increase the scaling factor S proportionally to improve the calculation accuracy.

7. The method for optimizing software performance under a domestic chip architecture according to claim 4, characterized in that: The S34 is specifically: Analyze and identify data with access frequency greater than a threshold in model calculation as high-frequency data; Divide memory into multiple levels, from the fastest and smallest L1 cache to the large main memory; Store high-frequency data in L1 or L2 cache and use the high-speed cache feature of Kylin to ensure access speed; Migrate intermediate results and other large-scale memory usage scenarios to main memory; pre-allocate dedicated buffers for input and output layers and key computing tasks to ensure that large-scale matrix operations do not cause bottlenecks due to improper memory allocation; implement memory pools, dynamically allocate and release memory, avoid frequent malloc / free operations, and reduce memory fragmentation; Adjust the data layout according to the data access pattern to maximize the cache hit rate.

8. The method for optimizing software performance under a domestic chip architecture according to claim 1, characterized in that: The S4 is specifically: Use PaddlePaddle's built-in model conversion tool to convert the model to a format suitable for the Kirin architecture in order to fully utilize the computing power of the hardware; Enable PaddlePaddle's AMP feature to use mixed precision computing in training and inference to improve performance and reduce memory usage; Analyze and customize the implementation of operators for the operations enabled on the Kirin chip, and optimize key operators using the instruction set and parallel acceleration features provided by the chip; Use PaddlePaddle's Operator Layer support to implement operator-level performance tuning based on actual needs; Use PaddleSlim to prune the model, reduce unimportant neurons or channels, maintain model effectiveness and reduce computational burden; Use PaddlePaddle's ParallelExecutor to perform model reasoning in parallel, taking advantage of NPU and DSP acceleration units; Implement layer fusion in the model execution link, reduce memory interaction and cache switching, use pipeline execution technology, improve data flow efficiency, and adapt to the hardware execution unit characteristics of the chip.

9. The method for optimizing software performance under a domestic chip architecture according to claim 8, characterized in that: The implementation of layer fusion in the model execution link to reduce memory interaction and cache switching is as follows: Use PaddlePaddle's static mode to customize operators or use supported built-in operators to try the fusion process; identify the connectable operation parts in the calculation graph optimization process and perform the corresponding operation fusion; Analyze the execution graph of the model and divide the model into multiple stages, each of which is executed in parallel on different computing resources; The producer-consumer model is adopted to ensure that the data flow at each stage can be transmitted synchronously; In the pipeline, each execution stage is submitted as an asynchronous task, and the parallelism of each stage is achieved through PaddlePaddle's Executor; Computing resources are bound to the hardware execution unit type of the Kirin chip to ensure that each pipeline stage matches the best resources, and a ring buffer is used to ensure that data is synchronized and passed without delay between execution stages.

10. A software performance optimization system under a domestic chip architecture, characterized in that: It includes a processor, a memory and a computer program stored in the memory. When the processor executes the computer program, it specifically performs the steps in the software performance optimization method under a domestic chip architecture as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Adaptive reasoning acceleration method and device, computer equipment and storage medium

    CN114997401A

  • Domestic deep learning algorithm and operation equipment thereof

    CN118397432A

  • Business application high-reliability optimization method and system based on micro-service architecture

    CN119211013A

  • Elastic concurrent AI model optimization productivity acceleration middle table

    CN119718639A

  • Inference service system based on kubernetes

    WO2021238251A1