Method, device and equipment for optimizing reasoning operator of large model based on mercuric chloride chip
By performing block processing and operator optimization on the Ascend chip, efficiency and compatibility issues in large-scale model reasoning are resolved, and real-time performance and accuracy are improved, making it suitable for scenarios such as medical image analysis and high-frequency financial transactions.
Patent Information
- Application Number
- CN202510710884.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-09-16
AI Technical Summary
When using Ascend chips for large-model inference, existing technologies have problems such as low efficiency in medical image processing, poor compatibility in genetic sequence analysis, conflicts in multimodal data processing, response delays in high-frequency transactions in the financial sector, and loss of data segmentation information. This makes it difficult for the real-time and accuracy requirements of high-demand medical and financial applications to be met.
By processing the original input data in blocks, block data adapted to the computing units of the Ascend chip is generated, the memory access path of the matrix multiplication operator is optimized, the sliding step size and padding method of the convolution operation are adjusted, and multiple operators are fused into composite operators. Computing resources are dynamically allocated for collaborative execution, and a closed-loop optimization is formed by combining the verification feedback mechanism.
It significantly improves the real-time performance and accuracy of large-scale model reasoning, improves resource utilization, solves cross-scenario compatibility issues, and meets the application needs of high-demand fields such as medical image analysis and high-frequency financial transactions.
Smart Images

Figure CN120654816A_ABST
Abstract
Claims
1. A large model inference operator optimization method based on Ascend chips, characterized by: The following steps are involved: Based on the parallel computing capabilities of the Ascend chip, the original input data is processed in blocks to generate block data adapted to the chip's computing units. Based on the block data, combined with the memory hierarchy and computing core type of the Ascend chip, the memory access path of the matrix multiplication operator is optimized, and the sliding step size and padding method of the convolution operation are adjusted; Fusion of multiple operators executed continuously in large models into composite operators, optimization of the data storage process of composite operators, and coordinated execution through dynamic allocation of computing resources; The block calculation results after collaborative optimization are integrated and decoded, and the input data block strategy and operator execution parameters are adjusted through the verification feedback mechanism to form a closed-loop optimization.
2. The large model inference operator optimization method based on the Ascend chip according to claim 1 is characterized in that: The memory access path of the optimized matrix multiplication operator specifically includes: After dividing the input matrix into blocks by rows or columns, the high-frequency access matrix blocks are dynamically allocated to the high-speed memory layer based on the high-speed to low-speed access characteristics of the Ascend chip memory layer and the access frequency statistics of historical computing data. Plan the order in which the compute cores access matrix blocks to reduce memory access latency.
3. The large model inference operator optimization method based on the Ascend chip according to claim 1 is characterized in that: The sliding step size and filling method of adjusting the convolution operation specifically include: Convert the data format of the input feature map and convolution kernel to the format supported by the chip hardware acceleration unit; Based on the parallel processing capabilities of the Ascend chip's computing core, the sliding step size and padding method of the convolution kernel on the input feature map are adjusted to ensure that the convolution kernel sliding matches the layout of the chip's parallel task processing units.
4. The large model inference operator optimization method based on the Ascend chip according to claim 1 is characterized in that: The process of fusing multiple operators executed continuously in a large model into a composite operator and optimizing the data storage of the composite operator specifically includes: Fusion of convolutional layers, batch normalization layers, and activation function layers into a composite operator; This is achieved by storing the convolutional output data directly in memory locations where batch normalization and activation layers can read them efficiently.
5. The large model inference operator optimization method based on the Ascend chip according to claim 1, characterized in that: The collaborative execution achieved by dynamically allocating computing resources specifically includes: Dynamically allocate computing resources for convolutional layers, fully connected layers, and pooling layers based on the large model inference task process and the total amount of chip computing resources.
6. The large model inference operator optimization method based on the Ascend chip according to claim 1, characterized in that: The block processing of the original input data specifically includes: Divide the text sequence into predefined block lengths to generate continuous text blocks with consistent lengths; The image data is divided into spatial regions to generate sub-image blocks whose sizes are adapted to the parallel computing units of the chip.
7. The large model inference operator optimization method based on the Ascend chip according to claim 6, characterized in that: The adjustment of input data partitioning strategy and operator execution parameters through the verification feedback mechanism includes: Based on the accuracy index of the output results, identify the source of the error as data partitioning strategy or operator execution parameters; If the error is caused by the data block strategy, adjust the block length or image sub-block size; If the error is caused by operator execution parameters, re-optimize memory allocation or collaborative resource allocation ratio.
8. A large model inference operator optimization device based on Ascend chip, characterized in that: include: The input data preprocessing module is used to process the original input data in blocks based on the parallel computing capabilities of the Ascend chip, generating block data adapted to the chip's computing units. An operator execution optimization module, which is used to optimize the memory access path of the matrix multiplication operator based on the block data and in combination with the memory hierarchy and computing core type of the Ascend chip, and adjust the sliding step size and padding method of the convolution operation; The operator fusion and collaborative optimization module is used to fuse multiple operators executed continuously in a large model into composite operators, optimize the data storage process of composite operators, and achieve collaborative execution by dynamically allocating computing resources; The output result post-processing module is used to integrate and decode the block calculation results after collaborative optimization, and adjust the input data block strategy and operator execution parameters through the verification feedback mechanism to form a closed-loop optimization.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the large model inference operator optimization method based on the Ascend chip are implemented as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the large model inference operator optimization method based on the Ascend chip are implemented as described in any one of claims 1 to 7.
Citation Information
Cited By
Fusion operator execution method, electronic device, storage medium and program product
CN121029432A
Fusion operator execution method, electronic device, storage medium, and program product
CN121029432B
Distributed operator optimization method, artificial intelligence chip, computer equipment, readable storage medium and program product
CN121364958A
Optimization methods for distributed operators, artificial intelligence chips, computer equipment, readable storage media, and program products
CN121364958B