Performance improvement method and electronic equipment
By loading the large AI model into the cache of the preset accelerator and utilizing its high-performance data processing and low-latency data transmission capabilities, the problem of insufficient computing power of the large AI model is solved, energy consumption and latency are reduced, and data processing efficiency is improved.
Patent Information
- Application Number
- CN202510739627.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-09-19
AI Technical Summary
Traditional computer hardware cannot meet the huge computing requirements of large AI models, resulting in slower computing speeds and low efficiency. How can we improve the computing power of large AI models and reduce energy consumption and delays in data processing?
Load the large AI model into the cache of the preset accelerator, utilize the high-performance data processing and low-latency data transmission capabilities of the preset accelerator, and perform data processing by storing historical data and processing results as a reference to reduce redundant calculations.
It improves the computing power of large AI models, reduces energy consumption and delays in data processing, and improves data processing efficiency.
Smart Images

Figure CN120671749A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a performance improvement method and electronic device. Background Art
[0002] With the rapid development of artificial intelligence (AI) technology, the computing power of large AI models has become increasingly important. With the surge in data volume and the increasing complexity of data types, large AI models require powerful computing power to efficiently process and analyze data. This powerful computing power not only improves the accuracy of model predictions but also accelerates model training and inference, enabling AI applications to respond and adapt to new data more quickly. However, traditional computer hardware can no longer meet the massive computing demands of large AI models, resulting in slow computing speeds and low efficiency.
[0003] Therefore, how to improve the computing power of large AI models and reduce energy consumption and delays in data processing is an urgent problem that needs to be solved. Summary of the Invention
[0004] This application provides a performance improvement method and electronic device to at least solve the problem of how to improve the computing power of large AI models and reduce energy consumption and delay in data processing in related technologies.
[0005] This application provides a method for improving performance, including:
[0006] Loading the obtained target model to be optimized into a cache of a preset accelerator to obtain a loaded target model, wherein the preset accelerator includes at least high-performance data processing attributes and low-latency data transmission attributes;
[0007] Inputting the acquired first data to be processed into the loaded target model for data processing to obtain a first processing result, and storing the first processing result and the first data to be processed in a cache of a preset accelerator, wherein the first data to be processed is the starting data input into the loaded target model;
[0008] By using the loaded target model, data processing is performed on the target data to be processed based on a plurality of historical data to be processed in a cache of a preset accelerator and a plurality of historical processing results corresponding to each of the plurality of historical data to be processed, to obtain a target processing result, wherein the target data to be processed is data input into the loaded target model after the plurality of historical data to be processed;
[0009] The target processing result and the target data to be processed are stored in a cache of a preset accelerator.
[0010] The present application also provides a device for improving performance, comprising:
[0011] A loading unit, configured to load the acquired target model to be optimized into a cache of a preset accelerator to obtain a loaded target model, wherein the preset accelerator includes at least high-performance data processing attributes and low-latency data transmission attributes;
[0012] a processing unit, configured to input the acquired first data to be processed into the loaded target model for data processing to obtain a first processing result;
[0013] a storage unit, configured to store the first processing result and first data to be processed in a cache of a preset accelerator, wherein the first data to be processed is the starting data input into the loaded target model;
[0014] The processing unit is further configured to, through the loaded target model, perform data processing on the target data to be processed based on a plurality of historical data to be processed in a cache of a preset accelerator and a plurality of historical processing results corresponding to each of the plurality of historical data to be processed, to obtain a target processing result, wherein the target data to be processed is data input into the loaded target model after the plurality of historical data to be processed;
[0015] The storage unit is further configured to store the target processing result and the target data to be processed in a cache of the preset accelerator.
[0016] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned performance improvement methods when executing the computer program.
[0017] The present application also provides a computer-readable storage medium, in which a computer program is stored, wherein when the computer program is executed by a processor, the steps of any of the above-mentioned performance improvement methods are implemented.
[0018] The present application also provides a computer program product, including a computer program, which implements the steps of any of the above-mentioned performance improvement methods when the computer program is executed by a processor.
[0019] The performance improvement method and electronic device of the present application can apply the high-performance data processing capabilities and low-latency data transmission capabilities of the preset accelerator to the improvement of the model's computing capabilities by loading the model into the preset accelerator. By storing historical data and processing results in the preset accelerator, when new data is processed through the model, data processing can be performed based on the historical data and processing results as a reference. If the data is repeated or similar, there is no need to reprocess the data, which can improve data processing efficiency. The cache technology is applied to the preset accelerator and model. While the computing power is improved, the delay and energy consumption of data access are reduced. Therefore, it can solve the technical problem of how to improve the computing power of large AI models and reduce energy consumption and delay in the data processing process, thereby achieving the technical effect of improving the computing power of large AI models and reducing energy consumption and delay in the data processing process. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0021] Figure 1 A flowchart of a performance improvement method provided in an embodiment of the present application;
[0022] Figure 2 A schematic diagram of the structure of a performance-enhancing device provided in an embodiment of the present application;
[0023] Figure 3 A schematic structural diagram of another performance-enhancing device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0024] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0025] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.
[0026] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0027] Figure 1 A flowchart of a method for improving performance provided in an embodiment of the present application is provided, and the method is described in detail in conjunction with the execution process of the method for improving performance.
[0028] like Figure 1 As shown, the performance improvement methods include:
[0029] Step 101 : Loading the acquired target model to be optimized into a cache of a preset accelerator to obtain a loaded target model, wherein the preset accelerator at least includes high-performance data processing attributes and low-latency data transmission attributes.
[0030] In the embodiments of the present application, the target model to be optimized refers to a large artificial intelligence model (such as a large language model (LLM) or a large visual model of the Transformer architecture) that needs to improve computing performance, which is characterized by a large number of parameters and intensive calculations. The preset accelerator refers to a device with hardware-level acceleration capabilities, such as an eXpress Data Path (XDP) accelerator card. For ease of understanding, the preset accelerators in this application are all explained using the XDP accelerator card as an example.
[0031] The preset accelerator has at least two key attributes: high-performance data processing attribute refers to the inherent ability to parallelize data calculations through dedicated hardware pipelines, which can significantly improve the efficiency of AI operations; low-latency data transmission attribute refers to the ability to bypass the operating system kernel and interact directly with the network card / storage, which can eliminate the memory copy delay in traditional data transmission.
[0032] The accelerator's cache is a high-speed storage area built into the accelerator, such as High Bandwidth Memory (HBM) or Static Random-Access Memory (SRAM). Access to the accelerator's cache is much faster than main memory. The loading process uses the Direct Memory Access (DMA) engine to migrate the parameters of the target model to be optimized from host memory to the accelerator's cache, ultimately creating a loaded target model—a model instance that has been hardware-adapted and can be executed directly on the accelerator.
[0033] In step 102, the acquired first data to be processed is input into the loaded target model for data processing to obtain a first processing result, and the first processing result and the first data to be processed are stored in the cache of a preset accelerator, wherein the first data to be processed is the starting data input into the loaded target model.
[0034] In the embodiments of this application, the data processing process for initial data is represented by the first batch of input data to be processed by the loaded target model. This initial data is unique in that it requires initializing the model's computation graph. During data processing, the data is directly injected into the loaded target model, utilizing the XDP accelerator's zero-copy mechanism to avoid transfer to host memory.
[0035] Data processing is the process by which the loaded target model performs inference (e.g., neural network layer calculations) on the input data to generate processing results (e.g., classification probabilities or generated text). The first processing results are stored together with the first data to be processed in the XDP accelerator card's cache, forming a historical benchmark dataset.
[0036] Step 103, through the loaded target model, data processing is performed on the target data to be processed based on multiple historical data to be processed in the cache of the preset accelerator and multiple historical processing results corresponding to the multiple historical data to be processed, to obtain a target processing result, wherein the target data to be processed is the data input into the loaded target model after the multiple historical data to be processed.
[0037] In an embodiment of the present application, optimized processing of incremental data is implemented. The target data to be processed refers to subsequent batches of data input after the historical data to be processed. By analyzing multiple historical data to be processed (i.e., past inputs) stored in the cache of a preset accelerator and their corresponding multiple historical processing results (i.e., historical outputs), reusable intermediate features (such as the key-value pair vector of the attention mechanism) are identified.
[0038] Cache-based data processing can be performed in the following ways, but is not limited to: When new data is input, the XDP accelerator card prioritizes matching similar historical data fragments in the cache, directly invoking the associated historical processing results for the current calculation (for example, skipping duplicate calculations through key-value caching), and only performing full calculations on the differences. The XDP accelerator card's near-storage computing architecture, which physically tightly couples the computing unit with the cache, ensures nanosecond response times for feature retrieval and calculations.
[0039] Step 104: Store the target processing result and the target data to be processed in a cache of a preset accelerator.
[0040] In an embodiment of the present application, the update and maintenance of the dynamic cache is completed. When the target processing results and the target data to be processed are written into the cache of the preset accelerator, an intelligent cache replacement strategy can be executed, that is, based on indicators such as data access frequency and timing correlation, low-frequency historical data (such as the Least Recently Used-K (LRU-K) algorithm) is automatically eliminated while retaining high-value intermediate features.
[0041] Cache updates can utilize write-merging technology to aggregate scattered writes into bursts, fully utilizing the XDP accelerator's highly concurrent input / output (I / O) channels (e.g., 1024 parallel queues). During storage, metadata tags (e.g., data hash values and timestamps) can be attached to target processing results to establish an index mapping with the target data to be processed, forming a continuously optimized, self-learning cache structure.
[0042] By leveraging the low-latency properties and cache reuse mechanism of the XDP accelerator card, the host memory interaction overhead is avoided, significantly shortening the end-to-end inference response time; by reducing redundant data transmission and calculations, PCIe bus activity and computing unit wake-up frequency are reduced, effectively saving system-level energy consumption; the synergistic effect of high-performance data processing properties and cache hit rate enables the computing unit to continuously work at saturation, significantly increasing the amount of data processed per unit time.
[0043] The performance improvement method of the present application, by loading the model into a preset accelerator, can apply the high-performance data processing capabilities and low-latency data transmission capabilities of the preset accelerator to the improvement of the model's computing capabilities. By storing historical data and processing results in the preset accelerator, when new data is processed through the model, data processing can be performed based on the historical data and processing results as a reference. If the data is repeated or similar, there is no need to reprocess the data, which can improve data processing efficiency. The cache technology is applied to the preset accelerator and model. While the computing power is improved, the delay and energy consumption of data access are reduced. Therefore, it can solve the technical problem of how to improve the computing power of large AI models and reduce energy consumption and delay in the data processing process, thereby achieving the technical effect of improving the computing power of large AI models and reducing energy consumption and delay in the data processing process.
[0044] In one implementable manner of an embodiment of the present application, after the first data to be processed (i.e., the starting data) is processed on the loaded target model, the following method may also be used but is not limited to: performing data processing on the second data to be processed based on the first processing result and the first data to be processed in the loaded target model to obtain a second processing result, and storing the second processing result and the second data to be processed in a cache of a preset accelerator, wherein the second data to be processed is data input into the loaded target model after the first data to be processed.
[0045] In an embodiment of the present application, it is necessary to further expand the data processing flow and introduce a second batch of data to be processed, that is, a subsequent input batch that follows the first batch of data to be processed in terms of timing logic (for example, the second user query of the same session in large model inference or the next key frame group of the video stream), which can have semantic continuity or computational dependency with the first batch of data. The second data to be processed is input into the loaded target model through the low-latency data injection channel of the preset accelerator. The low-latency data injection channel uses the user-mode ring buffer of the XDP accelerator card to directly receive data packets, avoiding the operating system kernel scheduling overhead.
[0046] When the second data to be processed is input into the loaded target model, the first processing result (including but not limited to intermediate states such as feature vectors and attention weights) and the first data to be processed (original input) stored in the cache of the preset accelerator are actively called to form a scalable computing context. Calling the data in the cache of the preset accelerator can rely on the state inheritance engine of the preset accelerator to automatically load the first data to be processed and the first processing result as initial conditions through hardware-level registers, so that the calculation process of the second data to be processed directly inherits the historical calculation state, eliminating repeated initialization overhead.
[0047] The preset accelerator uses a cache computing unit to fuse the features of the second data to be processed with the first processing result in real time (such as through gated weighting of the cross-attention mechanism), generating a second processing result that has both current input characteristics and historical semantic continuity, giving full play to the pipeline parallel capabilities of the preset accelerator, and making feature retrieval and model reasoning seamlessly connected at the hardware level.
[0048] After data processing is completed, the second processing result and the second data to be processed can be written into the cache of the preset accelerator through the cache block chain storage mechanism. The second processing result and the second data to be processed can be associated with the first data to be processed and the first processing result by establishing a physical address linked list to ensure the physical proximity of the logically associated data on the storage medium. The establishment of the physical address linked list association can utilize the atomic write merging unit of the preset accelerator to aggregate scattered write operations into burst transmissions, significantly improving storage efficiency. Ultimately, a progressive cache knowledge base is formed, which dynamically maintains the temporal dependencies between multiple batches of data through the metadata index table of the preset accelerator, providing topological support for the context calculation of any subsequent batch of data.
[0049] By inheriting the first batch of processing results, the initialization delay of inter-batch calculations is eliminated, significantly improving the processing efficiency of continuous data streams. Through the context binding mechanism, historical intermediate states are accurately reused, greatly reducing redundant calculations caused by repeated feature extraction. Through the chain storage structure, low-overhead expansion of unlimited batches of data is supported, providing hardware-level support for ultra-long sequence reasoning.
[0050] In one possible implementation of an embodiment of the present application, when storing the first processing result and the first data to be processed, it can also be implemented in, but not limited to, the following manner: extracting the first result element corresponding to the first processing result and the first data element corresponding to the first data to be processed; storing the first result element and the first data element in the cache of a preset accelerator.
[0051] In an embodiment of the present application, a refined storage strategy can be adopted to optimize cache utilization, specifically, including but not limited to the following methods: first, element-level feature extraction is performed, and the first result element is separated from the first processing result through the hardware feature filter built into the preset accelerator (the preprocessing unit integrated in the XDP accelerator card), that is, the key intermediate state with reuse value for subsequent calculations (for example: the attention weight matrix in the Transformer architecture, the hidden state vector of the Long Short-Term Memory (LSTM) unit, or the activation hot zone of the convolutional feature map). Simultaneously, the first data element is extracted from the first data to be processed, that is, the structured fragment with high reuse potential in the first data to be processed (such as: the semantic core word embedding vector of the text data, the local feature descriptor of the image data).
[0052] After extraction is complete, the first data element and the first result element can be optimized using a block compression engine. The first result element can be compressed using sparse coding (preserving non-zero values and their coordinates), and the first data element can be compressed using dictionary coding (constructing a high-frequency feature codebook). The compressed elements can be written to the cache via the cache controller of the preset accelerator. The cache controller uses a non-contiguous address mapping mechanism to disperse and store logically related elements in low-latency areas of the cache physical medium, while recording the topological relationship between elements through a metadata pointer table.
[0053] By storing only high-value elements and performing compression processing, cache capacity is effectively released, significantly improving the state retention capability of long-cycle calculations; hardware-level feature screening ensures that subsequent calculations can accurately call key intermediate states, avoiding redundant data contamination of the computing context; the collaborative design of block compression and non-contiguous storage enables storage operations and computing pipelines to be executed in parallel, significantly reducing the impact of data persistence on end-to-end latency.
[0054] In one implementable manner of an embodiment of the present application, the process of processing the second data to be processed to obtain a second processing result can be implemented by but not limited to the following manner: extracting the second data element of the second data to be processed, and performing element matching processing with the first data element in the cache of a preset accelerator based on the second data element to obtain the same data element and the matching hit rate, wherein the same data element is the element in the first data element that is the same as the second data element, and the matching hit rate is the ratio of the same data element in the second data element; performing element matching processing based on the same data element and the first result element to obtain a matching result element; when the matching hit rate is less than a preset hit rate threshold, data processing is performed on the second data to be processed through the loaded target model based on the matching result element to obtain a difference result element, wherein the difference result element is the difference element between the matching result element and the second processing result; performing data generation processing based on the difference result element and the matching result element to obtain the second processing result; when the matching hit rate is greater than or equal to the preset hit rate threshold, generating the second processing result based on the matching result element.
[0055] In an embodiment of the present application, an intelligent element reuse mechanism can be used to achieve a leap in computing efficiency. Through the hardware feature parser integrated in the preset accelerator, structured feature fragments (such as word embedding vectors of text or local descriptors of images) are separated from the second data to be processed to form a second data element. Subsequently, cross-batch element matching is performed, and the near cache comparison unit of the preset accelerator is used to perform a similarity comparison between the second data element and the first data element stored in the cache (such as feature matching based on Hamming distance). The same data elements (i.e., completely consistent or semantically equivalent feature fragments in the two batches of inputs) are output and the matching hit rate is calculated. The matching hit rate represents the proportion of the same data elements to the total amount of the second data elements, reflecting the degree of repetitiveness between the second data to be processed and the data stored in the cache of the preset accelerator.
[0056] During the element matching process, a result element association search is synchronously triggered. Based on the index identifier of the same data element, the corresponding first result element is accurately located from the cache of the preset accelerator. After topological mapping, the two are formed into a matching result element (the matching result element represents a historical calculation result that can be directly reused). Then, a dynamic calculation path decision is made based on the matching hit rate. When the matching hit rate is lower than the preset hit rate threshold (the preset hit rate threshold can be a custom threshold or dynamically set by the preset accelerator based on the model type and data characteristics, such as 99% or 100%), it indicates that the second to-be-processed data is significantly different from the first to-be-processed data. At this time, the matching result element is input as the basic condition into the loaded target model, and incremental calculation is performed on the second to-be-processed data to generate a difference result element representing the new feature. Finally, the difference result element can be combined with the matching result element through the result fusion to generate a complete second processing result. When the matching hit rate is equal to or higher than the threshold, the second processing result is directly generated by the matching result element through result reconstruction, completely skipping the model calculation step.
[0057] By dynamically switching the calculation path based on the hit rate threshold, the invalid calculation overhead in low-duplicate data scenarios can be significantly reduced; element-level matching ensures that historical results are accurately called, avoiding the risk of semantic distortion caused by overall data reuse; the difference result element generation mechanism minimizes the activation range of the computing unit and greatly improves the energy efficiency ratio.
[0058] In one possible implementation of an embodiment of the present application, before loading the target model to be optimized into the cache of a preset accelerator, the following method may also be used but is not limited to: initializing the cache of the preset accelerator, wherein the initialization process includes setting the cache of the preset accelerator to a preset size and setting the organization mode of the cache of the preset accelerator to a preset organization mode.
[0059] In an embodiment of the present application, during initialization processing, the cache of the preset accelerator is set to a preset size, which means that the physical capacity ratio of the high-speed storage area is dynamically allocated according to the parameter magnitude and computational complexity of the target model. For example, the low-latency SRAM block in the hybrid storage architecture of the XDP accelerator card is expanded to a preset proportion of the full cache (such as 70%) to ensure that high-frequency access data resides preferentially in the near computing unit, wherein the preset size is a size that is custom-set according to the parameter magnitude and computational complexity of the target model to be optimized.
[0060] Configuring the organization mode of the cache of the preset accelerator means defining the physical arrangement logic of the data in the storage medium, including but not limited to: adopting an interleaved memory body mapping strategy to improve the access bandwidth, while enabling a non-uniform storage access architecture to allow computationally intensive units to exclusively occupy exclusive cache blocks, and setting a cache line replacement strategy to optimize space utilization, wherein the preset organization mode is a customized organization mode. Specifically, this application does not impose any restrictions on the organization mode.
[0061] During initialization, predictive data access pattern templates (such as stride prefetching or adjacent block prefetching) can be preloaded based on model structural characteristics (e.g., the temporal dependencies of recurrent neural networks or the spatial locality of attention mechanisms), precisely aligning cache behavior with subsequent computational needs. The resulting hardware-ready cache environment features adaptive data layout capabilities, laying the foundation for efficiently hosting the target model to be optimized.
[0062] By jointly configuring preset sizes and organizational methods, the matching degree between cache space and model requirements is significantly improved, avoiding capacity waste or bottlenecks; optimized storage body mapping and non-uniform architecture design significantly reduce data retrieval latency, providing hardware support for high-throughput computing; early configuration of prefetch strategies enables cache behavior and model calculation mode to form an inherent synergy, effectively eliminating the overhead of runtime adaptive tuning.
[0063] In one possible implementation of the embodiment of the present application, before processing the first data to be processed, the following method may be used but is not limited to: performing data compression processing on the initial data to be processed to obtain compressed data to be processed; performing data encryption processing on the compressed data to be processed to obtain the first data to be processed.
[0064] In an embodiment of the present application, end-to-end data optimization preprocessing can be performed to adapt to the hardware characteristics of a preset accelerator. The initial data to be processed represents a data source (such as a distributed storage system or an edge device). Data compression processing is performed through a hardware compression pipeline integrated in the preset accelerator. Redundant information that does not contribute to the reasoning of the loaded target model (such as high-frequency image noise or text stop words) is dynamically identified and eliminated to generate high-fidelity compressed data to be processed. The data structure of the compressed data to be processed is aligned with the cache line granularity of the preset accelerator.
[0065] The preset accelerator can perform data encryption processing on the compressed data to be processed through the cryptographic acceleration engine, using a dynamic session key bound to the loaded target model (dynamically generated by the secure enclave of the preset accelerator) for encryption. The encryption process can simultaneously inject tamper-proof watermarks (such as blockchain-style hash chains) to ensure the integrity of the data during transmission and caching.
[0066] Compression processing significantly reduces the amount of data moved, fully leveraging the low-latency properties of the preset accelerator to accelerate data injection; hardware-level encryption establishes a chain of trust at the source of data input, effectively defending against side-channel attacks and man-in-the-middle hijacking; hardware collaborative processing of compression and encryption eliminates the additional computing load caused by traditional software preprocessing, allowing system resources to focus on core inference tasks.
[0067] In one implementable manner of an embodiment of the present application, the following method may also be used but is not limited to: calculating the operation time and energy consumption data corresponding to each of multiple historical data to be processed; sorting and processing the multiple historical data to be processed in the order of input into the loaded target model to obtain sorted data to be processed, and determining the sorted operation time and sorted energy consumption data corresponding to the sorted data to be processed based on the operation time and energy consumption data; when it is determined that the sorted operation time and the sorted energy consumption data do not meet the decreasing rule, determining that a fault has occurred in the preset accelerator, and generating a fault prompt message.
[0068] In an embodiment of the present application, it is necessary to perform hardware health status diagnosis to ensure computing reliability. Through a preset accelerator, the computing time and energy consumption data corresponding to multiple historical data to be processed are accurately calculated. Then, the historical data to be processed are strictly sorted and processed in the order of input into the loaded target model to form a sorted sequence of data to be processed. The corresponding performance indicators are synchronously reorganized into a sorted computing time and a sorted energy consumption data sequence, forming a performance change curve in the time dimension.
[0069] As cache history accumulates and hardware temperature rise stabilizes, computing time and energy consumption data should show a gradual downward trend (due to improved cache hit rates and optimized hardware computing pipelines). Therefore, during verification, it is necessary to check whether the sorted sequence satisfies the joint decreasing nature of computing time and energy consumption data. When a non-decreasing anomaly is detected (such as a sudden increase in computing time or a fluctuating increase in energy consumption), it is determined that the preset accelerator has a fault (such as a cache bit flip or computing unit aging), and the generation of a fault prompt message is immediately triggered. The fault prompt message contains at least a fault type code (such as a timing anomaly pattern fingerprint) and an impact range assessment, which is directly transmitted to the host alarm system through the preset accelerator's out-of-band management channel.
[0070] Through performance trend violation detection, signs of hardware degradation are discovered significantly earlier than traditional threshold alarm mechanisms; the execution of low-confidence calculations in abnormal hardware environments is blocked to prevent erroneous results from contaminating business flows; fault determination instantly freezes the computing pipeline, effectively preventing the irreversible spread of hardware damage.
[0071] In one implementable method of an embodiment of the present application, when calculating the operation time and energy consumption data, the following methods may also be used but are not limited to: obtaining the data size, number of operations, and cache bandwidth corresponding to each of multiple historical data to be processed, wherein the number of operations is the number of operations of the loaded target model to process the multiple historical data to be processed respectively, and the cache bandwidth is the bandwidth of the cache of the preset accelerator when storing the multiple historical data to be processed respectively; obtaining the processor performance corresponding to the processor where the loaded target model is located, and calculating the operation time corresponding to each of the multiple historical data to be processed respectively based on the data size, number of operations, cache bandwidth, and processor performance; obtaining the first energy consumption corresponding to the processor where the loaded target model is located, and calculating the energy consumption data corresponding to each of the multiple historical data to be processed respectively based on the data size, number of operations, cache bandwidth, processor performance, and the first energy consumption.
[0072] In an embodiment of the present application, a multi-dimensional performance modeling mechanism can be used to accurately quantify hardware resource consumption. The data size corresponding to each of multiple historical data to be processed can be obtained through the data feature probe of the preset accelerator. The data size is the number of bytes physically occupied by a single batch of inputs in memory (such as the product of the tensor dimension and the precision bit width). The number of operations is simultaneously extracted. The number of operations refers to the total number of floating-point instructions actually executed when the loaded target model processes a single batch of data. At the same time, the cache bandwidth is monitored, that is, the effective throughput of the cache subsystem of the preset accelerator when storing the corresponding historical data.
[0073] When calculating the operation time, you need to first obtain the processor performance. The processor performance can represent the real-time effective computing power of the preset accelerator computing core (such as the number of instructions per cycle (IPC) value affected by chip temperature, voltage and frequency regulation).
[0074] At this time, the calculation of the operation time can be performed by, but not limited to, formula (1):
[0075]
[0076] Among them, T is the operation time, D is the data size, N is the number of operations, B is the cache bandwidth, and P is the processor performance.
[0077] The computation time of each historical data is solved in parallel by the vector computing unit of the preset accelerator. Formula (1) essentially reflects that the computation delay is determined by the data movement efficiency (dominated by cache bandwidth) and the processing capability (dominated by processor performance).
[0078] When calculating energy consumption data, you also need to obtain the first energy consumption. The first energy consumption refers to the baseline energy consumption of the preset accelerator processing a single basic computing instruction, which is calculated by the board-level energy management unit based on the current voltage / current measured values.
[0079] At this time, the calculation of energy consumption data can be performed by, but not limited to, formula (2):
[0080]
[0081] Among them, E is energy consumption data, and V is the first energy consumption.
[0082] Formula (2) reveals that energy consumption is linearly proportional to computing time and is jointly controlled by cache bandwidth and processor performance.
[0083] Through direct hardware acquisition of underlying parameters, the physical reality and diagnostic credibility of performance indicators are significantly enhanced; the separate modeling of computing time and energy consumption enables independent analysis of the bottlenecks of the cache subsystem and computing unit, providing clear guidance for optimization; real-time acquisition of processor performance and primary energy consumption parameters ensures that the model continuously tracks changes in hardware operating status, eliminating the lag bias of static modeling.
[0084] In one possible implementation of the embodiment of the present application, the following method may also be used but is not limited to: obtaining the remaining storage space of the cache of the preset accelerator, and when the remaining storage space is less than or equal to the preset remaining space threshold, deleting the stored data to be processed and the processing results in the cache of the preset accelerator in order of storage time from longest to shortest.
[0085] In an embodiment of the present application, a dynamic cache space management mechanism is used to maintain the continuous computing performance of the preset accelerator, and the remaining storage space monitoring can be periodically triggered. The remaining storage space, that is, the total amount of unoccupied physical storage units in the cache medium, is obtained in real time through the status register of the preset accelerator cache controller. The remaining storage space is obtained by subtracting the actual occupied space of the currently stored data to be processed, processing results and metadata from the total capacity.
[0086] When it is monitored that the remaining storage space is less than or equal to the preset remaining space threshold (the preset remaining space threshold is a custom-set threshold, for example: the initial value is a preset proportion of the total cache capacity (such as 10%), which is automatically adjusted according to the model calculation characteristics in the later stage), the time-priority deletion strategy is immediately started. The time-priority deletion strategy includes but is not limited to: sorting the timeliness of all cached data to be processed and processing results through the storage timestamp index table (based on the absolute timestamp of the data written to the cache), strictly following the order from long to short storage time (that is, the earliest stored data is eliminated first), and then the physical block erase instruction can be used to directly release the storage block where the target data is located, and the cache metadata mapping table is synchronously updated to maintain the logical consistency of the storage space. High-value data identifiers (such as high-frequency access feature tags) are retained during the deletion process to ensure that the core computing state is not affected by accidental deletion.
[0087] Dynamic space reclamation significantly mitigates the risk of storage overflow, ensuring uninterrupted execution of long-cycle inference tasks; the time-series deletion strategy effectively suppresses the dilution of cache efficiency caused by obsolete data, and increases the retention probability of high-reuse value data; threshold adaptive adjustment and hardware-level execution of deletion operations completely avoid computing pipeline pauses caused by traditional software cleanup mechanisms.
[0088] Furthermore, in one possible implementation of the embodiment of the present application, the present application improves the computing power of large AI models and reduces energy consumption and latency by using the cache technology of the XDP accelerator card. Specifically, the following steps may be included:
[0089] (1) Initialize the cache of the XDP accelerator card: First, you need to initialize the cache of the XDP accelerator card, including setting the cache size, cache organization, etc.
[0090] (2) Loading the AI large model: Then, the AI large model needs to be loaded into the cache of the XDP accelerator card.
[0091] (3) Data preprocessing: Next, data preprocessing is required, including data compression and encryption.
[0092] (4) Computing large AI models: Then, it is necessary to compute large AI models, using the high-performance data processing capabilities and low-latency data transmission capabilities of the XDP accelerator card.
[0093] (5) Storing results: Finally, the calculation results need to be stored, including storing the results in a cache or transmitting them to other devices.
[0094] By applying the XDP accelerator card's cache technology to improve the computing power of large AI models, energy consumption and latency can be reduced. At the same time, the XDP accelerator card's high-performance data processing capabilities and low-latency data transmission capabilities can be applied to improve the computing power of large AI models.
[0095] Among them, when improving the computing power of large AI models, multiple algorithms can also be used, for example:
[0096] Algorithm 1: Cache algorithm for XDP accelerator card. Input: data, cache size. Output: cache hit ratio. It includes: initializing the cache, loading data into the cache, calculating the cache hit ratio, and returning the cache hit ratio.
[0097] Algorithm 2: AI large model calculation algorithm, input: data, AI large model, output: calculation result, including: loading the AI large model into the XDP accelerator card, preprocessing data, calculating the AI large model, storing the calculation result, and returning the calculation result.
[0098] In summary, this application can achieve the following technical effects:
[0099] 1. By loading the model into the preset accelerator, the high-performance data processing capability and low-latency data transmission capability of the preset accelerator can be used to improve the model's computing capability. By storing historical data and processing results in the preset accelerator, when new data is processed through the model, the historical data and processing results can be used as a reference for data processing. If the data is repeated or similar, there is no need to reprocess the data, which can improve data processing efficiency. By applying caching technology to the preset accelerator and model, while improving computing power, it reduces data access latency and energy consumption. Therefore, it can solve the technical problem of how to improve the computing power of large AI models and reduce energy consumption and latency in the data processing process, thereby achieving the technical effect of improving the computing power of large AI models and reducing energy consumption and latency in the data processing process.
[0100] 2. Methods for improving the computing power of large AI models using XDP accelerator cards, reducing energy consumption and latency. The high-performance data processing and low-latency data transmission capabilities of XDP accelerator cards can be applied to improve the computing power of large AI models. Furthermore, caching technology should be applied to XDP accelerator cards to enhance the computing power of large AI models, reducing data access latency and energy consumption.
[0101] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.
[0102] The embodiment of the present application also provides a device for improving performance. Figure 2 A schematic diagram of the structure of a performance-enhancing device provided in this application, such as Figure 2 Shown, including:
[0103] A loading unit 21 is configured to load the acquired target model to be optimized into a cache of a preset accelerator to obtain a loaded target model, wherein the preset accelerator includes at least high-performance data processing attributes and low-latency data transmission attributes;
[0104] The processing unit 22 is configured to input the acquired first data to be processed into the loaded target model for data processing to obtain a first processing result;
[0105] A storage unit 23 is configured to store the first processing result and first data to be processed in a cache of a preset accelerator, wherein the first data to be processed is the starting data input into the loaded target model;
[0106] The processing unit 22 is further configured to, through the loaded target model, perform data processing on the target data to be processed based on a plurality of historical data to be processed in a cache of a preset accelerator and a plurality of historical processing results corresponding to each of the plurality of historical data to be processed, to obtain a target processing result, wherein the target data to be processed is data input into the loaded target model after the plurality of historical data to be processed;
[0107] The storage unit 23 is further configured to store the target processing result and the target data to be processed in a cache of a preset accelerator.
[0108] In one embodiment of the present application, the processing unit 22 is also used to perform data processing on the second data to be processed based on the first processing result and the first data to be processed in the loaded target model to obtain a second processing result, and store the second processing result and the second data to be processed in the cache of a preset accelerator, wherein the second data to be processed is data input into the loaded target model after the first data to be processed.
[0109] In one embodiment of the present application, the storage unit 23 is further configured to:
[0110] Extracting a first result element corresponding to the first processing result and a first data element corresponding to the first data to be processed;
[0111] The first result element and the first data element are stored in a cache of a preset accelerator.
[0112] In one embodiment of the present application, the processing unit 22 is further configured to:
[0113] Extracting a second data element of the second data to be processed, and performing element matching processing with the first data element in a cache of a preset accelerator based on the second data element, to obtain identical data elements and a matching hit rate, wherein the identical data element is an element in the first data element that is identical to the second data element, and the matching hit rate is a ratio of the identical data element to the second data element;
[0114] Perform element matching processing based on the same data element and the first result element to obtain a matching result element;
[0115] When the matching hit rate is less than a preset hit rate threshold, data processing is performed on the second to-be-processed data using the loaded target model according to the matching result element to obtain a difference result element, wherein the difference result element is a difference element between the matching result element and the second processing result;
[0116] Performing data generation processing based on the difference result element and the matching result element to obtain a second processing result;
[0117] When the matching hit rate is greater than or equal to a preset hit rate threshold, a second processing result is generated according to the matching result element.
[0118] In one embodiment of the present application, Figure 3 As shown, the performance-enhancing device also includes:
[0119] The operating unit 24 is configured to perform initialization processing on the cache of the preset accelerator, wherein the initialization processing includes setting the cache of the preset accelerator to a preset size and setting the organization mode of the cache of the preset accelerator to a preset organization mode.
[0120] In one embodiment of the present application, Figure 3 As shown, the performance-enhancing device also includes:
[0121] The compression unit 25 is used to perform data compression processing on the initial data to be processed to obtain compressed data to be processed;
[0122] The encryption unit 26 is configured to perform data encryption processing on the compressed data to be processed to obtain first data to be processed.
[0123] In one embodiment of the present application, Figure 3 As shown, the performance-enhancing device also includes:
[0124] A calculation unit 27 is used to calculate the operation time and energy consumption data corresponding to each of the plurality of historical data to be processed;
[0125] a determination unit 28 for sorting the plurality of historical data to be processed in the order in which they were input into the loaded target model to obtain sorted data to be processed, and determining sorted computing time and sorted energy consumption data corresponding to the sorted data to be processed based on the computing time and energy consumption data;
[0126] The determining unit 28 is further configured to determine that a fault occurs in the preset accelerator and generate fault prompt information when it is determined that the sorted operation time and the sorted energy consumption data do not satisfy a decreasing rule.
[0127] In one embodiment of the present application, the calculation unit 27 is further configured to:
[0128] Obtain the data size, number of operations, and cache bandwidth corresponding to each of the multiple historical data to be processed, where the number of operations refers to the number of operations that the loaded target model performs on the multiple historical data to be processed, and the cache bandwidth refers to the bandwidth of the cache of the preset accelerator when storing the multiple historical data to be processed.
[0129] Obtain the processor performance corresponding to the processor where the loaded target model is located, and calculate the operation time corresponding to each of multiple historical data to be processed based on the data size, number of operations, cache bandwidth and processor performance;
[0130] The first energy consumption corresponding to the processor where the loaded target model is located is obtained, and energy consumption data corresponding to each of multiple historical data to be processed is calculated according to the data size, number of operations, cache bandwidth, processor performance and the first energy consumption.
[0131] In one embodiment of the present application, Figure 3 As shown, the performance-enhancing device also includes:
[0132] The deletion unit 29 is used to obtain the remaining storage space of the cache of the preset accelerator. When the remaining storage space is less than or equal to the preset remaining space threshold, the stored data to be processed and the processing results are deleted in the cache of the preset accelerator in the order of storage time from longest to shortest.
[0133] For the description of the features in the embodiment corresponding to the performance improvement device, please refer to the relevant description of the embodiment corresponding to the performance improvement method, and no further details will be given here.
[0134] An embodiment of the present application further provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above-mentioned performance improvement method embodiments.
[0135] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above-mentioned performance improvement method embodiments when running.
[0136] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.
[0137] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any of the above-mentioned performance improvement method embodiments are implemented.
[0138] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps in any of the above-mentioned performance improvement method embodiments.
[0139] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0140] The above is a detailed introduction to a performance improvement method and electronic device provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core ideas of the present application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.
Claims
1. A method for improving performance, characterized in that: include: Loading the obtained target model to be optimized into a cache of a preset accelerator to obtain a loaded target model, wherein the preset accelerator includes at least high-performance data processing attributes and low-latency data transmission attributes; Inputting the acquired first data to be processed into the loaded target model for data processing to obtain a first processing result, and storing the first processing result and the first data to be processed in a cache of the preset accelerator, wherein the first data to be processed is the starting data input into the loaded target model; By means of the loaded target model, data processing is performed on the target data to be processed based on a plurality of historical data to be processed in the cache of the preset accelerator and a plurality of historical processing results corresponding to each of the plurality of historical data to be processed, to obtain a target processing result, wherein the target data to be processed is data input into the loaded target model after the plurality of historical data to be processed; The target processing result and the target data to be processed are stored in a cache of the preset accelerator.
2. The performance improvement method according to claim 1, characterized in that: After storing the first processing result and the first to-be-processed data in the cache of the preset accelerator, the method further includes: In the loaded target model, data processing is performed on second data to be processed based on the first processing result and the first data to be processed to obtain a second processing result, and the second processing result and the second data to be processed are stored in the cache of the preset accelerator, wherein the second data to be processed is data input into the loaded target model after the first data to be processed.
3. The performance improvement method according to claim 2, characterized in that: The storing the first processing result and the first to-be-processed data in the cache of the preset accelerator includes: Extracting a first result element corresponding to the first processing result and a first data element corresponding to the first data to be processed; The first result element and the first data element are stored in a cache of the preset accelerator.
4. The performance improvement method according to claim 3, characterized in that: The performing data processing on the second data to be processed based on the first processing result and the first data to be processed in the loaded target model to obtain the second processing result includes: Extracting a second data element of the second data to be processed, and performing element matching processing with the first data element in the cache of the preset accelerator based on the second data element to obtain identical data elements and a matching hit rate, wherein the identical data element is an element in the first data element that is identical to the second data element, and the matching hit rate is a ratio of the identical data element in the second data element; Perform element matching processing on the identical data element and the first result element to obtain a matching result element; When the matching hit rate is less than a preset hit rate threshold, data processing is performed on the second to-be-processed data using the loaded target model according to the matching result element to obtain a difference result element, wherein the difference result element is a difference element between the matching result element and the second processing result; Performing data generation processing according to the difference result element and the matching result element to obtain the second processing result; In a case where the matching hit rate is greater than or equal to the preset hit rate threshold, the second processing result is generated according to the matching result element.
5. The performance improvement method according to claim 1, characterized in that: Before loading the acquired target model to be optimized into the cache of the preset accelerator to obtain the loaded target model, the method further includes: Initialization processing is performed on the cache of the preset accelerator, wherein the initialization processing includes setting the cache of the preset accelerator to a preset size and setting the organization mode of the cache of the preset accelerator to a preset organization mode.
6. The performance improvement method according to claim 1, characterized in that: Before inputting the acquired first data to be processed into the loaded target model for data processing to obtain a first processing result, the method further includes: Performing data compression processing on the initial data to be processed to obtain compressed data to be processed; The compressed data to be processed is encrypted to obtain the first data to be processed.
7. The performance improvement method according to claim 1, characterized in that: Before obtaining the target processing result by performing data processing on the target data to be processed based on the multiple historical data to be processed in the cache of the preset accelerator and the multiple historical processing results corresponding to the multiple historical data to be processed using the loaded target model, the method further includes: Calculating the operation time and energy consumption data corresponding to each of the plurality of historical data to be processed; Sorting the plurality of historical data to be processed in the order of inputting them into the loaded target model to obtain sorted data to be processed, and determining sorted computing time and sorted energy consumption data corresponding to the sorted data to be processed based on the computing time and the energy consumption data; When it is determined that the sorted operation time and the sorted energy consumption data do not satisfy a decreasing rule, it is determined that a fault occurs in the preset accelerator, and fault prompt information is generated.
8. The performance improvement method according to claim 7, characterized in that: The calculating of the operation time and energy consumption data corresponding to each of the plurality of historical data to be processed includes: Obtaining data size, number of operations, and cache bandwidth corresponding to each of the plurality of historical data to be processed, wherein the number of operations is the number of operations performed by the loaded target model to process the plurality of historical data to be processed, and the cache bandwidth is the bandwidth of the cache of the preset accelerator when storing the plurality of historical data to be processed; Obtaining the processor performance corresponding to the processor where the loaded target model is located, and respectively calculating the operation time corresponding to each of the plurality of historical data to be processed according to the data size, the number of operations, the cache bandwidth, and the processor performance; Obtain a first energy consumption corresponding to the processor where the loaded target model is located, and respectively calculate the energy consumption data corresponding to each of the multiple historical data to be processed according to the data size, the number of operations, the cache bandwidth, the processor performance and the first energy consumption.
9. The performance improvement method according to claim 1, characterized in that: The method further comprises: Obtain the remaining storage space of the cache of the preset accelerator. When the remaining storage space is less than or equal to a preset remaining space threshold, delete the stored data to be processed and the processing results in the cache of the preset accelerator in descending order of storage time.
10. An electronic device, characterized in that: include: memory for storing computer programs; A processor, configured to implement the steps of the performance improvement method according to any one of claims 1 to 9 when executing the computer program.