An ai accelerator card inference computing power benchmark test method and device

By using business scenario feature annotation and cross-scenario load transfer, the accumulated accuracy deviation is decomposed, the computing power sensitive component and bandwidth sensitive component are identified, the false verification effect is eliminated, and the benchmark test of AI accelerator cards is optimized. This solves the problems of underestimation and false reporting of computing power loss in existing testing methods and achieves accurate computing power comparison.

CN122470484BActive Publication Date: 2026-08-25JIANGSU HAINA ELECTRONICS TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610948197.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-29
Publication Date
2026-08-25
Estimated Expiration
2046-06-29

AI Technical Summary

Technical Problem

Existing benchmark testing methods for AI accelerator cards have large differences in load characteristics under different business scenarios, which leads to a decrease in the ability to distinguish the test results from the actual performance differences. In mixed precision inference scenarios, the computing power consumption is underestimated, and existing methods lack identification and elimination mechanisms.

Method used

By jointly annotating business scenario features, load anchor sequence is generated, and a computing power benchmark map is constructed through cross-scenario load transfer. Accumulated accuracy deviation is decomposed to identify computing power-sensitive and bandwidth-sensitive components, memory access conflicts are removed, accuracy recalibration and bandwidth bottleneck location are performed, false verification effects are eliminated, and benchmark test results are optimized.

Benefits of technology

A benchmark map is constructed that can truly reflect the effective computing power boundary of accelerator cards in various business scenarios, eliminating the influence of false reporting, solving the problem of underestimated computing power loss, and providing a reliable basis for computing power comparison.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122470484B_ABST
    Figure CN122470484B_ABST
Patent Text Reader

Abstract

The application discloses an AI acceleration card reasoning computing power benchmark test method and device, obtains reasoning load data and hardware configuration parameters, generates a load anchor point sequence through business scene feature joint labeling, constructs an initial load test set accordingly, quantifies concurrent throughput backoff features to generate a computing power benchmark atlas, directionally decomposes cumulative accuracy deviation in the test set, identifies two components of computing power sensitivity and bandwidth sensitivity, respectively obtains computing power correction and bandwidth bottleneck anchor points through utilization rate false reporting traceability and memory conflict frequency domain association, generates accuracy correction instructions and implements accuracy re-correction on the initial test set, performs cold start path loss detection and accuracy gradient monotonicity failure identification based on the corrected test set, locates false reporting consistency blind area anchor points, executes global optimization on the computing power benchmark atlas, and finally outputs benchmark test results, thereby improving the accuracy and business adaptability of AI acceleration card reasoning computing power benchmark test results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of integrated circuit testing technology, and in particular to a method and apparatus for benchmarking the inference computing power of an AI accelerator card. Background Technology

[0002] AI accelerator cards support various business scenarios such as search recommendation and dialogue generation in large model inference services. The load characteristics of different scenarios vary significantly. Existing benchmark testing methods adopt static testing strategies with fixed models and fixed concurrency scales. The test data deviates significantly from the actual business load distribution, and the test conclusions are difficult to reflect the computing power level of the accelerator cards in the actual deployment environment.

[0003] As concurrency increases, accelerator card throughput exhibits a rollback phenomenon. Test data near the rollback boundary suffers from spurious verification due to overfitting, reducing the ability of computational power assessments to distinguish true performance differences. In mixed-precision inference scenarios, pipeline pauses during precision level switching lead to false positives in driver layer utilization reports. Existing methods lack corresponding identification and removal mechanisms, resulting in an underestimation of precision-related computational power losses. These issues collectively constrain the accuracy of benchmark results, necessitating a testing method that eliminates rollback interference and false positives. Summary of the Invention

[0004] This invention discloses a benchmark testing method and apparatus for inference computing power of AI accelerator cards. It aims to construct a computing power benchmark map that covers the real business boundary and eliminates the false verification effect by jointly annotating business scenario features, cross-scenario load transfer and concurrent throughput backoff quantization. It eliminates the underestimation of computing power loss in mixed precision scenarios by decomposing the directionality of precision deviation and tracing the source of false utilization. Finally, it drives the global optimization of the benchmark map by jointly identifying the cold start path missing and precision gradient failure, providing a reliable basis for horizontal computing power comparison of different models of AI accelerator cards.

[0005] The first aspect of this invention proposes a benchmark testing method for the inference computing power of an AI accelerator card, comprising the following steps: Obtain the inference load data and hardware configuration parameters of the AI ​​acceleration card, and perform joint annotation of the inference load data and the hardware configuration parameters with business scenario features to generate a load anchor sequence; Based on the load anchor sequence, cross-scenario load transfer is performed to generate an initial load test set. Concurrent throughput backoff measurement is performed on the inference load data to generate a throughput deviation map. The initial load test set and the throughput deviation map are fused with backoff segment size constraints to construct a computing power benchmark map. The cumulative accuracy deviation in the initial load test set is decomposed directionally to identify computing power sensitive components and bandwidth sensitive components. The computing power sensitive components are subjected to accuracy level switching and utilization back-off identification to obtain computing power correction amount. Based on the bandwidth sensitive components and the memory access channel characteristics in the hardware configuration parameters, memory access conflict is correlated to obtain bandwidth bottleneck anchor points. Accuracy channel weight allocation is performed according to the computing power correction amount and the bandwidth bottleneck anchor points to generate accuracy correction instructions. According to the accuracy correction instruction, the initial load test set is recalibrated scene by scene to generate a calibrated load test set. Based on the calibrated load test set, cold start path missing detection and accuracy gradient monotonicity failure identification are performed to generate false alarm consistency blind zone anchor points. Based on the false alarm consistency blind zone anchor points, the computing power benchmark map is globally optimized and the benchmark test results are output.

[0006] A second aspect of this invention provides an AI accelerator card inference computing power benchmark testing device, comprising: Anchor point annotation unit is used to obtain inference load data and hardware configuration parameters of AI accelerator card, and perform joint annotation of business scenario features on the inference load data and hardware configuration parameters to generate load anchor point sequence; The graph construction unit is used to generate an initial load test set by performing cross-scenario load transfer based on the load anchor point sequence, perform concurrent throughput backoff measurement on the inference load data to generate a throughput deviation map, and construct a computing power benchmark graph by fusing the initial load test set and the throughput deviation map with backoff segment size constraints. The deviation correction unit is used to perform directional decomposition and identification of computing power sensitive components and bandwidth sensitive components on the cumulative accuracy deviation in the initial load test set, perform accuracy level switching and utilization back-off identification on the computing power sensitive components to obtain computing power correction amount, perform memory access conflict association based on the bandwidth sensitive components and the memory access channel characteristics in the hardware configuration parameters to obtain bandwidth bottleneck anchor points, and perform accuracy channel weight allocation based on the computing power correction amount and the bandwidth bottleneck anchor points to generate accuracy correction instructions. The benchmark optimization unit is used to perform scene-by-scene precision recalibration on the initial load test set according to the precision correction instruction to generate a calibrated load test set, perform cold start path missing detection and precision gradient monotonicity failure identification on the calibrated load test set to generate false alarm consistency blind zone anchor points, and perform global optimization on the computing power benchmark map based on the false alarm consistency blind zone anchor points to output benchmark test results.

[0007] The beneficial effects of this invention are reflected in the following aspects: First, by jointly annotating inference load data and hardware configuration parameters with business scenario characteristics, load records of different concurrency scales and accuracy types are anchored to the corresponding scenario categories according to resource constraint boundaries. Furthermore, cross-scenario load transfer fills the load combination gap existing in single-scenario collection, effectively solving the problem of large deviations between existing static testing strategies and real business load distribution. Based on this, the concurrency throughput backoff characteristics are directionally quantified, identifying false verification distribution areas and correcting the deviation direction. This allows test data near the backoff boundary to escape the interference of excessive linear fitting, constructing a benchmark map that truly reflects the effective computing power boundary of the accelerator card in various business scenarios. Second, this invention performs directional decomposition on the accumulated accuracy deviation of the test set, independently correcting the two sources: computing unit utilization loss and memory access bandwidth conflict. False jumps caused by pipeline pauses during precision level switching in the computing power-sensitive component are eliminated after scenario-by-scenario source tracing. For the bandwidth-sensitive component, the bottleneck is precisely located to a specific storage level by correlating the periodicity of memory access density with the frequency domain of bandwidth occupancy. The two correction results are combined to generate a precision correction instruction, driving the test set to complete precision recalibration, fundamentally solving the problem of systematically underestimating computing power loss in mixed-precision inference scenarios. Finally, addressing potential residual defects in the cold start path and precision gradient monotonicity of the test set after precision recalibration, this invention jointly encapsulates these two types of defects as false alarm consistency blind zone anchors. Using these anchors as indexes, the computing power benchmark map undergoes step-by-step global optimization in three stages: reprojection, false alignment region identification, and perturbation resampling. The bias effect of local extremum locking on the computing power benchmark value is eliminated layer by layer. The final benchmark test results can truly reflect the computing power difference between different accelerator card models in horizontal comparisons, providing a reliable basis for procurement decisions and deployment planning. Attached Figure Description

[0008] The accompanying drawings illustrate specific examples of the technical solutions described in this invention and, together with the detailed embodiments, form part of the specification, serving to explain the technical solutions, principles, and effects of this invention.

[0009] Figure 1 This is a flowchart illustrating a benchmark testing method for AI accelerator card inference computing power according to the present invention.

[0010] Figure 2 This is a structural block diagram of an AI accelerator card inference computing power benchmark testing device according to the present invention. Detailed Implementation

[0011] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, circuits, and methods are omitted so as not to obscure the description of this application with unnecessary detail.

[0012] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.

[0013] It should also be understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0014] The technical solutions of the embodiments of this application will be described below.

[0015] like Figure 1 As shown, this embodiment of the invention provides a benchmark testing method for AI accelerator card inference computing power, including the following steps S11-S14: Step S11: Obtain the inference load data and hardware configuration parameters of the AI ​​acceleration card, and perform joint annotation of business scenario features on the inference load data and hardware configuration parameters to generate a load anchor sequence.

[0016] Specifically, the inference load data and hardware configuration parameters of the AI ​​accelerator card are obtained. The actual request traffic carried by the AI ​​accelerator card during inference service operation is captured by a sampling window to form inference load data. The sampling window covers both peak and off-peak periods to ensure the inference load data reflects the complete load intensity distribution range. The inference load data uses each inference request as a record unit, covering four dimensions: input sequence length, batch size, precision type, and end-to-end latency. The batch size spans from a single request to 256 concurrent requests, fully covering the full intensity range under low-concurrency warm-up and high-concurrency saturation states. Hardware configuration parameters are obtained from the AI ​​accelerator card driver layer interface, covering three dimensions: peak computing power, memory capacity, and memory access bandwidth. The memory access bandwidth field is further subdivided into three levels: L1 cache bandwidth, L2 cache bandwidth, and global memory bandwidth. The bandwidth of these three levels, along with their corresponding capacity and scheduling cycle, constitute the memory access channel characteristics. The ratio of peak computing power to memory access bandwidth constitutes the card's compute-to-memory ratio. This ratio defines the bottleneck boundary between matrix multiplication-intensive and memory read-intensive operators. Different AI accelerator cards exhibit drastically different saturation characteristics under the same inference task due to variations in compute-to-memory ratio. The memory capacity determines the maximum sequence size that the KV-cache can support. The boundary values ​​of these three parameters together constitute the hardware constraint benchmark for scene category determination. Inference load data and hardware configuration parameters are linked through a unique device identifier field to ensure consistency between the two types of data during joint annotation.

[0017] The inference load data and hardware configuration parameters are jointly labeled with business scenario features to generate a load anchor sequence. Business scenario categories are divided into four types based on the resource constraint matching relationship between inference load data and hardware configuration parameters: computing power saturation scenario, bandwidth saturation scenario, memory-limited scenario, and balanced scenario. A computing power saturation scenario is defined as one where the concurrency scale is large, the precision is mainly INT8, and the throughput is close to the peak computing power conversion limit; a bandwidth saturation scenario is defined as one where the computation memory access intensity is lower than the hardware computation memory access ratio threshold and the measured memory access bandwidth utilization exceeds 80% of the hardware limit; if only the former condition is met, it is marked as a bandwidth-sensitive candidate scenario and does not trigger bandwidth saturation determination; a memory-limited scenario is defined as one where the average sequence length is high, causing KV-cache usage to approach the memory capacity limit and the P99 tail latency to suddenly increase after memory paging; and a balanced scenario is defined as one where the utilization rate of each resource does not exceed the corresponding threshold of 80% and the throughput maintains a near-linear growth with the concurrency scale. Anchor points are extracted for each scenario using uniform step sizes based on latency gradients. For scenarios with saturated computing power and bandwidth, where bottleneck boundaries are most sensitive to changes in concurrency scale, 20 anchor points are selected for each. For memory-constrained scenarios, 15 anchor points are selected, and for balanced scenarios, 10 anchor points are selected. Anchor point records are encapsulated with five fields: input sequence length, batch size, precision type, latency, and scenario category. These fields are then concatenated in scenario order to form a load anchor point sequence. Scenario boundary positions are recorded in an index table. Records with unclear resource constraint boundary matching relationships are marked as composite constraint states, excluding them from the source anchor point candidate set for cross-scenario load transfer, and are retained only as boundary reference entries in the load anchor point sequence.

[0018] Step S12: Based on the load anchor sequence, perform cross-scenario load transfer to generate an initial load test set, perform concurrent throughput backoff measurement on the inference load data to generate a throughput deviation map, and construct a computing power benchmark map by fusing the initial load test set and the throughput deviation map with backoff segment size constraints.

[0019] Specifically, an initial load test set is generated by performing cross-scenario load transfer based on the load anchor sequence. Each anchor in the load anchor sequence is divided into source scenario anchors and target scenario anchors according to scenario category labels. The input sequence length, batch size, and precision type fields of the source scenario anchors are mapped to the target scenario load parameters according to the hardware constraint boundaries of the target scenario. The mapping rules are determined based on the differences in resource constraints between the source and target scenarios: when an anchor from a computing power saturated scenario is transferred to a bandwidth saturated scenario, the batch size is compressed by the upper limit of memory access bandwidth; when transferred to a memory-limited scenario, the input sequence length is truncated according to the KV-cache threshold corresponding to the upper limit of memory capacity; when an anchor from a bandwidth saturated scenario is transferred to a computing power saturated scenario, the batch size is expanded by the upper limit of computing power peak utilization; in other directions, the batch size and input sequence length are compressed or truncated respectively based on the hardware constraint boundaries corresponding to the target scenario. After all load anchor sequence passes across scenarios, the passing results for each target scenario are deduplicated and merged with the original anchors based on two dimensions: input sequence length and batch size. If the two dimensions are completely identical, the passing result overwrites the original anchor (the passing result has been remapped according to the target scenario constraints, and the matching degree is higher than the original anchor collected in a single scenario). The merged result forms a two-dimensional load matrix. The initial load test set is composed of two-dimensional load matrices from each scenario. Each record contains four fields: input sequence length, batch size, precision type, and source scenario. The initial load test set covers load combinations of all scenario categories for the load anchor sequence. Extreme cross-scenario load combinations are supplemented to the initial load test set through a passing mechanism.

[0020] In some embodiments, the step of performing concurrent throughput backoff measurement on the inference load data to generate a throughput deviation map includes: performing scenario-by-scenario concurrent scale throughput detection on the inference load data to generate a scenario-by-scenario throughput set; performing linearity deviation evaluation calculation on the scenario-by-scenario throughput set to generate a scenario-by-scenario pseudo-linear parameter set; performing overfitting segment extraction and identification of pseudo-validation distribution areas on the scenario-by-scenario pseudo-linear parameter set; and performing throughput deviation direction calibration based on the pseudo-validation distribution areas and the scenario-by-scenario throughput set to generate a throughput deviation map.

[0021] A scenario-by-scenario concurrency scale throughput test is performed on the inference load data to generate a scenario-by-scenario throughput set. Records for each scenario in the inference load data are grouped by the batch size field. Within each group, the batch size values ​​are arranged in ascending order of concurrency levels, from smallest to largest, forming a concurrency scale gradient sequence. The intervals between these levels are kept uniform to ensure consistent distribution of sampling points for subsequent linear fitting. The measured throughput T_real (data / ms) at each concurrency level is obtained by dividing the batch size (data) by the end-to-end latency (ms), and paired with the concurrency scale gradient sequence to form scenario throughput pairs. The scenario throughput pairs for all scenarios in the inference load data are summarized to form the scenario-by-scenario throughput set. The scenario-by-scenario throughput set shows near-linear growth in the low-concurrency segment and a continuously decreasing growth rate in the high-concurrency segment. Bandwidth-sensitive scenarios, due to their memory-intensive characteristics, reach a non-linear transition at lower concurrency levels, while matrix multiplication-intensive scenarios only enter the non-linear growth rate decline phase at higher concurrency levels. The difference in the transition positions between these two types of scenarios serves as the basis for subsequent scenario-level pseudo-linear region labeling. The length of each sequence in the per-scenario throughput set is determined by the number of concurrency scale tiers covered by the inference load data. If the coverage of a tier is less than 30% of the total tiers, it is marked as low confidence. If the extrapolated interval exceeds 50% of the measured interval, the pseudo-linear region determination result is downgraded to a reference value. This downgraded result is simultaneously marked with a low confidence label during subsequent pseudo-validation distribution region construction and scale-constrained residual field generation. Each sequence in the per-scenario throughput set is stored as its original sampled value without smoothing to preserve the true concurrency-sensitive boundary location information.

[0022] A linearity deviation assessment is performed on the throughput set for each scenario to generate a pseudo-linear parameter set for each scenario. Due to differences in load characteristics, each scenario in the throughput set exhibits a distinctly different throughput growth slope. Least-squares linear fitting is performed on the throughput sequence of each scenario in the throughput set against the concurrency scale. The slope k (1 / ms) and intercept b (segments / ms) of the fitted line constitute the initial linear parameters, which describe the throughput growth rate and zero-concurrency intercept of that scenario under the ideal linear assumption. The difference between the measured throughput sequence and the fitted line at each concurrency level forms a residual sequence. The sign change pattern of the residual sequence reveals the direction and location of the throughput deviation from linearity. A positive residual value in the low-concurrency segment indicates that the measured value is higher than the linear extrapolation value, while a negative residual value in the high-concurrency segment indicates a fallback to below the linear extrapolation value. When the throughput sequence has a significant nonlinear segment, linear fitting leads to overfitting: the slope of the fitted line is determined by the entire data segment. In the low-concurrency linear growth segment, the slope is too low, resulting in positive residuals; in the high-concurrency regression segment, the slope is too high, resulting in negative residuals. The positive and negative residuals of the two segments cancel each other out, minimizing the absolute value of the residuals in the middle concurrency segment. This local minimum segment coincides with the local tangent of the nonlinear curve. When the absolute value of the residuals in the middle concurrency segment is less than 30% of the average absolute value of the residuals across the entire segment, it is identified as a pseudo-linear region. The scenario-specific pseudo-linear parameter set consists of the initial linear parameters, residual sequence, and the start and end concurrency scales of the pseudo-linear region for each scenario. The start and end concurrency scales of the pseudo-linear region calibrate the boundary of the concurrency range where the linear fitting model completely fails in that scenario. In the INT8 quantization scenario, due to the shift of the saturation point of the dedicated acceleration engine, the pseudo-linear region usually appears in the high-concurrency segment. The right boundary of its pseudo-linear region is significantly higher than that of the FP16 scenario. The difference between the two directly quantizes the effect of the quantization engine on the release of concurrent scheduling resources. This difference is recorded in the scenario-by-scenario pseudo-linear parameter group for subsequent comparison. The residual sequence of the scenario-by-scenario pseudo-linear parameter group retains the positive and negative signs for the calibration of the throughput deviation direction.

[0023] For each scenario-specific pseudo-linear parameter group, overfitting segment extraction and pseudo-verification distribution areas are identified. The throughput curve exhibits local flattening at the on-chip cache capacity boundary. This flattening naturally reduces the linear fitting error for this concurrency level segment, making it difficult to distinguish from the overfitting state in terms of residual shape. This leads to their overlap on the concurrency scale axis. The start and end concurrency levels of each scenario's pseudo-linear region within the scenario-specific pseudo-linear parameter group determine the candidate range for overfitting segments. If the mean absolute value of the residual sequence within the candidate range is less than 30% of the mean absolute value of the residuals across the entire segment, the segment is confirmed as an overfitting segment. If the threshold is not reached, the candidate range is narrowed down to the consecutive levels with the smallest residuals for re-evaluation. The linear fitting model prediction error is numerically smallest within the overfitting segment, but this low error stems from a local tangency between the fitted curve and the nonlinear throughput curve, rather than a true linear relationship. Throughput test data within the overfitting segment has extremely low performance differentiation capabilities. Two different AI accelerator cards with a 20% difference in computing power exhibit similar throughput values ​​within the overfitting segment, with the differences masked by the local tangent effect. The concentration of overfitting segments in different scenarios on the concurrency scale axis reflects the architectural characteristics of the AI ​​accelerator card. The pseudo-verification distribution area is composed of the union of all overfitting segments in all scenarios on the concurrency scale axis (taking the union rather than the intersection to ensure that all concurrency levels with overfitting risk in any scenario are marked, avoiding omissions that could lead to misjudgments). All concurrency levels within the pseudo-verification distribution area are marked as pseudo-verification. The start and end concurrency scale boundaries and the number of covered levels are recorded in the structure. When the number of covered levels exceeds 40% of the total number of levels (corresponding to more than half of the concurrency levels in the test interval being in pseudo-verification, and the peak throughput indicator can no longer distinguish the difference in computing power), a wide pseudo-verification blind zone alarm is triggered. The alarm indicates that this model of AI accelerator card has a systematic evaluation failure risk in the normal concurrency scale range, and relying on the peak throughput indicator for computing power comparison will result in serious misjudgments.

[0024] A throughput deviation map is generated by calibrating the direction of throughput deviation based on the pseudo-validation distribution region and the scenario-by-scenario throughput set. The backlash characteristics of the throughput curve in the pseudo-validation distribution region are systematically underestimated due to the overfitting effect of linear fitting. It is necessary to calibrate the deviation direction of this section separately to restore the true backlash magnitude. The original deviation sequence is formed by subtracting the measured throughput sequence of each scenario in the scenario-by-scenario throughput set from the linear fitting parameters of each scenario associated with the pseudo-validation distribution region at each concurrency scale level. The pseudo-validation distribution region is derived from the scenario-by-scenario overfitting section. The linear fitting slope k (1 / ms) and intercept b (strings / ms) of the corresponding scenario are recorded together with the pseudo-validation distribution region structure. The deviation value formula is D(c)=T_real(c)-(k×c+b), where c is the concurrency scale (strings), T_real(c) is the measured throughput at the concurrency scale c (strings / ms), and D(c) and T_real(c) have the same dimensions. Within the pseudo-validation distribution area, the original deviation value corresponding to the concurrency level is systematically underestimated due to overfitting. The deviation value in this segment is corrected towards a point away from zero. The correction amount is equal to 50% of the mean absolute value of the residuals in the pseudo-linear region of this scenario (this proportion is half of the mean residual value to partially restore the true deviation suppressed by overfitting, while avoiding overcorrection introduced by full correction). The correction direction has the same sign as the current deviation value, and the corrected deviation magnitude is restored to the true deviation level. The corrected deviation sequence is arranged by scenario category to form a deviation matrix. The rows of the deviation matrix correspond to the scenario category, and the columns correspond to the concurrency level. The throughput deviation graph is constructed using the deviation matrix as the data source. The deviation curve for each scenario marks the starting point of throughput rollback for that scenario from the position where it crosses the zero axis in the positive value region. The depth of the negative value region to the right of the rollback starting point reflects the degree of computing power loss. In long-sequence inference scenarios, the memory access occupancy of the attention layer increases with the square of the sequence length. After the memory access bandwidth hard cap is triggered, the throughput continues to decrease monotonically. The depth of the negative value region can be several times that of computing power saturation scenarios. Therefore, the throughput deviation graph has the ability to directly distinguish between the two types of bottleneck sources.

[0025] In some embodiments, the step of constructing a computing power benchmark map by fusing the initial load test set with the throughput deviation map to meet the backoff segment size constraints includes: extracting backoff segment size feature points from the throughput deviation map to generate a size anchor point set; constructing a size constraint residual field by constraining the size anchor point set with the initial load test set; identifying scenarios where the residual is continuously zero based on the size constraint residual field to generate test blind zone labels; and constructing a computing power benchmark map by fusing the backoff segment size constraints based on the test blind zone labels.

[0026] The scale feature points of the rollback segment are extracted from the throughput deviation map to generate a scale anchor point set. The deviation curves of each scenario in the throughput deviation map exhibit a three-segment change on the concurrency scale axis: a near-linear segment, a deep rollback segment, and a segment that stabilizes again under extremely high concurrency. These correspond to three types of feature positions: the rollback start point, the minimum deviation point, and the deviation recovery point, respectively. The rollback start point is the first position where the deviation value turns from positive to negative and remains continuously negative at consecutive levels. The minimum deviation point is the position where the global minimum value of the deviation curve for that scenario is located. The deviation recovery point is the position where the first difference of the deviation curve approaches zero after the minimum deviation point. The rollback start point marks the effective concurrency upper limit for the AI ​​accelerator card to maintain linear throughput growth. The minimum deviation point marks the concurrency scale level where the rollback is most severe. The deviation recovery point marks the position where the throughput stabilizes after the deep rollback. The deviation recovery point usually corresponds to the AI ​​accelerator card's internal scheduling mechanism entering a stable rate-limiting state under extremely high concurrency. The concentrated distribution of concurrent scale values ​​at the rollback starting point across different scenarios indicates that different business scenarios share the same concurrent scheduling saturation boundary. When multi-scenario large-model inference services are deployed in a mixed manner, this saturation boundary constitutes a unified upper limit constraint on the concurrent scale. The three types of scale feature points for each scenario, along with the scenario category label, are summarized to form a scale anchor point set. Each record contains three fields: scenario category, feature point type, and corresponding concurrent scale value. In some scenarios, the deviation curve shows no obvious trend of deviation recovery at the maximum concurrency level, resulting in missing deviation recovery points. This state typically occurs in scenarios with extremely limited memory access bandwidth, where throughput continues to decrease monotonically after reaching the bandwidth limit until the test boundary without a plateau. In the scale anchor point set, the deviation recovery point field for this scenario is marked with a null value.

[0027] A scale constraint residual field is generated by constructing constraint residuals from the scale anchor set and the initial load test set. The concurrent scale value of the fallback starting point for each scenario in the scale anchor set serves as the upper limit of the scale constraint for that scenario. The concurrent scale of the minimum deviation point in the scale anchor set is recorded synchronously for subsequent blind zone correction. The constraint residual value is formed by subtracting the upper limit of the scale constraint from the batch processing scale field of each record in the same scenario in the initial load test set. A positive residual value indicates that the batch processing scale is lower than the fallback starting point and belongs to the effective non-fallback interval. A negative residual value indicates that the batch processing scale has entered the fallback segment. The larger the absolute value of the residual value, the further it deviates from the fallback boundary. All scenario constraint residual values ​​are arranged according to two dimensions: scenario category and batch processing scale to form the scale constraint residual field. The positive value region corresponds to the record distribution that can stably reflect the linear computing power characteristics, and the negative value region corresponds to the oversaturated record distribution that has entered the fallback. The boundary between positive and negative values, i.e., the fallback boundary, is marked with a zero-value contour line in the scale constraint residual field. In the scale-constrained residual field, the width of the transition interval from positive to negative for constraint residual values ​​within the same scenario reflects the steepness of the throughput backoff for that scenario. Long-sequence inference scenarios dominated by attention computation typically have narrower transition intervals and clearer backoff boundaries, while short-sequence batch inference scenarios dominated by matrix multiplication have wider transition intervals and blurred backoff boundaries. When the difference in transition interval width exceeds 20% of the total concurrency level, the blurring of the backoff boundaries for the two types of scenarios is no longer on the same order of magnitude, and the scale constraint upper limit needs to be calibrated separately rather than using a uniform threshold. The number of rows in the scale-constrained residual field matrix equals the number of scenario categories in the initial load test set, and the number of columns equals the number of batch size levels. Each record in the scale-constrained residual field has a unique corresponding residual value cell.

[0028] Based on the scale-constrained residual field, scenarios where residuals are consistently zero are identified, generating test blind zone labels. In actual operation and maintenance, the concurrency scale of AI accelerator cards is often maintained around a fixed level due to load balancing strategies that evenly distribute requests across multiple cards. When this fixed level happens to fall at the fallback initiation point, the throughput test results remain in a critical neutral region for a long time. The interval in the scale-constrained residual field where the absolute value of the constraint residual is consistently lower than one level step size across consecutive concurrency levels corresponds to this state. Within this interval, the batch processing scale fluctuates slightly around the fallback initiation point without clearly falling into the fallback or non-fallback segment. The throughput recorded in this interval does not show linear growth characteristics nor obvious fallback characteristics. The scenario of continuously zero residual blocks combined with the concurrency scale corresponds to the AI ​​accelerator card being in a critical state where scheduling resources are just saturated. The actual concurrency scale in the inference service is often affected by the randomness of user request arrivals and fluctuates around this critical region. Therefore, the throughput test results show random jitter unrelated to computing power. The wider the coverage of persistently zero residual blocks in the scale-constrained residual field, the higher the probability that the total concurrency scale will fall within this range when multiple large model inference services are deployed in a mixed manner. The actual observed throughput fluctuation can be more than three times the actual computing power difference, masking the performance differentiation between different AI accelerator cards. The test blind zone label encapsulates the scenario attribution field and the concurrency scale boundary field of all persistently zero residual blocks. The scenario attribution field is used to locate the initial load test set record that needs to be corrected when constructing the subsequent computing power benchmark map. The concurrency scale boundary field is used to limit the correction range of the batch processing scale field when making subsequent corrections. The number of test blind zone label entries is equal to the total number of persistently zero residual blocks identified in the scale-constrained residual field. Residual blocks with a width value greater than 10 concurrency levels are marked as wide test blind zones in the test blind zone label. The computing power benchmark evaluation conclusion of the wide test blind zone scenario is accompanied by a credibility limitation label.

[0029] A benchmark computing power map is constructed by fusing backoff segment size constraints based on test blind zone labels. When cloud inference services elastically scale up or down, they often maintain the single-card concurrency level near the backoff boundary to maximize resource utilization. At this time, the throughput benchmark value is most severely affected by blind zone interference. The blind zone scenarios and concurrency ranges marked by test blind zone labels are located and recorded in the initial load test set. The test blind zone label structure synchronously records the minimum deviation point concurrency scale corresponding to each blind zone scenario. The batch processing scale field of the located record is corrected according to this minimum deviation point concurrency scale value. After correction, the batch processing scale deviates from the blind zone center interval, making the record clearly fall into the backoff segment or non-backoff segment, eliminating the interference of the critical state on the throughput benchmark value. After correction, the batch size value of each record is read according to the deviation curve of each scenario, and the deviation correction value at the corresponding concurrency scale is read. After deviation correction, the throughput is arranged according to the scenario category and concurrency scale level to form a scenario computing power benchmark sequence. The peak point of each scenario computing power benchmark sequence corresponds to the effective computing power upper limit of the AI ​​acceleration card in the scenario. The peak point calculation formula is B_peak=max{T_real(c)-D(c)|c∈[c_min,c_ret]}, where c_min is the lower limit of the effective concurrency scale of the current scenario (number of records), which is the minimum value of the batch size field of the scenario in the initial load test set; c_ret is the concurrency scale at the rollback starting point (number of records); T_real(c) is the measured throughput at the concurrency scale c (number of records / ms); D(c) is the throughput deviation value at the corresponding concurrency scale (number of records / ms). The rollback segment D(c) is negative, and subtracting D(c) is equivalent to the theoretical throughput after restoring the rollback loss. The computing power benchmark sequence for all scenarios is arranged in rows by scenario category and columns by concurrency level to form a computing power benchmark map. Each cell of the computing power benchmark map stores the throughput benchmark value after deviation correction.

[0030] Step S13: Directional decomposition is performed on the cumulative accuracy deviation in the initial load test set to identify the computing power sensitive component and the bandwidth sensitive component. Accuracy level switching and utilization backoff identification are performed on the computing power sensitive component to obtain the computing power correction amount. Based on the memory access channel characteristics in the bandwidth sensitive component and the hardware configuration parameters, memory access conflict is correlated to obtain the bandwidth bottleneck anchor point. Accuracy channel weight allocation is performed according to the computing power correction amount and the bandwidth bottleneck anchor point to generate accuracy correction instructions.

[0031] Specifically, the cumulative precision deviation in the initial load test set is directionally decomposed to identify computationally sensitive and bandwidth-sensitive components. In mixed-precision inference scenarios, the AI ​​accelerator card simultaneously carries INT8 quantization operators and FP16 activation operators. INT8 matrix multiplication primarily consumes on-chip computing power, while the FP16 attention layer primarily consumes memory access bandwidth. The end-to-end latency deviation recorded in each scenario of the initial load test set is formed by the cumulative sum of these two types of precision deviations with different directions. The cumulative precision deviation is defined as the difference between the measured latency of each scenario in the initial load test set and the baseline latency of single-precision full inference in the same scenario. The baseline latency of single-precision full inference is taken as the measured latency under the corresponding concurrency scale when only the FP32 precision level is enabled in the initial load test set. Directional decomposition is achieved by performing principal component decomposition on the cumulative accuracy deviation sequence of each scenario in the initial load test set. After decomposition, the cosine similarity between the direction of each principal component and the vectors of peak computing power utilization change and memory bandwidth utilization change is calculated. The principal component with the highest absolute value of cosine similarity with the vector of peak computing power utilization change is extracted as the computing power sensitive component, and the principal component with the highest absolute value of cosine similarity with the vector of memory bandwidth utilization change is extracted as the bandwidth sensitive component. Both types of components are considered valid when the absolute value of cosine similarity exceeds 0.7; otherwise, the corresponding directional component is set to zero. The two types of components are expanded into a component matrix according to the two-dimensional structure of scenario-concurrency scale in the initial load test set. The component value of each cell is filled with the projection value of the corresponding principal component at the concurrency scale of that scenario and carries a positive or negative sign. A positive value indicates that the accuracy loss under this condition is higher than the baseline level, and a negative value indicates that it is lower than the baseline level.

[0032] In some embodiments, the step of performing precision level switching utilization rollback identification on the computing power sensitive component to obtain computing power correction includes: extracting scene-by-scene precision level utilization from the computing power sensitive component to generate a utilization sequence; performing switching instantaneous jump detection on the utilization sequence to generate a utilization false alarm jump sequence; performing scene-by-scene false alarm tracing detection based on the utilization sequence and the utilization false alarm jump sequence to generate a valid scene marker; and anchoring the computing power sensitive component with computing power through the valid scene marker to obtain the computing power correction amount.

[0033] Utilization rates for each scene's precision level are extracted from the computing power sensitive component to generate a utilization rate sequence. The AI ​​accelerator card driver layer reports the utilization rate of computing power units at each precision level with a fixed sampling period of 100ms. After aligning the sampled data with the component matrix of the computing power sensitive component according to the concurrency level, the actual computing power consumption of the computing power unit at each concurrency level is converted into actual computing power (FLOPs / s) by multiplying the measured throughput (lines / s) by the single inference computation (FLOPs / line). The single inference computation is obtained by querying the operator computation statistics interface of the inference framework layer under the corresponding precision level. The ratio of actual computing power to the rated peak computing power (FLOPs / s) of that precision level is extracted as the precision level utilization rate. The ratio ranges from 0 to 1. A ratio close to 1 indicates that the computing power unit is close to full load, while a ratio significantly lower than 1 indicates that there is idle and wasted computing power unit. The precision level utilization rates are arranged in ascending order of concurrency level to form the utilization rate sequence for each scene. The number of elements in the utilization rate sequence is equal to the number of concurrency levels covered by that scene. Matrix multiplication-intensive scenarios maintain high utilization rates during high-concurrency periods, while memory-intensive scenarios exhibit significantly lower utilization rates due to memory access bandwidth limitations. The difference in utilization rate sequence patterns between these two scenarios directly reveals the source of computing power bottlenecks. The utilization rate sequence shows abrupt changes before and after the precision level switching position. The direction of these changes depends on the difference between the rated peak computing power of the precision level after the switch and the value before the switch. When switching from INT8 to FP16, the rated peak computing power decreases, leading to a smaller denominator. Even if the actual throughput decreases simultaneously, the utilization rate sequence value may temporarily appear artificially high. The length of this artificially high interval is positively correlated with the pipeline refill time. For AI accelerator cards with deep pipeline stages, the duration of the artificially high interval is positively correlated with the pipeline refill cycle, potentially lasting for several driver layer sampling cycles. All scenario utilization sequences are indexed and summarized by scenario category. Each utilization sequence includes a list of corresponding scenario precision level switching positions. The switching position list records the concurrency level number where the precision level switch occurred. An empty list indicates that the scenario maintains a single precision level throughout.

[0034] A false utilization jump sequence is generated by performing instantaneous jump detection on the utilization rate sequence during the switch. When the precision level switches, the hardware pipeline is refilled. During this refilling period, the computing units are in a paused waiting state. The driver layer still reports the cumulative average utilization rate before the switch, resulting in abrupt peaks in the utilization rate sequence near the switch position that do not match the actual computing power state. Instantaneous jump detection is achieved by performing a first-order difference on the utilization rate sequence and extracting positions where the absolute value of the difference exceeds three times the average of adjacent segments. These threshold positions correspond to false utilization jump points caused by precision level switching. In the first-order difference calculation, the average of adjacent segments is taken as the average of the absolute values ​​of the differences of several levels before and after the current level (the window width is rounded down to 10% of the total number of concurrent scale gradient levels to ensure sufficient local mean estimation coverage while avoiding contamination by the jump points themselves). The jump amplitude, jump direction, and corresponding concurrent scale level of each false jump point are encapsulated as jump event records. All jump event records within the same scenario are arranged in order of concurrent scale to form the false utilization jump sequence for that scenario. The sign of the jump amplitude in the utilization false alarm jump sequence reflects the switching direction: a negative jump occurs when switching from low precision to high precision, and a positive jump occurs when switching from high precision to low precision. The absolute value of the amplitude reflects the false alarm intensity caused by the switch. When precision level switching is frequently triggered near the concurrency scale critical zone, the number of elements in the utilization false alarm jump sequence is significantly higher. The utilization sequence exhibits continuous sawtooth-shaped fluctuations in this concurrency scale range, with the sawtooth amplitude increasing with the switching frequency. When the switching interval is shorter than two pipeline fill cycles, adjacent jump events overlap in the utilization false alarm jump sequence. The jump amplitude at the overlapping position is the sum of the absolute values ​​of the two jump amplitudes to retain the actual false alarm intensity after superposition. When the two jump directions are opposite, the difference is taken to reflect the net false alarm amplitude after mutual cancellation.

[0035] For example, the step of generating valid scene markers by performing scene-by-scene false alarm source tracing detection based on the utilization rate sequence and the utilization rate false alarm jump sequence includes: performing scene-by-scene synchronization alignment of the utilization rate sequence and the utilization rate false alarm jump sequence to generate a scene-level deviation sequence; extracting the zero-deviation persistence interval of the scene-level deviation sequence to generate a pseudo-steady-state deviation interval; performing constant value locking scanning to identify an abnormal deviation scene set on the scene-level deviation sequence and the pseudo-steady-state deviation interval; and performing invalid scene removal based on the abnormal deviation scene set to generate valid scene markers.

[0036] The utilization rate sequence and the utilization rate false jump sequence are synchronized and aligned on a scenario-by-scenario basis to generate a scenario-level deviation sequence. There is a phase deviation between the utilization sampling clock and the concurrent scale gradient scan stepping clock in the driving layer. Direct alignment by sequence number would cause a systematic offset in the jump deduction position. Synchronization and alignment are achieved by locating the clock deviation through cross-correlation calculation. During cross-correlation peak detection, positions with zero hysteresis are excluded to avoid interference from the global maximum value of zero hysteresis in identifying the peak corresponding to the scheduling cycle. The effective peak is the point with the highest cross-correlation amplitude within the range where the hysteresis is greater than zero. The corresponding offset is used as the time index for applying the alignment compensation to the utilization rate false jump sequence. The difference between the aligned utilization rate sequence and the compensated utilization rate false jump sequence is calculated point-by-point at each concurrent scale level. The difference sequence is the scenario-level deviation sequence. The physical meaning of each element in the scenario-level deviation sequence is the residual deviation after false deduction at that concurrent scale level. A scene-level deviation sequence value close to zero indicates that false alarms at that level have been fully captured by the utilization false alarm jump sequence. A significantly non-zero value indicates the presence of hidden false alarm components that were not captured. When there is a drift in the sampling phase of the driving layer between different batches, the alignment compensation amount needs to be re-determined before each batch test to eliminate the impact of phase shift on subtraction accuracy. When the deviation value in the scene-level deviation sequence fluctuates randomly, sampling noise is suppressed by mean filtering. When the deviation value is continuously locked in a step-like manner, false alarms caused by discrete jumps in the driving layer state machine are identified by subsequent constant value locking scans. All scene-level deviation sequences are summarized by scene category index, and the number of elements in the scene-level deviation sequence is strictly equal in length to that of the utilization sequence.

[0037] Zero-deviation persistence intervals are extracted from the scene-level deviation sequence to generate pseudo-steady-state deviation intervals. Continuous levels with absolute values ​​below 0.02 in the scene-level deviation sequence constitute candidate zero-deviation intervals. This threshold corresponds to the lower limit of the driver layer utilization measurement accuracy; deviations below this threshold are considered within the measurement noise range rather than true false residuals. The threshold setting refers to the utilization measurement resolution specified in the driver layer technical documentation. A candidate interval with a continuous length greater than or equal to 5% of the total concurrent level count is confirmed as a zero-deviation persistence interval (this proportion corresponds to the shortest continuous sampling length required for steady-state determination of the utilization sequence). If the length is insufficient, it is considered an occasional zeroing-out rather than a true steady state. The start and end concurrent level numbers and interval length of the zero-deviation persistence interval are encapsulated into interval records. All interval records within the same scene are summarized to form the pseudo-steady-state deviation interval for that scene. The record with the largest interval length in the deviation interval corresponds to the concurrent scale segment with the highest reliability of the utilization sequence data. When the proportion of the total number of coverage levels in the pseudo-steady-state deviation interval to the total number of concurrent levels is lower than the minimum sampling density requirement for linear fitting, the scenario is marked as a low-confidence scenario. The computing power anchoring result of the low-confidence scenario is downgraded to a reference value. The scenario is still included in the effective scenario candidate set, but a confidence-limited label is attached to the corresponding entry of the precision correction instruction. When allocating precision channel weights, the weight combination of this scenario is constrained by the interpolation results of adjacent effective scenarios as the upper bound. When the aggregation logic of the framework layer and the sampling rhythm of the driving layer are out of sync, the values ​​of the entire segment of the scenario-level deviation sequence are difficult to return to zero. In this state, the low-confidence label is passed to the corresponding entry of the precision correction instruction along with the computing power anchoring result.

[0038] A constant-value locking scan is performed on the scene-level deviation sequence and the pseudo-steady-state deviation interval to identify abnormal deviation scene sets. When the driver layer status register refresh stagnates, the utilization reported value is fixed at a certain value. Therefore, the deviation value of the corresponding concurrency scale segment in the scene-level deviation sequence remains constant between consecutive levels. This constant locking mode is a typical feature of fixed bias false alarms, which is completely different from the occasional zeroing pattern caused by random noise. The constant-value locking scan traverses the non-zero deviation segment outside the pseudo-steady-state deviation interval through a sliding window with a width equal to 10% of the full concurrency scale. When the range of deviation values ​​within the window is less than the lower limit of the driver layer utilization measurement resolution, it is determined to be a constant locking state. A concurrency scale segment with more than three consecutive locking windows is determined to be a constant deviation segment, and the average deviation value within this segment is the fixed bias false alarm amplitude. There may be multiple constant deviation segments in the scene-level deviation sequence. When the fixed bias false alarm amplitudes of each segment are different, it indicates that there is a multi-level register refresh delay in the driver layer. Different concurrency scales trigger different levels of register stagnation. Scenarios containing constant deviation segments and whose fixed bias false alarm magnitude exceeds twice the nominal measurement error limit of the driver layer are identified as abnormal scenarios. The scenario category number, constant deviation segment location, and fixed bias false alarm magnitude of all abnormal scenarios are summarized to form an abnormal deviation scenario set. Each entry in the abnormal deviation scenario set is accompanied by a coverage ratio field, which is obtained by dividing the number of constant deviation segment bits by the number of concurrent bits for the scenario.

[0039] Invalid scenarios are removed from the abnormal deviation scenario set to generate valid scenario markers. Scenarios in the abnormal deviation scenario set whose fixed bias false alarm magnitude exceeds twice the standard deviation of the utilization rate sequence of adjacent valid scenarios are removed from the valid scenario candidate set because the deviation of the utilization rate data exceeds the range of interpolation correction capabilities. Moderately distorted scenarios whose amplitude does not exceed this threshold and whose constant deviation segment coverage ratio is less than 20% of the full range are allowed to be retained with correction annotations. Removal decisions are made under dual constraints of the amplitude and coverage ratio fields. Scenarios are retained if both conditions are met, and removed if either condition is not met. These dual constraints prevent scenarios with moderate bias amplitude but extremely wide coverage under a single metric from being incorrectly retained. Scenarios whose precision switching frequency increases with concurrency scale have a correspondingly larger constant deviation segment coverage ratio; such scenarios are usually correctly removed under the dual constraints of amplitude and coverage ratio. Scenes outside the abnormal deviation scene set, together with the retained moderately distorted scenes, constitute the effective scene candidate set. The scene category number and effective concurrency range of each scene in the effective scene candidate set are encapsulated as effective scene marker entries. The effective concurrency range excludes the concurrency gap left after the removal of highly distorted scenes. The number of effective scene marker entries equals the number of scenes in the effective scene candidate set. The scene category number corresponding to the removed scene is marked with a gap identifier in the effective scene markers. The gap identifier also records two auxiliary fields: the scene category number of adjacent effective scenes and the concurrency distance. These are directly read during the subsequent computational power anchoring stage when performing linear interpolation to fill the gap. The distance value participates in the interpolation weight calculation to ensure a smooth transition between the gap filling result and adjacent effective scenes.

[0040] The computing power correction amount is obtained by anchoring the computing power sensitive components through effective scene marking. The row component values ​​corresponding to the effective scenes identified by the effective scene markings are free of false alarms in the computing power sensitive component matrix and can be directly anchored. The row component values ​​of invalid scenes corresponding to the gaps in the effective scene markings are contaminated by false alarms and need to be replaced by interpolation of the effective scene component values. The interpolation is based on the corresponding component values ​​read from the adjacent effective scene category numbers recorded in the effective scene markings. The weight is determined by the reciprocal of the concurrency scale distance recorded in the effective scene markings. When adjacent effective scenes are missing, the average value of the computing power sensitive components of all effective scenes is used to fill the gap. After anchoring, the scene component values ​​of the computing power sensitive components reflect the deviation of the actual computing power unit utilization rate after eliminating false alarms. The difference between the average values ​​of the components of adjacent concurrency levels before and after the precision level switch in each scene is extracted from the anchored scene component values ​​as the computing power loss amplitude for that scene. The computing power loss amplitude reflects the net decrease in computing power unit utilization caused by the precision switch. The larger the value, the more severe the disturbance to the computing power unit caused by the precision switch in that scene. The computing power loss amplitudes of all scenes are arranged by scene category to constitute the computing power correction amount. Scenarios in the computing power correction quantity whose value exceeds one standard deviation of the average computing power loss magnitude across all scenarios are marked as significant regression scenarios, corresponding to the hardware behavior of discontinuous computing power scheduling during precision switching of the corresponding AI accelerator card; scenarios in which the difference in the average number of component levels before and after switching continuously exceeds 20% of the total number of concurrent levels are marked with a high loss in the computing power correction quantity, indicating that there is a systematic waste of computing power in the precision switching mechanism of this scenario.

[0041] In some embodiments, the step of obtaining a bandwidth bottleneck anchor point by associating memory access conflicts based on the bandwidth-sensitive component and the memory access channel characteristics in the hardware configuration parameters includes: performing periodic extraction of memory access density to generate a memory access conflict feature spectrum on the memory access channel characteristics in the hardware configuration parameters; performing spectral analysis on the bandwidth-sensitive component to generate a bandwidth occupancy frequency sequence; performing frequency domain correlation matching between the bandwidth occupancy frequency sequence and the memory access conflict feature spectrum to identify conflict-associated memory access layers; and performing bandwidth bottleneck tracing and localization based on the conflict-associated memory access layers to generate bandwidth bottleneck anchor points.

[0042] Memory access channel characteristics in hardware configuration parameters are periodically extracted to generate a memory access conflict feature spectrum. On-chip memory level cache refresh and video memory data prefetch follow a fixed scheduling cycle, which exhibits regular fluctuations in bandwidth utilization time-series data. The L1 cache scheduling cycle corresponds to the operator execution time of a single inference request, the L2 cache scheduling cycle corresponds to the pipeline refresh interval of a complete inference batch, and the global video memory scheduling cycle corresponds to the video memory scheduling interval triggered by KV-cache paging. The magnitude difference between the three cycle values ​​can reach two orders of magnitude, with the L1 scheduling frequency typically being more than a hundred times higher than the global video memory frequency. After excluding the global maximum value where the hysteresis is zero, the autocorrelation function of the bandwidth utilization time-series data for each level shows an effective peak at the point where the hysteresis equals the scheduling cycle. The hysteresis corresponding to the effective peak is the memory access density cycle for that level. The calculation window length for autocorrelation analysis must ensure that it includes at least three complete scheduling cycles to ensure stable detection of the autocorrelation peak. If the window length is insufficient, the memory access density cycle for that level is marked as low confidence. The reciprocals of the three memory access density cycles (unit: sampling intervals) form the cycle frequency field (unit: sampling interval). -1 The memory access conflict feature spectrum is composed of the periodic frequency and the corresponding bandwidth upper limit. This spectrum is constructed with rows representing levels and columns representing periodic frequency and bandwidth upper limit, with a fixed number of three rows. The bandwidth upper limit field for each level in the memory access conflict feature spectrum originates from the driver layer query value of the hardware configuration parameters, reflecting the physical bandwidth upper limit of the chip design rather than the available bandwidth visible at the software layer. Due to driver scheduling overhead, there is a fixed difference between the two. The memory access conflict feature spectrum uses the physical bandwidth upper limit to ensure consistency in the conflict saturation judgment benchmark. The periodic frequency span of the memory access conflict feature spectrum dictates that subsequent frequency domain correlation matching needs to be performed on the logarithmic frequency axis to ensure equal matching sensitivity at each level. An excessively wide frequency difference threshold at low-frequency levels on the linear frequency axis will lead to mismatches.

[0043] For example, the step of performing spectral analysis on the bandwidth-sensitive component to generate a bandwidth-occupied frequency sequence includes: performing a discrete Fourier transform on the bandwidth-sensitive component to generate a bandwidth-occupied spectrum; extracting spectral flattening segments from the bandwidth-occupied spectrum to generate a flattened frequency set; performing full-band saturation identification on the flattened frequency set to generate bandwidth latent saturation feature points; and performing associated frequency band suppression on the bandwidth-occupied spectrum based on the bandwidth latent saturation feature points to generate a bandwidth-occupied frequency sequence.

[0044] A Discrete Fourier Transform (DFT) is performed on the bandwidth-sensitive component to generate a bandwidth occupancy spectrum. During batch scheduling, the inference framework periodically loads the weight matrices of multiple requests from video memory to the on-chip cache. This loading period coincides with the batch scheduling interval, generating regular bandwidth pressure pulses in the time-domain sequence of the bandwidth-sensitive component. The DFT transforms this time-domain regularity into significant peaks in the spectrum, with the peak positions shifting with the batch scheduling interval. The sequence of bandwidth-sensitive component values ​​for each scenario is expanded along the concurrency scale level, with the distance between adjacent levels constituting the sampling interval. The sampling frequency is the reciprocal of the sampling interval. Before performing the DFT, the sequence of bandwidth-sensitive component values ​​for each scenario undergoes Hanning window weighting to suppress spectral leakage. After windowing, the sequence is padded with zeros to the nearest power of 2 length to improve frequency resolution. After modulo-taking the transform output, the DC component is removed to obtain the AC amplitude spectrum at each frequency point. The AC amplitude spectra of each scenario are superimposed and averaged according to scenario category to obtain a cross-scenario comprehensive amplitude spectrum. This superposition and averaging eliminates the interference of single-scenario sampling noise on the spectral shape. The frequency resolution of the comprehensive amplitude spectrum is determined by the length of the bandwidth-sensitive component sequence after zero-padding. The bandwidth occupancy spectrum is constructed with frequency on the horizontal axis and amplitude on the vertical axis. The upper limit of the frequency axis is half of the sampling frequency, i.e., the Nyquist frequency. The bandwidth occupancy spectrum saves two sets of data: independent amplitude spectrum for each scene and comprehensive amplitude spectrum across scenes. The independent amplitude spectrum retains scene-specific frequency characteristics for scene-level conflict analysis, while the comprehensive amplitude spectrum reflects common bandwidth saturation patterns across scenes for global bottleneck location. The two sets of data are used respectively in subsequent flattening segment extraction and implicit saturation identification.

[0045] The spectral flattening segments of the bandwidth utilization spectrum are extracted to generate a flattened frequency set. When memory access bandwidth is close to saturation, the memory tier can no longer respond to higher frequency memory access requests. The bandwidth utilization tends to be evenly distributed across a wide bandwidth, and the amplitude distribution of the corresponding frequency range in the bandwidth utilization spectrum therefore exhibits a flattening characteristic. This flattening is a frequency domain indicator of bandwidth saturation, contrasting with the peak shape where amplitudes are concentrated at a few dominant frequencies under normal conditions. The spectral flattening segments are detected by scanning a sliding window with a width of 10% of the total number of frequency points in the comprehensive amplitude spectrum (the window width is chosen to cover enough frequency points to stabilize the range estimation). When the amplitude range within the window is less than 20% of the average amplitude within the window, it is identified as a flattening candidate segment (this proportion corresponds to the minimum sensitivity for judging the uniformity of amplitude distribution). When the frequency interval between adjacent flattening candidate segments is less than 3 times the frequency resolution, they are merged into a continuous flattened frequency range. All frequency points within the flattened frequency range are extracted to form a flattened frequency set. Each element contains two fields: frequency value and corresponding amplitude value. The elements are arranged in ascending order of frequency. When the flattened frequency set is empty, it indicates that memory access pressure is concentrated on a few dominant frequencies, and the AI ​​accelerator card still has a large margin in memory access bandwidth. The number of elements in the flattened frequency set increases sharply with the increase in concurrency. After KV-cache penetrates the L2 cache and enters global video memory for access, the number of elements in the flattened frequency set increases sharply. The jump position corresponds to the actual effective capacity boundary of the L2 cache. The deviation from the hardware nominal value comes from the actual eviction overhead of the cache replacement strategy. When the flattened frequency set covers more than 50% of the entire frequency band, a global bandwidth saturation warning is triggered.

[0046] The flattened frequency set is subjected to full-band saturation identification to generate bandwidth implicit saturation feature points. The difference between implicit saturation and explicit saturation is that the former has not yet triggered the hardware bandwidth limiting mechanism, but the memory access pressure has entered the non-linear growth region. The ratio of the amplitude of the corresponding frequency point in the flattened frequency set to the maximum value of the entire band is defined as the relative saturation. When it exceeds 0.6, it is judged as a saturation candidate point (this threshold is taken as the median level of the amplitude distribution of the entire band; the amplitude intensity of frequency points with a ratio below this value is insufficient to trigger the bandwidth implicit saturation judgment). The memory access bandwidth utilization intensity of the frequency points in the flattened frequency set with a relative saturation of more than 0.6 has entered the pre-saturation region. Continuing to increase the concurrency in this concurrency scale range will quickly trigger the bandwidth hard cap. Among the saturation candidate points, points with adjacent frequency intervals less than 3 times the frequency resolution are merged into a single bandwidth implicit saturation feature point. The frequency of the merged feature point is the amplitude-weighted average frequency of all points before merging, and the amplitude is the maximum value of all points before merging. Low-frequency feature points correspond to slow bandwidth accumulation and saturation due to long-cycle, low-frequency patterns. These feature points are triggered by the continuous loading pressure of the weight matrix during long-batch continuous inference services. High-frequency feature points are triggered by fast bandwidth pulses with short-cycle, high-frequency patterns. The two types of feature points are configured with bandwidth constraint weights for different concurrency scales in the precision correction instruction. If the set of latent bandwidth saturation feature points contains more than three elements, it indicates that the AI ​​accelerator card has a risk of latent bandwidth saturation across multiple frequency rhythms. Bandwidth constraint weights need to be configured for each concurrency scale in the precision correction instruction. If the number of elements is zero, it indicates that the risk of latent bandwidth saturation does not currently exist.

[0047] Based on the bandwidth implicit saturation feature points, a bandwidth occupancy frequency sequence is generated by performing correlation band suppression on the bandwidth occupancy spectrum. The amplitude of the implicit saturation band in the bandwidth-sensitive component is nonlinearly amplified by the saturation effect. When the bandwidth-sensitive component directly participates in frequency domain correlation matching, it suppresses the relative contribution of the unsaturated band, causing bandwidth conflict attribution to the implicit saturation band while ignoring the real conflict signals of other memory access levels. The correlation band suppression operation reduces the weight of the implicit saturation band amplitude in the bandwidth-sensitive component to ensure a balanced contribution of each memory access level in the matching process. The correlation band is determined by expanding the range of each bandwidth implicit saturation feature point frequency value by 10% to both sides (the expansion ratio is consistent with the frequency domain correlation matching threshold to ensure the suppression range covers the entire matching interval). The amplitude within the correlation band is multiplied by a suppression coefficient of 0.3 (suppressing the dominant effect of the implicit saturation band while retaining some amplitude contribution). Frequency points whose suppressed amplitude exceeds twice the average of the entire frequency band (corresponding to an effective peak value significantly higher than the noise floor) are extracted as effective bandwidth occupied frequencies, and arranged in descending order of amplitude to form a bandwidth occupied frequency sequence. When the frequency features of bandwidth-sensitive components are concentrated, the sequence has fewer elements; when they are dispersed, a few elements with the highest amplitudes are selected to control the scale of subsequent matching calculations. Each element includes a frequency value, the suppressed amplitude, and a flag indicating whether it has been suppressed. The suppression flag indicates that the amplitude of that frequency point has been downweighted, and its contribution weight in frequency domain association matching is correspondingly reduced. When the number of elements in the bandwidth occupied frequency sequence is less than 3, it indicates that the memory access bandwidth pressure is highly concentrated and the bandwidth bottleneck source is singular. In this case, the matching results of a few frequency points in the bandwidth occupied frequency sequence with the memory access conflict feature spectrum have high confidence, and the conflict level positioning accuracy is better than when there are more elements.

[0048] Frequency domain correlation matching is used to identify conflict-related memory access layers based on the bandwidth occupancy frequency sequence and memory access conflict feature spectrum. Frequency matching is categorized based on the absolute value of the frequency difference being less than 10% of the smaller frequency value (a relative threshold ensures that the matching sensitivity of the L1 high-frequency band and the global memory low-frequency band is equal on the logarithmic frequency axis; under a linear threshold, the matching interval for low-frequency layers will be too wide, leading to mismatches). When this condition is met, the memory access pressure fluctuation of the bandwidth-sensitive component at that frequency is highly synchronized with the scheduling cycle rhythm of the corresponding memory layer, indicating that this layer is the source layer of memory access conflicts at that frequency. The frequency matching relationship is represented by an correlation matrix. The matrix rows correspond to each frequency point in the bandwidth occupancy frequency sequence, and the matrix columns correspond to the three memory layers in the memory access conflict feature spectrum. The matrix elements are the absolute values ​​of the frequency differences at the time of frequency matching, with unmatched positions filled with infinity. The number of rows in the correlation matrix equals the number of elements in the bandwidth occupancy frequency sequence, and the number of columns is fixed at three. The minimum element in each column of the correlation matrix corresponds to the frequency point with the highest matching degree for that storage level. When the minimum element is less than the matching threshold, that level is determined to be a conflict-related memory access level. Each entry contains three fields: level name, matching frequency value, and frequency difference. When a conflict-related memory access level includes L1, L2, and global memory levels, all three levels generate bandwidth bottleneck anchor point records. The difference in the concurrent scale of each level quantifies the hierarchical buffering capacity of the AI ​​accelerator card's cache level for inference memory access pressure. When the number of conflict-related memory access level entries is zero, it indicates that the frequency characteristics of the bandwidth-sensitive component do not match the scheduling rhythm of the three storage levels, and the bandwidth pressure comes from non-periodic burst memory access. Manual verification is required in conjunction with the time-domain waveform of the bandwidth-sensitive component.

[0049] The bandwidth bottleneck is traced and located based on the conflict-correlated memory access layer, generating bandwidth bottleneck anchor points. The mapping position of the bandwidth saturation trigger moment of each memory level on the concurrency scale axis is distributed at different concurrency levels due to the difference between cache capacity and bandwidth limit. The L2 cache penetration point appears at medium concurrency levels, and the global memory bandwidth hard cap trigger point appears at high concurrency levels. The difference between the two positions directly reflects the buffering capacity of the AI ​​accelerator card's cache level for inference memory access pressure. The larger the difference, the more abundant the L2 cache capacity. The bandwidth saturation is determined by the ratio of the bandwidth limit of each memory level in the conflict-correlated memory access layer to the amplitude of the bandwidth sensitive component at the corresponding matching frequency. When the saturation exceeds 0.9, the level is determined to be a bandwidth saturation level (0.9 corresponds to the conservative lower bound of the bandwidth utilization entering the hardware speed limiting mechanism trigger range). The concurrency scale level triggered by the bandwidth saturation level is obtained by the conversion formula c_bottleneck=round(f_match / f_step), where f_match is the matching frequency value (sampling interval). -1 f_step=1 / N (N is the total number of gradient bits for concurrent operation, and the unit sampling interval) -1`round` is the rounding operation, and `c_bottleneck` is the unit of the bandwidth saturation level. The trigger position of the bandwidth saturation level on the concurrency scale axis is extracted as the bandwidth bottleneck anchor point. Each record contains three fields: level name, trigger concurrency scale, and bandwidth saturation level. Storage levels with a saturation level below 0.9 do not generate bandwidth bottleneck anchor point records. When the bandwidth bottleneck anchor point positions triggered by the same storage level differ in different scenarios, the weighted median of the trigger concurrency scale for each scenario is taken as the representative anchor point position for that level (the weighted median is more robust to extreme trigger positions in individual scenarios than the weighted mean). The weight is determined by the magnitude of each scenario in the bandwidth-sensitive component. Scenarios with larger magnitudes contribute more to bandwidth conflicts, and their trigger positions have higher weights among the representative anchor points.

[0050] Precision correction instructions are generated by allocating precision channel weights based on the computing power correction amount and bandwidth bottleneck anchor points. The computing power loss magnitude of each scenario in the computing power correction amount and the bandwidth saturation of each scenario in the bandwidth bottleneck anchor points together constitute the two-dimensional constraints for that scenario. Scenarios with high computing power loss magnitude and low bandwidth saturation shift their precision channel weights towards FP16, while scenarios with high bandwidth saturation and low computing power loss magnitude shift their weights towards INT8 to reduce memory access density and alleviate bandwidth conflicts. The precision channel weight values ​​for each scenario are determined by performing linear programming in a two-dimensional coordinate system composed of computing power loss magnitude and bandwidth saturation. The decision variables for the linear programming are the INT8 weight w_8, the FP16 weight w_16, and the FP32 weight w_32. The constraints are w_8 + w_16 + w_32 = 1, w_8 ≥ 0, w_16 ≥ 0, w_32 ≥ 0, and Σ(w_p × acc_p) ≥ acc_min (where acc_p is the inference precision index of precision level p, and acc_min is the acceptable precision for the business). The lower limit of accuracy is input from the test configuration file); the objective function is minΣ_c[T_corr(c)-T_ref(c)]², where T_corr(c) is the end-to-end latency of accuracy recalibration at concurrency level c (ms), T_ref(c) is the baseline latency of storing the benchmark map at the corresponding concurrency level (ms), acc_p and acc_min are both dimensionless inference accuracy indicators, acc_p is obtained by the inference framework running on the standard evaluation set at accuracy level p, and acc_min is input from the test configuration file. The optimal accuracy channel weight combination for each scenario is arranged according to scenario category to form a weight allocation matrix, in which the sum of INT8 weight, FP16 weight and FP32 weight in the weight allocation matrix is ​​equal to 1. The precision correction instruction encapsulates two components: the weight allocation matrix and the precision switching trigger concurrency threshold for each scenario. The precision switching trigger concurrency threshold is determined by the smaller of the bandwidth bottleneck anchor point position and the starting level of significant rollback in computing power correction (taking the smaller value ensures that precision switching is completed before any resource dimension triggers a bottleneck, following a conservative principle to avoid oversaturation). The number of precision correction instructions is equal to the number of scenario categories in the initial load test set.

[0051] Step S14: Based on the accuracy correction instruction, perform scene-by-scene accuracy recalibration on the initial load test set to generate a calibrated load test set. Based on the calibrated load test set, perform cold start path missing detection and accuracy gradient monotonicity failure identification to generate false alarm consistency blind zone anchor points. Based on the false alarm consistency blind zone anchor points, perform global optimization on the computing power benchmark map and output benchmark test results.

[0052] In some embodiments, the step of performing scene-by-scene precision recalibration on the initial load test set according to the precision correction instruction to generate a calibrated load test set includes: extracting computing power correction components and bandwidth correction components based on the precision correction instruction to generate a fractal dimension correction vector; superimposing the loads of each scenario in the initial load test set according to the fractal dimension correction vector to generate a scene-by-scene recalibrated load; performing sign inversion detection between adjacent scenarios to eliminate correction oscillations and generate an oscillation-suppressed load sequence on the scene-by-scene recalibrated load; and performing global load consistency verification based on the oscillation-suppressed load sequence to generate a calibrated load test set.

[0053] Based on the precision correction instructions, the computing power correction component and bandwidth correction component are extracted to generate a fractal correction vector. In the mixed precision deployment mode, the contribution directions of computing power bottleneck and bandwidth bottleneck to end-to-end latency of the inference service are independent. Each column of the weight allocation matrix encapsulated in the precision correction instructions corresponds to a different precision level. The weight contribution of each scenario in the computing power sensitive direction is restored from the weight allocation matrix. The weight value of each column is multiplied by the computing power utilization correction coefficient of the corresponding precision level and then summed to form the computing power correction component of that scenario. The computing power utilization correction coefficient is defined as the value of the scenario component in the computing power correction amount divided by the rated peak computing power of the precision level. It reflects the net change in computing power unit utilization caused by the unit weight adjustment. This component reflects the net change in computing power unit utilization after the weight adjustment of the precision level. The precision switching trigger concurrency threshold encapsulated in the precision correction instruction implicitly contains bandwidth constraint boundary information. A bandwidth margin estimate is extracted from the difference between this threshold and the bandwidth saturation trigger concurrency scale determined during the precision correction instruction generation stage. This estimated bandwidth margin is weighted by the memory access density coefficient for each precision level to form the bandwidth correction component for that scenario. The memory access density coefficient is defined as the ratio of the number of bytes accessed per inference attempt to the corresponding peak computing power for each precision level, expressed in bytes per FLOP. The memory access density coefficient for the INT8 precision level is significantly lower than that for FP16 due to its halved data width. The number of bytes accessed per inference attempt is determined by the sum of the weight matrix loading and KV-cache read / write operations of the inference framework layer at each precision level. The computing power correction component and the bandwidth correction component are concatenated to form a fractal correction vector for each scenario. This fractal correction vector is a two-dimensional vector, with the first dimension being the computing power correction component and the second dimension being the bandwidth correction component. Each component carries a positive or negative sign; a positive value indicates that the resource utilization of that dimension has increased due to precision recalibration, while a negative value indicates a decrease. The fractal correction vectors of all scenarios are arranged according to scenario category to form a fractal correction matrix. The number of rows in the fractal correction matrix is ​​equal to the number of scenario categories in the initial load test set, and the number of columns is fixed at two.

[0054] A scenario-by-scenario recalibrated load is generated by superimposing the loads of each scenario in the initial load test set based on the fractal correction vector. The end-to-end latency field recorded in each scenario of the initial load test set serves as the base load value. The first dimension of the fractal correction vector, the computing power correction component, is multiplied by the computing power unit occupancy latency coefficient under the current concurrency scale of that scenario and then superimposed onto the base load value. The computing power unit occupancy latency coefficient is defined as the impact of a unit computing power correction component on the end-to-end latency, determined by the least-squares linear fit slope of the sequence of end-to-end latency values ​​of the computing power-sensitive component for each recorded scenario in the initial load test set. The second dimension of the fractal correction vector, the bandwidth correction component, is multiplied by the memory access latency coefficient under the current concurrency scale of that scenario and then superimposed onto the intermediate load value. The memory access latency coefficient is defined as the impact of a unit bandwidth correction component on the end-to-end latency, determined by the least-squares linear fit slope of the sequence of end-to-end latency values ​​of the bandwidth-sensitive component for each recorded scenario in the initial load test set. The latency value after superposition is used as the recalibration latency of that record. The recalibration latency, along with the original input sequence length, batch size, and precision type fields, is encapsulated into a recalibration record. The recalibration results of all records for each scenario are arranged by scenario category to form a scenario-by-scenario recalibration load. The changing trend of the recalibration latency sequence of each scenario in the scenario-by-scenario recalibration load should be consistent with the changing trend of the computing power benchmark value of the corresponding scenario in the computing power benchmark map. Records with inconsistent directions are marked as superposition direction anomalies. Superposition direction anomalies usually appear in scenarios where the signs of the fractal correction vectors are opposite. The superposition direction anomaly information is passed along with the recalibration record to the subsequent sign inversion detection stage between adjacent scenarios for reference in determining oscillation segments.

[0055] The recalibration load is processed scene-by-scene, and sign inversion detection between adjacent scenes is performed to eliminate calibration oscillations and generate an oscillation-suppressed load sequence. Adjacent scenes are defined as two consecutive scene categories arranged in the order of computing power saturation, bandwidth saturation, memory limitation, and balanced load in the load anchor sequence. Oscillation detection is performed pairwise in this order. Calibration oscillation refers to the phenomenon where the recalibration delay change direction of adjacent scenes repeatedly alternates between positive and negative on the concurrency scale axis. The fractal correction vector signs of computing power saturation scenes and bandwidth saturation scenes are often in opposite directions; the computing power correction component is positive and the bandwidth correction component is negative in the former, while the latter is in the opposite direction. The superposition of the two scenes compensates for each other at the same concurrency scale, causing oscillations. Sign inversion detection compares the sign of the recalibration delay difference between scene i and scene i+1 at each concurrency scale level. When the sign of the difference is opposite to that of the previous level, it is recorded as a sign inversion event. A concurrency scale segment with more than three consecutive sign inversion events is determined to be an oscillation segment. The lower bound of three excludes occasional jitter caused by a single sign inversion. Within the oscillation zone, the recalibration latency at each concurrency level is replaced by the arithmetic mean of the recalibration latency at the corresponding level in adjacent scenarios. After averaging, the recalibration latency of the two scenarios within the oscillation zone tends to be consistent, reflecting that the accuracy recalibration effect of the two scenarios is similar in this concurrency level. Outside the oscillation zone, the scene-by-scene recalibration load records remain unchanged. The averaged records within the oscillation zone are merged with the original records and arranged by scene category and concurrency level to form an oscillation suppression load sequence. Each record contains four fields: scene category, batch size, accuracy type, and recalibration latency. The total number of records is equal to the scene-by-scene recalibration load.

[0056] A global load consistency check is performed based on the oscillation suppression load sequence to generate a corrected load test set. The global load consistency check is executed sequentially in three dimensions according to priority; if a previous dimension fails, it is marked as an anomaly and not proceeds to subsequent dimensions. The first dimension checks the monotonicity of the recalibration latency sequence for each scenario as concurrency increases. Non-monotonic positions correspond to correction stacking directions at that concurrency level that contradict physical expectations; this has the highest priority because latency monotonicity is a fundamental physical constraint. The second dimension checks whether the recalibration latency order across scenarios at the same concurrency level is consistent with the computing power benchmark value order in the corresponding concurrency level column of the computing power benchmark map. Inconsistent ordering indicates a systematic deviation in the correction stacking magnitude of a certain scenario compared to other scenarios. The third dimension checks whether the distribution ratio of each precision type field in the oscillation suppression load sequence matches the weighting ratio of the precision correction instruction weight allocation matrix. If the deviation exceeds the floating-point truncation error tolerance limit when generating the precision correction instruction weight allocation matrix, it is determined to be a precision allocation inconsistency. Records that pass all three dimensions are marked as consistent and qualified. Records that fail any dimension have their latency field replaced by the median recalibrated latency of consistent and qualified records in the same scenario and concurrency level. After replacement, the record re-enters the first dimension for verification. If it still fails, the replacement value is retained and a secondary anomaly label is added, and it is no longer checked in a loop. After replacement, all scenario records are arranged by scenario category to form a calibration load test set. Each record retains five fields: scenario category, batch size, calibration accuracy type, calibration latency value, and consistency status label. The consistency pass rate is defined as the number of consistent and qualified records divided by the total number of records. When the consistency pass rate is lower than the minimum effective record ratio required for cross-scenario comparison, the entire calibration load test set is marked with low confidence. The minimum effective record ratio is determined by the number of scenario categories and the three verification dimensions. The evaluation conclusion of the corresponding scenario in the benchmark test results is downgraded accordingly.

[0057] Based on the calibration load test set, cold start path missing detection and accuracy gradient monotonicity failure identification are performed to generate false alarm consistency blind zone anchor points. When the AI ​​accelerator card starts the inference service for the first time, it needs to complete three stages: operator compilation, weight loading, and accuracy level initialization. The sum of the latency of the three stages constitutes the cold start overhead. When the cold start path is complete, the calibration latency value of the lowest concurrency level in the calibration load test set is significantly higher than that of the subsequent levels. When the cold start path is missing, there is no significant difference between the two. Cold start path missing detection is achieved by calculating the cold start ratio R_cold=T_corr(c0) / T_corr(c0+1), where c0 is the lowest concurrency level number in the corrected load test set for this scenario, and T_corr(c0) and T_corr(c0+1) are the corrected latency (ms) of the lowest and second highest concurrency levels, respectively. When R_cold is lower than 1.3, it is determined that the cold start path is missing. 1.3 is taken from the nominal lower bound ratio of cold start overhead to steady-state inference latency in the AI ​​accelerator card driver layer technical document. A ratio lower than this indicates that the initialization phase is truncated by the test framework or the cache is hit, causing the overhead to be skipped. Precision gradient monotonicity failure is defined as an abnormal state in which the expected value of the correction latency does not monotonically increase as the precision level progresses from INT8 to FP16 and then to FP32. Monotonicity failure indicates a directional error in the estimation of the computing power correction amount at a certain precision level. Precision gradient monotonicity failure identification is achieved by checking the sorting relationship of the correction latency values ​​of the three precision levels under each concurrency level in each scenario. When the sorting violates monotonically increasing, the scenario category and concurrency level of the violation location are recorded. Locations where R_cold is lower than 1.3 and the location of precision gradient monotonicity failure are marked as blind zone candidate points in the scenario-concurrency level two-dimensional space of the correction load test set. The scenario category number, concurrency level, and defect type of each blind zone candidate point are encapsulated and summarized to form false alarm consistency blind zone anchor points. False alarm consistency blind zone anchor points drive the blind zone correction of the subsequent computing power benchmark map and directly affect the effective coverage of the benchmark test results.

[0058] In some embodiments, the step of performing global optimization on the computing power benchmark map based on the false alarm consistency blind zone anchor point and outputting benchmark test results includes: performing scene accuracy test value reprojection on the computing power benchmark map through the false alarm consistency blind zone anchor point to generate a reprojection test set; performing reprojection error convergence rate detection on the reprojection test set to identify false alignment regions that converge too quickly; performing test value resampling to eliminate local extremum locking for the false alignment regions to generate an optimized test map; and performing global benchmark unification on the optimized test map to output benchmark test results.

[0059] A reprojection test set is generated by reprojecting scenario accuracy test values ​​onto the computing power benchmark graph using false consistency blind zone anchors. The blind zone cells marked by the false consistency blind zone anchors correspond to computing power benchmark values ​​in the computing power benchmark graph affected by cold start deficiencies or accuracy gradient failures. Cold start deficiencies cause the values ​​of cells in the first concurrency level to be lower, while accuracy gradient failures cause the values ​​of cells in specific accuracy levels to be higher or lower. The reprojection operation maps each blind zone cell to the cell with the closest weighted Euclidean distance within the effective region. The effective region is defined as the set of cells in the computing power benchmark graph that are neither marked by false consistency blind zone anchors nor by subsequent false alignment detections. The mapping distance is calculated by weighting the scenario category coding difference and the concurrency level difference. The scenario category weight is set to 2, and the concurrency level weight is set to 1. Scenario categories are assigned values ​​from 0 to 3 in the order of the load anchor sequence, and the absolute value of the coding difference is used in the calculation. The weight ratio of 2:1 reflects that the hardware constraint error introduced by scenario category mismatch is significantly greater than the load difference introduced by concurrency level mismatch. The benchmark computing power value of the blind zone cell is replaced by the benchmark computing power value of the valid target cell to form the reprojected value. Each entry in the reprojection test set records four fields: scenario category, concurrency level, original benchmark computing power value, and reprojected value. These four fields are directly used in the subsequent convergence rate detection and resampling stages. The original benchmark computing power value and reprojected value fields of the valid cells are the same, and the blind zone type field is marked as valid. The reprojection test set covers all cells of the benchmark computing power map, and there are no longer any blind zone gaps in the benchmark computing power sequences of each scenario, ensuring full scenario coverage for subsequent benchmark test results.

[0060] The reprojection error convergence rate is used to detect and identify false alignment regions on the reprojection test set. The reprojection error E(s,c) is defined as the absolute value of the difference between the reprojection value at scene s and concurrency level c in the reprojection test set and the original computing power baseline value. The reprojection error of non-blind zone cells is zero, while the reprojection error of blind zone cells reflects the blind zone correction magnitude. The reprojection error convergence rate V(s,c) is defined as the decrease in reprojection error between adjacent concurrency levels, calculated as V(s,c) = E(s,c-1) - E(s,c). A positive V(s,c) indicates error convergence, while a negative V(s,c) indicates error amplification. Within the normal correction region, the V(s,c) value is relatively stable. The blind zone correction value smoothly transitions horizontally along the concurrency level axis of the effective region, and the transition gradient is jointly determined by the blind zone width and the gradient of the computing power baseline value of adjacent effective regions. False alignment occurs when a blind zone cell is mapped to a valid cell with a similar actual computing power level. In this case, V(s,c) spikes sharply at the blind zone boundary and then quickly returns to zero, a stark contrast to the gradual decline of true convergence. Different AI accelerator cards can exhibit similar normalized baseline values ​​within the false alignment region, masking the true performance differences through local alignment. The convergence rate threshold is set to three times the mean convergence rate of the reprojection error sequence for that scene (three times is taken from a commonly used upper bound for outlier identification in statistics; local convergence rates exceeding this value significantly deviate from the normal correction gradient). Continuous segments exceeding the threshold are identified as false alignment regions. The scene affiliation and concurrency scale boundaries of each region are encapsulated as region records, and all region records are aggregated to form a set of false alignment regions. Regions in the false alignment region set whose width exceeds the average span of the false reporting consistency blind zone anchor point are marked as wide false alignment regions. Wide false alignment regions indicate that the computing power baseline value of this AI accelerator card has systematic distortion within a wider concurrency scale range, and the corresponding concurrency scale segment is marked with a systematic false reporting risk label in the benchmark test results.

[0061] To eliminate local extremum locking and generate an optimized test map, test values ​​are resampled for false alignment regions. Local extremum locking is defined as the state where the benchmark computing power value at a certain concurrency level within the false alignment region of the reprojected test set is skewed and locked near that extremum by the extremum of adjacent valid cells. Extremum locking causes the benchmark computing power value within the false alignment region to abnormally concentrate around a certain value, thus losing its responsiveness to changes in concurrency level. The resampling operation performs perturbation sampling on the benchmark computing power values ​​at each concurrency level within the false alignment region. The perturbation amount is determined by the formula ΔB=σ_valid×rand(-1,1)×α, where σ_valid is the standard deviation of the benchmark computing power value in the valid region of the scenario (strips / ms), rand(-1,1) is a uniformly random number between -1 and 1, and α is a dimensionless perturbation coefficient set to 0.15. This value strikes a balance between introducing sufficient perturbation to break the extremum locking and avoiding disrupting the overall monotonic trend. ΔB has the same dimensions as σ_valid (strips / ms). The perturbed benchmark value is superimposed on the original reprojection value to form a resampled value. This resampled value replaces the reprojection value of the corresponding cell in the false alignment region of the reprojection test set. After resampling, monotonicity verification is performed on the optimized test graph to check whether the benchmark sequence after resampling in each scenario satisfies the same monotonic trend as the original effective region. When the perturbation causes local destruction of monotonicity, the resampled value of the corresponding level is corrected to a linear interpolation of the adjacent effective level value to restore monotonicity. This correction only applies to a few levels where the perturbation amplitude exceeds the difference between adjacent effective levels and does not affect the effect of eliminating local extreme value locking in other levels. Only after the monotonicity verification is passed can the global benchmark unification stage be entered. The cell values ​​of the reprojection test set outside the false alignment region are directly copied to the optimized test graph without perturbation processing. The dimensions of the optimized test graph are completely consistent with the benchmark graph, with rows corresponding to scenario categories and columns corresponding to concurrency scale levels. The optimized test graph includes a false alignment region coverage location marker, and the marker field records whether each cell has been resampled and the corresponding perturbation amplitude ΔB value.

[0062] The benchmark test results are output using a global benchmark unification method for the optimized test graph. First, cross-scenario dimensional alignment is performed on the computing power benchmark sequences for each scenario in the optimized test graph. The normalization divisor is then taken as the peak cell value within the effective region of each scenario. The effective region is defined as the set of cells in the optimized test graph that are neither marked by falsely reported consistency blind zone anchor points nor by falsely aligned region sets. After normalization, the computing power benchmark value for each scenario is uniformly expressed as a proportion relative to the effective peak value of that scenario, eliminating the interference of cross-scenario dimensional inconsistencies on horizontal comparisons. After normalization, the value range of each cell in the optimized test graph converges to the interval between 0 and 1. Cells with values ​​close to 1 correspond to the AI ​​accelerator card operating close to the effective computing power limit under that scenario and concurrency scale, while cells with values ​​significantly lower than 1 indicate computing power loss under that condition. The benchmark test results use the normalized optimized test graph as the core matrix, supplemented by four auxiliary fields: blind zone anchor point correction records, false alignment region annotations, resampling perturbation amplitude distribution, and global consistency verification pass rate. The global consistency verification pass rate is calculated by summarizing the consistency status markers of each record in the calibrated load test set, and is equal to the number of consistent records divided by the total number of records. These fields together constitute a multi-dimensional quantitative description of the AI ​​accelerator card's inference computing power. The extreme values ​​within the effective region of each scenario correspond to the optimal concurrency scale for that scenario, which is listed by scenario in the auxiliary fields of the benchmark test results. The difference in normalized computing power benchmark values ​​for the same scenario and the same concurrency scale level for different AI accelerator card models directly reflects the true computing power difference between the two devices after eliminating backoff effects, blind zone interference, and false alignment.

[0063] To implement the AI ​​accelerator card inference computing power benchmark test method corresponding to the above method embodiments, in order to achieve the corresponding functions and technical effects. See [link to documentation]. Figure 2 , Figure 2 This diagram illustrates a structural block diagram of an AI accelerator card inference computing power benchmark testing device 200 provided in an embodiment of this application. For ease of explanation, only the parts relevant to this embodiment are shown. The AI ​​accelerator card inference computing power benchmark testing device 200 provided in this embodiment includes: Anchor point annotation unit 201 is used to obtain the inference load data and hardware configuration parameters of the AI ​​accelerator card, and perform joint annotation of the inference load data and the hardware configuration parameters with business scenario features to generate a load anchor point sequence. The graph construction unit 202 is used to generate an initial load test set by performing cross-scenario load transfer based on the load anchor sequence, perform concurrent throughput backoff measurement on the inference load data to generate a throughput deviation map, and construct a computing power benchmark graph by fusing the initial load test set and the throughput deviation map with backoff segment size constraints. The deviation correction unit 203 is used to perform directional decomposition and identification of computing power sensitive components and bandwidth sensitive components on the cumulative accuracy deviation in the initial load test set, perform accuracy level switching and utilization back-off identification on the computing power sensitive components to obtain computing power correction amount, perform memory access conflict association based on the bandwidth sensitive components and the memory access channel characteristics in the hardware configuration parameters to obtain bandwidth bottleneck anchor points, and perform accuracy channel weight allocation based on the computing power correction amount and the bandwidth bottleneck anchor points to generate accuracy correction instructions. The benchmark optimization unit 204 is used to perform scene-by-scene precision recalibration on the initial load test set according to the precision correction instruction to generate a calibrated load test set, perform cold start path missing detection and precision gradient monotonicity failure identification on the calibrated load test set to generate false alarm consistency blind zone anchor points, and perform global optimization on the computing power benchmark map based on the false alarm consistency blind zone anchor points to output benchmark test results.

[0064] The AI ​​accelerator card inference computing power benchmark testing device 200 described above can implement the AI ​​accelerator card inference computing power benchmark testing method of the above method embodiments. The options in the above method embodiments are also applicable to this embodiment, and will not be detailed here. The remaining content of this application embodiment can be referred to the content of the above method embodiments, and will not be repeated in this embodiment.

[0065] The purpose of the above embodiments is to reproduce and derive the technical solution of the present invention by way of example, and to fully describe the technical solution, purpose and effect of the present invention. The purpose is to enable the public to have a more thorough and comprehensive understanding of the disclosure of the present invention, and not to limit the scope of protection of the present invention.

Claims

1. A benchmark testing method for the inference computing power of an AI accelerator card, characterized in that, include: Obtain the inference load data and hardware configuration parameters of the AI ​​acceleration card, and perform joint annotation of the inference load data and the hardware configuration parameters with business scenario features to generate a load anchor sequence; Based on the load anchor sequence, cross-scenario load transfer is performed to generate an initial load test set. Concurrent throughput backoff measurement is then performed on the inference load data to generate a throughput deviation map. This includes: performing scenario-by-scenario concurrent scale throughput detection on the inference load data to generate a scenario-by-scenario throughput set; performing linearity deviation evaluation calculation on the scenario-by-scenario throughput set to generate a scenario-by-scenario pseudo-linear parameter set; extracting and identifying overfitting segments and pseudo-validation distribution areas for the scenario-by-scenario pseudo-linear parameter set; calibrating the throughput deviation direction based on the pseudo-validation distribution area and the scenario-by-scenario throughput set to generate a throughput deviation map; and constructing a computing power benchmark map by fusing the initial load test set and the throughput deviation map with backoff segment scale constraints. The cumulative accuracy deviation in the initial load test set is decomposed directionally to identify computing power sensitive components and bandwidth sensitive components. The computing power sensitive components are subjected to accuracy level switching and utilization back-off identification to obtain computing power correction amount. Based on the bandwidth sensitive components and the memory access channel characteristics in the hardware configuration parameters, memory access conflict is correlated to obtain bandwidth bottleneck anchor points. Accuracy channel weight allocation is performed according to the computing power correction amount and the bandwidth bottleneck anchor points to generate accuracy correction instructions. According to the accuracy correction instruction, the initial load test set is recalibrated scene by scene to generate a calibrated load test set. Based on the calibrated load test set, cold start path missing detection and accuracy gradient monotonicity failure identification are performed to generate false alarm consistency blind zone anchor points. Based on the false alarm consistency blind zone anchor points, the computing power benchmark map is globally optimized and the benchmark test results are output.

2. The method according to claim 1, characterized in that, The step of constructing a computing power benchmark map by fusing the initial load test set with the throughput deviation map and the backoff segment size constraint includes: Extract the scale feature points of the backtracking section from the throughput deviation map to generate a scale anchor point set; A scale-constrained residual field is generated by constructing constrained residuals from the scale anchor point set and the initial load test set. Based on the scale-constrained residual field, test blind zone labels are generated for scenarios where the residual is continuously zero. Based on the test blind zone labels, a computing power benchmark map is constructed by fusing backtracking segment size constraints.

3. The method according to claim 1, characterized in that, The step of performing precision level switching and utilization backoff identification on the computing power sensitive component to obtain the computing power correction amount includes: Extract the utilization rate of each scene's precision level from the computing power sensitive components to generate a utilization rate sequence; The utilization rate sequence is subjected to instantaneous jump detection during switching to generate a false utilization rate jump sequence; Based on the utilization rate sequence and the utilization rate false alarm jump sequence, perform scene-by-scene false alarm tracing detection to generate valid scene labels; The computing power correction amount is obtained by anchoring the computing power sensitive component through the effective scene marker.

4. The method according to claim 1, characterized in that, The process of obtaining bandwidth bottleneck anchors by associating memory access conflicts based on the bandwidth-sensitive component and the memory access channel characteristics in the hardware configuration parameters includes: The memory access channel characteristics in the hardware configuration parameters are periodically extracted to generate a memory access conflict feature spectrum. Spectral analysis is performed on the bandwidth-sensitive components to generate a bandwidth occupancy frequency sequence; Frequency domain correlation matching is performed between the bandwidth occupancy frequency sequence and the memory access conflict feature spectrum to identify conflict-related memory access layers; Based on the conflict-related memory access layer, bandwidth bottleneck tracing and location are implemented to generate bandwidth bottleneck anchor points.

5. The method according to claim 1, characterized in that, The step of performing scene-by-scene accuracy recalibration on the initial load test set according to the accuracy correction instruction to generate a calibrated load test set includes: Based on the accuracy correction instruction, the computing power correction component and bandwidth correction component are extracted to generate a fractal correction vector; Based on the fractal correction vector, the loads of each scenario in the initial load test set are superimposed to generate a scenario-by-scenario recalibration load; The scene-by-scene recalibration load is subjected to sign inversion detection between adjacent scenes to eliminate calibration oscillations and generate an oscillation-suppressed load sequence. Based on the oscillation-suppressed load sequence, a global load consistency check is performed to generate a corrected load test set.

6. The method according to claim 1, characterized in that, The step of performing global optimization on the computing power benchmark map based on the false inconsistency blind zone anchor point and outputting benchmark test results includes: The scenario accuracy test values ​​of the computing power benchmark map are reprojected using the false reporting consistency blind zone anchor point to generate a reprojection test set. The reprojection test set is subjected to reprojection error convergence rate detection to identify false alignment regions that converge too quickly; For the false alignment region, test values ​​are resampled to eliminate local extremum locking and generate an optimized test map; The optimized test map is then subjected to global benchmark unification to output benchmark test results.

7. The method according to claim 3, characterized in that, The step of generating valid scene labels by performing scene-by-scene false alarm source tracing detection based on the utilization rate sequence and the utilization false alarm jump sequence includes: The utilization rate sequence and the utilization rate false jump sequence are synchronized and aligned on a scene-by-scene basis to generate a scene-level deviation sequence. Extract the zero-deviation duration interval of the scene-level deviation sequence to generate a pseudo-steady-state deviation interval; Perform constant value locking scans on the scene-level deviation sequence and the pseudo-steady-state deviation interval to identify abnormal deviation scene sets; Based on the set of abnormal deviation scenarios, invalid scenarios are removed to generate valid scenario markers.

8. The method according to claim 4, characterized in that, The step of generating a bandwidth-occupied frequency sequence by performing spectral analysis on the bandwidth-sensitive components includes: Perform a Discrete Fourier Transform on the bandwidth-sensitive components to generate a bandwidth occupancy spectrum. Extract the spectrum flattening segments from the bandwidth occupancy spectrum to generate a flattened frequency set; Full-band saturation identification is performed on the flattened frequency set to generate bandwidth implicit saturation feature points; Based on the bandwidth latent saturation feature points, the bandwidth occupancy spectrum is correlated with frequency band suppression to generate a bandwidth occupancy frequency sequence.

9. A benchmark testing device for AI accelerator card inference computing power, characterized in that, include: Anchor point annotation unit is used to obtain inference load data and hardware configuration parameters of AI accelerator card, and perform joint annotation of business scenario features on the inference load data and hardware configuration parameters to generate load anchor point sequence; The graph construction unit is used to generate an initial load test set by performing cross-scenario load transfer based on the load anchor sequence, and to generate a throughput deviation graph by performing concurrent throughput backoff measurement on the inference load data. This includes: performing scenario-by-scenario concurrent scale throughput detection on the inference load data to generate a scenario-by-scenario throughput set; performing linearity deviation evaluation calculation on the scenario-by-scenario throughput set to generate a scenario-by-scenario pseudo-linear parameter set; extracting and identifying overfitting segments and pseudo-validation distribution areas for the scenario-by-scenario pseudo-linear parameter set; calibrating the throughput deviation direction based on the pseudo-validation distribution area and the scenario-by-scenario throughput set to generate a throughput deviation graph; and constructing a computing power benchmark graph by fusing the initial load test set and the throughput deviation graph with backoff segment scale constraints. The deviation correction unit is used to perform directional decomposition and identification of computing power sensitive components and bandwidth sensitive components on the cumulative accuracy deviation in the initial load test set, perform accuracy level switching and utilization back-off identification on the computing power sensitive components to obtain computing power correction amount, perform memory access conflict association based on the bandwidth sensitive components and the memory access channel characteristics in the hardware configuration parameters to obtain bandwidth bottleneck anchor points, and perform accuracy channel weight allocation based on the computing power correction amount and the bandwidth bottleneck anchor points to generate accuracy correction instructions. The benchmark optimization unit is used to perform scene-by-scene precision recalibration on the initial load test set according to the precision correction instruction to generate a calibrated load test set, perform cold start path missing detection and precision gradient monotonicity failure identification on the calibrated load test set to generate false alarm consistency blind zone anchor points, and perform global optimization on the computing power benchmark map based on the false alarm consistency blind zone anchor points to output benchmark test results.

Citation Information

Patent Citations

  • AI chip test parameter adaptive optimization method based on deep learning

    CN120872714A

  • One-stop large model agent development, operation and maintenance platform integrating computing power scheduling and model management

    CN122019199A