Neural network reasoning performance analysis method for NPU computing architecture

Through standardized processing and multi-dimensional performance analysis of the NPU computing architecture, the problems of inconsistent data quality and incomplete feature extraction in existing technologies have been solved, more accurate performance analysis and optimization have been achieved, and the stability and efficiency of neural network reasoning have been improved.

CN120611352AActive Publication Date: 2025-09-09ZHONGYING QINGCHUANG TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202511097657.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-06
Publication Date
2025-09-09
Estimated Expiration
2045-08-06

AI Technical Summary

Technical Problem

The existing neural network inference performance analysis method of the NPU computing architecture lacks standardized processing methods, cannot comprehensively and accurately extract computing unit characteristics and task characteristics, is difficult to perform multi-dimensional performance analysis, and lacks real-time monitoring and optimization methods, affecting the reliability of analysis decisions and the stability of neural network inference.

Method used

By standardizing the raw data of the NPU computing architecture, extracting computing unit characteristics, task characteristics, and architecture load characteristics, generating performance feature vectors, and performing multi-dimensional performance analysis, an enhanced computing feature space is constructed for computing matching and parameter optimization, and real-time monitoring, anomaly detection, and adaptive learning are performed.

Benefits of technology

It improves data quality and the scientific nature of analytical decisions, ensures the feasibility and stability of optimization plans, and enhances the efficiency and reliability of neural network reasoning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120611352A_ABST
    Figure CN120611352A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of neural network reasoning performance analysis, and discloses an NPU computing architecture-oriented neural network reasoning performance analysis method, which comprises the following steps of: acquiring and standardizing original data, extracting computing unit, task and architecture load characteristics to generate performance characteristic vectors, and identifying reasoning performance requirements to obtain a preliminary analysis result; integrating the preliminary result and the standardized data to generate an enhanced analysis matrix, analyzing the performance in multiple dimensions, calculating a necessity score and an architecture influence value, and obtaining decision data through multiple verifications; combining an NPU calculation feature library and decision data to construct an enhanced calculation feature space, performing calculation matching and parameter optimization, optimizing a reasoning path and performing pre-inspection; and collecting execution data streams and historical monitoring data to generate a comprehensive monitoring data packet, and carrying out anomaly detection, early warning and dynamic optimization. According to the method, the accuracy and efficiency of network reasoning performance analysis are improved, and dynamic optimization and adaptive adjustment of the reasoning process are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of neural network reasoning performance analysis, and specifically to a neural network reasoning performance analysis method for NPU computing architecture. Background Art

[0002] With the advancement of intelligent transformation and upgrading in power plants, power plant management relies on information management systems to achieve precise monitoring and control of power plant operations and management. With the rapid development of artificial intelligence technology, neural networks are one of its core technologies. Introducing neural networks into information management systems can meet the intelligent management requirements of power plant management.

[0003] As a hardware architecture specifically designed for neural network computing, the NPU (Neural Network Processing Unit) possesses unique computing modes and architectural features. However, current neural network inference performance analysis methods for NPU computing architectures have numerous shortcomings. For one thing, existing methods lack effective standardized processing methods for processing raw data from NPU computing architectures, resulting in uneven data quality and affecting subsequent analysis results. Furthermore, the feature extraction process often fails to comprehensively and accurately extract computing unit features, task features, and architecture load features, resulting in incomplete performance feature vectors that fail to accurately reflect the actual neural network inference performance.

[0004] During performance analysis, most existing methods lack multi-dimensional performance analysis capabilities, making it impossible to comprehensively analyze neural network inference performance from multiple perspectives. Furthermore, when calculating the necessity score and architectural impact value for inference performance analysis, there is a lack of scientific and rational calculation methods, resulting in low reliability of analytical decision data. During optimization, existing methods also struggle to effectively match calculations and optimize parameters based on the characteristics of the NPU computing architecture, making it impossible to develop an optimal inference solution.

[0005] Real-time monitoring and anomaly detection are also crucial components of neural network reasoning. However, existing methods lack effective data fusion and analysis methods when collecting execution data streams and historical monitoring data, making it difficult to detect anomalies and perform dynamic optimization in a timely manner, thus affecting the stability and reliability of neural network reasoning.

[0006] Existing neural network inference performance analysis methods for NPU computing architectures have certain shortcomings in data processing, feature extraction, performance analysis, optimization decision-making, and real-time monitoring. They cannot meet the current demand for neural network inference performance analysis in the development of artificial intelligence technology. Therefore, a new neural network inference performance analysis method for NPU computing architectures is urgently needed to address these issues. Summary of the Invention

[0007] The purpose of the present invention is to provide a neural network reasoning performance analysis method for NPU computing architecture to solve the problems raised in the above background technology.

[0008] To achieve the above objectives, the present invention provides the following technical solution: a neural network reasoning performance analysis method for NPU computing architecture, the method comprising: S1. Obtain the original data of the NPU computing architecture and standardize it to obtain standardized computing data. Based on the standardized computing data, extract the computing unit characteristics, computing reasoning task characteristics, and architecture load characteristics to generate a performance feature vector. Based on the performance feature vector, identify the reasoning performance requirements and determine the analysis dimensions to obtain preliminary analysis result data. S2. Integrate the preliminary analysis result data and the standardized calculation data to generate an enhanced analysis matrix; perform multi-dimensional performance analysis based on the enhanced analysis matrix to obtain performance analysis result data; analyze the performance analysis result data to calculate the reasoning performance analysis necessity score and architectural impact value to form analysis decision data; perform multiple verifications on the analysis decision data to generate verified decision data; S3. Based on the pre-stored original information of the NPU computing feature library and the verified decision data, an enhanced computing feature space is constructed; based on the enhanced computing feature space, computing matching and parameter optimization are performed to generate an optimized computing selection plan; the reasoning path of the optimized computing selection plan is optimized to form an optimized reasoning plan; based on the optimized reasoning plan, a multi-dimensional pre-inspection is performed before analysis, and finally the pre-inspection report data is output.

[0009] Preferably, step S1 is further: S11. Obtaining NPU computing architecture raw data including computing unit configuration, timestamp, core identifier, and task identifier; converting the computing unit configuration in the NPU computing architecture raw data into a unified coding format, removing outliers and redundant information, and obtaining processed input data; performing data length normalization processing based on the processed input data to generate standardized computing data; S12. Based on the standardized computing data, calculate the unit parameter values ​​of the computing mode, extract the key computing feature set, identify the reasoning task type, and generate basic feature data; based on the basic feature data, use a preconfigured task analysis model to calculate the performance requirement probability distribution and task vector of the reasoning task to obtain task feature data; obtain the historical records of the NPU, and extract the architecture status information based on the historical records; combine the architecture status information with the task feature data to form a performance feature vector; S13. Based on the performance feature vector, use the preset feature-requirement mapping rules to identify the reasoning performance requirement type, calculate the requirement priority score, and generate the requirement feature data; analyze the requirement feature data, estimate the computing resource requirement value and the time window requirement value, and obtain the resource requirement data; based on the requirement feature data and resource requirement data, identify the dependency relationship between the requirements and construct a dependency relationship graph; based on the reasoning performance requirement type, priority score, resource requirement data and dependency relationship graph, generate preliminary analysis result data.

[0010] Preferably, step S2 is further: S21. Converting the preliminary analysis result data into a feature matrix, converting the standardized calculation data into a vector representation, and combining the feature matrix and the vector representation to form an initial analysis matrix; extracting relevant historical calculation records and calculating historical information weights; fusing the historical information weights with the initial analysis matrix to generate an enhanced analysis matrix; S22. Analyze the enhanced analysis matrix to identify the primary performance target; decompose the primary performance target into a set of sub-targets, construct a target dependency graph, and obtain target structure data; calculate the resource requirement vector and target priority matrix for each sub-target to generate target resource data; analyze the computing feature requirements based on the primary performance target and historical computing records, and calculate the importance weight of the computing mode; based on the calculated importance weight, integrate the target structure data and target resource data into performance analysis result data; S23. Based on the performance analysis result data, calculate the reasoning performance analysis necessity score, evaluate the analysis risk value, and obtain analysis evaluation data; based on the analysis evaluation data, determine the analysis timing and generate an analysis priority list; based on the analysis priority list, formulate an adjustment strategy and form analysis strategy data; integrate the analysis evaluation data and the analysis strategy data to generate an analysis path map; based on the analysis path map, calculate the confidence score and ultimately form analysis decision data; S24. Perform internal consistency verification on the analysis and decision-making data to generate consistency verification data; based on the consistency verification data, verify resource availability, check computing constraints and timing restrictions, and obtain feasibility assessment data; based on the consistency verification data and feasibility assessment data, calculate the verification score, mark risk points, and generate optimization suggestions; based on the optimization suggestions, form verified decision data.

[0011] Preferably, step S3 is further: S31. Based on the pre-stored original information of the NPU computing feature library and the verified decision data, extract the functional feature vector, attribute index vector, and resource requirement vector of each computing mode to generate static feature data; calculate the historical effect matrix, average inference time vector, and resource consumption distribution of the computing mode to form dynamic feature data; construct a computing dependency graph, calculate the compatibility matrix of the computing mode and the computing combination effect tensor to obtain associated feature data; integrate the static feature data, dynamic feature data, and associated feature data into an enhanced computing feature space; S32. Calculate the functional applicability based on the enhanced computing feature space; perform attribute constraint filtering based on the functional applicability to generate an initial candidate computing set; obtain the context features of the current NPU and calculate the context relevance score; adjust the candidate computing weights based on the context relevance score, and reorder the initial candidate computing set to obtain an optimized candidate computing set; construct a computing feasible combination set and calculate the combination synergy score; select the optimal combination scheme from the optimized candidate computing set based on the combination synergy score to form an optimized computing selection scheme; S33. Based on the optimized computation selection scheme, construct an inference dependency graph, calculate the critical path, and generate a parallel inference scheme; based on the parallel inference scheme, form inference sequence data; based on the inference sequence data, construct a resource allocation matrix and optimize the inference timing; based on the optimized inference timing, construct a cache strategy and obtain resource optimization data; based on the resource optimization data, construct a failure handling strategy, alternative schemes, and monitoring point set and generate fault tolerance mechanism data; integrate the inference sequence data, resource optimization data, and fault tolerance mechanism data into the optimized inference scheme; S34. Perform an online status check on the computing model in the optimized inference scheme to verify resource adequacy, test interface response, and generate availability verification data; based on the availability verification data, perform permission check, risk assessment, and compliance verification to form security assessment data; based on the security assessment data, estimate response time, predict resource consumption, calculate success probability, and obtain performance prediction data; integrate the availability verification data, security assessment data, and performance prediction data into pre-inspection report data.

[0012] Preferably, step S12 is further: S121. Read the computing unit segments in the standardized computing data, calculate the parameter value, computing value, and load value of each computing unit segment, and generate computing statistical data; based on the computing statistical data, use an analysis tool to segment the computing unit segments, obtain computing statistics, and generate computing feature data; combine the computing statistical data and the computing feature data to construct complete basic feature data; S122. Obtain computation information from the basic feature data. Based on the computation information, calculate the architectural correlation strength of each computation using the bidirectional correlation network of the model to generate computation-level task correlation data. Based on the computation-level task correlation data, construct a task similarity matrix, calculate key task units, and form task unit data. Map the task unit data to a predefined demand space, calculate the demand probability distribution, and obtain task feature data. S123. Obtain the operation sequence in the pre-stored NPU historical records, construct a timing feature vector, and generate historical calculation data; obtain and analyze the current architecture state, including architecture usage time, operation rounds, and calculation continuity, to form architecture state data; perform feature fusion on the task feature data, historical calculation data, and architecture state data, and output the final performance feature vector.

[0013] Preferably, step S22 is further: S221. Read the performance description information in the enhanced analysis matrix, construct an architecture dependency tree using a model based on the performance description information, extract core analysis nodes, and generate analysis sequence data; identify key performance targets based on the analysis sequence data, calculate the strength of logical relationships between the targets, and form target association data; combine the analysis sequence data and the target association data to output performance target data; S222. Based on the performance target data, pattern matching is performed using a predefined model target decomposition template library to identify decomposable sub-target units and generate an initial sub-target set; the execution conditions and completion criteria of each sub-target in the initial sub-target set are analyzed, a sub-target constraint relationship diagram is constructed, and target constraint data is obtained; based on the target constraint data, the initial sub-target set is optimized and reorganized to output sub-target sequence data; S223. Based on the sub-goal sequence data, extract the input-output dependency relationship of each sub-goal, construct a data flow graph, and generate data dependency data; based on the data dependency data, analyze the execution order constraints, identify parallel execution opportunities, construct a target execution network, and generate execution dependency data; integrate the data dependency data and execution dependency data to construct a complete target dependency graph; S224. Obtain historical execution records, and based on the sub-goal sequence data and historical execution records, calculate the processing complexity and resource consumption characteristics of each sub-goal to generate resource characteristic data; based on the resource characteristic data, analyze the time sensitivity and priority factors of the sub-goals, construct the target scheduling weight matrix, and form scheduling characteristic data; combine the resource characteristic data, scheduling characteristic data and the target dependency graph to output the final performance analysis result data.

[0014] Preferably, step S23 is further: S231. Read resource requirement information from the performance analysis result data, calculate resource utilization thresholds based on the resource requirement information and pre-stored historical analysis records, and generate resource evaluation data; analyze target completion time requirements based on the resource evaluation data, calculate the time pressure coefficient based on the current system load status, and generate time evaluation data; calculate a reasoning performance analysis necessity score matrix based on the resource evaluation data and the time evaluation data, and output the analysis necessity data; S232. Based on the analysis necessity data, extract characteristic patterns of historical analysis failure cases, construct risk characteristic vectors, and generate risk pattern data; based on the risk pattern data, analyze the similarity between the current target and pre-stored historical high-risk scenarios, calculate multi-dimensional risk coefficients, and generate risk assessment data; based on the risk pattern data and risk assessment data, construct a risk-benefit assessment matrix and output the analysis risk data; S233. Based on the analysis necessity data and analysis risk data, construct a reasoning performance analysis time series network, calculate the optimal analysis time window, and generate analysis time series data; based on the analysis time series data, analyze the priority dependency relationship between the calculation modes, establish an analysis priority queue, and generate priority data; combine the analysis time series data and priority data to construct an analysis execution plan and output analysis strategy data; S234. Based on the analysis strategy data and pre-stored historical failure processing records, a fault processing decision tree is constructed to generate fault recovery data; based on the fault recovery data, a multi-level adjustment plan is constructed, including alternative calculation chains and simplified strategies, to form adjustment strategy data; the analysis strategy data, fault recovery data, and adjustment strategy data are integrated to calculate the strategy reliability score, construct a complete analysis path map, and finally output the analysis decision data.

[0015] Preferably, step S32 is further: S321. Based on the functional feature vectors and decision requirements in the enhanced computing feature space, the model is used to calculate the functional applicability score of each computing mode to generate functional applicability data. Based on the functional applicability data, the attribute indicator vector is used to perform constraint filtering to select a set of computing modes that meet the attribute requirements to form attribute filtering data. Combining the functional applicability data and the attribute filtering data, a preliminary selection computing list and its scoring matrix are constructed to output initial candidate data. S322. Obtain context features of the current NPU, including timing windows, resource status, and target priority, and generate context feature data. Based on the context feature data, analyze the computing usage effects in historical similar scenarios, construct a scenario correlation matrix, and form scenario matching data. Calculate a context adjustment coefficient based on the context feature data and the scenario matching data, and output context score data. S323. Based on the initial candidate data and contextual scoring data, a dynamic weighting algorithm is used to adjust the calculation score to generate adjusted weighting data. The calculation credibility score is updated based on the historical success rate and stability index of the calculation model to generate credibility data. Based on the adjusted weighting data and credibility data, the preliminary calculation list is reordered to output optimized candidate data. S324. Based on the computational combination feature tensors in the optimization candidate data, a feasible computational combination scheme set is constructed to generate combination scheme data; based on the combination scheme data, synergy effect scores of different combination schemes are calculated, including functional complementarity and attribute gain, to form synergy evaluation data; based on the synergy evaluation data, the complexity and risk factors of the combination scheme data are analyzed, a comprehensive evaluation matrix is ​​constructed, and scheme evaluation data is obtained; S325. Based on the combination scheme data, collaborative evaluation data and scheme evaluation data, a multi-objective optimization algorithm is used to calculate the comprehensive score of each combination scheme and generate optimized scoring data; based on the optimized scoring data, the optimal combination scheme is selected, and a detailed calculation reasoning sequence is constructed to form reasoning sequence data; the optimized scoring data and reasoning sequence data are integrated, and finally the optimized calculation selection scheme is output.

[0016] Preferably, the method further comprises: S4. Collect the execution data flow and historical monitoring data during the neural network reasoning process to generate a comprehensive monitoring data packet; based on the comprehensive monitoring data packet, perform anomaly detection and early warning analysis, and output the anomaly analysis result data; based on the anomaly analysis result data and the comprehensive monitoring data packet, generate a dynamic optimization strategy to form an optimization strategy set data; based on the optimization strategy set data and the pre-stored historical optimization effect data, perform adaptive learning, and finally output the optimization update data packet.

[0017] Preferably, step S4 is further: S41. Read the real-time computing parameters in the execution data stream, extract the timestamp, computing load, and resource usage information, and generate real-time monitoring data; read the abnormal pattern library in the historical monitoring data, build an abnormal feature dictionary, and form historical monitoring data; perform feature fusion on the real-time monitoring data and the historical monitoring data, and output a comprehensive monitoring data package; S42. Based on the comprehensive monitoring data packet, using a pre-trained anomaly detection model, calculate the data deviation and generate deviation data; based on the deviation data, identify the anomaly type and severity level, construct an anomaly classification tree, and generate anomaly classification data; integrate the deviation data and the anomaly classification data, and output anomaly analysis result data; S43. Based on the anomaly analysis result data, extract the anomaly impact scope and duration to generate impact assessment data; based on the impact assessment data, call a pre-stored policy template library, match corresponding adjustment rules, and generate an initial optimization policy; based on the initial optimization policy, analyze the policy feasibility and adjustment cost, construct a policy evaluation matrix, and form optimization policy set data; S44. Based on the optimization strategy set data, extract the historical application effect records of each strategy, calculate the strategy effectiveness score, and generate effect evaluation data; based on the effect evaluation data, use the adaptive learning algorithm to update the strategy weights, build the learning model parameters, and form learning update data; integrate the effect evaluation data and the learning update data, and finally output the optimization update data package.

[0018] Compared with the prior art, the present invention has the following beneficial effects: During data processing, by standardizing the raw data from the NPU computing architecture, converting the computing unit configuration into a unified encoding format, removing outliers and redundant information, and performing data length standardization, data quality can be effectively improved, laying a solid foundation for subsequent analysis. This standardized processing ensures data consistency and reliability, avoiding analytical errors caused by inconsistent data formats or noise.

[0019] In terms of feature extraction, this method is based on standardized computing data, comprehensively extracting computing unit features, computing inference task features, and architecture load features to generate performance feature vectors. By calculating the unit parameter values ​​of the computing mode, extracting key computing feature sets, identifying inference task types, and other operations, it is possible to deeply explore the potential features in the data, so that the performance feature vectors more accurately reflect the actual operation of the NPU computing architecture. This helps to more accurately identify inference performance requirements and provide strong support for subsequent performance analysis.

[0020] During the performance analysis process, preliminary analysis results and standardized calculation data are integrated to generate an enhanced analysis matrix, which then conducts multi-dimensional performance analysis. This multi-dimensional analysis approach comprehensively evaluates neural network inference performance from different perspectives, avoiding the limitations of single-dimensional analysis. Simultaneously, the inference performance analysis necessity score and architectural impact value are calculated to generate analytical decision data, which is then validated through multiple stages. This improves the scientific nature and reliability of analytical decisions, providing valuable guidance for subsequent optimization efforts.

[0021] During the optimization phase, an enhanced computing feature space is constructed based on the pre-stored raw information of the NPU computing feature library and verified decision data. Computation matching and parameter optimization are performed to generate an optimized computing selection plan and optimize the inference path. This optimization method fully utilizes the characteristics of the NPU computing architecture and historical data to find the optimal computing mode and inference path, improving the efficiency and performance of neural network inference. In addition, multi-dimensional pre-inspection before analysis can identify potential problems in advance and ensure the feasibility and stability of the optimized inference plan.

[0022] In terms of real-time monitoring and dynamic optimization, the system collects execution data streams and historical monitoring data from the neural network inference process, performs anomaly detection and early warning analysis, generates dynamic optimization strategies, and performs adaptive learning. This real-time monitoring and dynamic optimization mechanism promptly detects anomalies during neural network inference and adjusts the optimization strategy accordingly, improving the stability and reliability of neural network inference. Adaptive learning also enables continuous optimization of strategies, enhancing system performance and adaptability. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 This is a diagram showing the working principle of the neural network reasoning performance analysis method for NPU computing architecture described in the present invention; Figure 2 Design diagram for standardization processing and feature extraction. DETAILED DESCRIPTION

[0024] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0025] See also Figure 1 - Figure 2 The present invention relates to a neural network reasoning performance analysis method for NPU computing architecture, and the specific implementation steps are as follows: S1. Obtain the original data of the NPU computing architecture and standardize it to obtain standardized computing data. Based on the standardized computing data, extract the computing unit characteristics, computing reasoning task characteristics, and architecture load characteristics to generate a performance feature vector. Based on the performance feature vector, identify the reasoning performance requirements and determine the analysis dimensions to obtain preliminary analysis result data. S2. Integrate the preliminary analysis result data and the standardized calculation data to generate an enhanced analysis matrix; perform multi-dimensional performance analysis based on the enhanced analysis matrix to obtain performance analysis result data; analyze the performance analysis result data to calculate the reasoning performance analysis necessity score and architectural impact value to form analysis decision data; perform multiple verifications on the analysis decision data to generate verified decision data; S3. Based on the pre-stored original information of the NPU computing feature library and the verified decision data, an enhanced computing feature space is constructed; based on the enhanced computing feature space, computing matching and parameter optimization are performed to generate an optimized computing selection plan; the reasoning path of the optimized computing selection plan is optimized to form an optimized reasoning plan; based on the optimized reasoning plan, a multi-dimensional pre-inspection is performed before analysis, and finally the pre-inspection report data is output.

[0026] Example 1: This embodiment refines step S1: See also Figure 2 In step S11, when obtaining the original data of the NPU computing architecture, specifically obtain the original data including the computing unit configuration, timestamp, core identifier and task identifier. Among them, the computing unit configuration covers parameters such as the type, quantity, and operating frequency of the computing unit, such as the specific models of different types of convolution computing units and fully connected computing units and the corresponding operating frequency range; the timestamp is used to accurately record the time sequence of each computing operation, accurate to the nanosecond level, to reflect the timing characteristics of the computing task; the core identifier corresponds to the different computing cores in the NPU, which can clearly define the physical core assigned to each computing operation; the task identifier is used to distinguish different inference tasks, such as image classification tasks, speech recognition tasks, etc. Each task identifier contains the type code and unique serial number of the task. The computing unit configuration in the original data is converted into a unified encoding format, and the JSON format is used for standardized encoding. For example, the parameters such as the type, quantity, and operating frequency of the computing unit are converted into key-value pairs in the JSON object to ensure the consistency and parsability of the data format. At the same time, a data cleaning algorithm removes outliers and redundant information. Outlier detection uses a statistically distributed approach, setting a reasonable range for parameters such as operating frequency. Values ​​outside this range are considered outliers and removed. Redundant information, primarily invalid timestamps of duplicate records and duplicate identifiers for the same task, is processed using a hash deduplication algorithm. After processing, the input data is length-normalized, determining a uniform standard length based on the maximum length of the data. Insufficient input data is padded with zeros. Data exceeding the standard length is truncated at the end to generate standardized calculated data, ensuring consistency in subsequent processing.

[0027] In step S12, based on the standardized calculation data, the parameter values, calculation values, and load values ​​of each calculation unit segment are first calculated. Taking the matrix multiplication unit as an example, the parameter values ​​include the number of rows, columns, and channels of the input matrix; the calculation value is the number of multiplication operations and addition operations actually completed; the load value is determined by the ratio of the amount of operations per unit time to the maximum computing power of the unit, and calculation statistics are generated. The calculation unit segment is segmented using an analysis tool. A sliding time window method is used, and the window size is set to 100 clock cycles. The calculation unit segment within each window is segmented to obtain calculation statistical features, such as the average number of operations in each window and the fluctuation range of the load rate, to generate calculation feature data. Combining the calculation statistics and calculation feature data, complete basic feature data is constructed, which contains detailed parameters, calculation process characteristics, and load status of each calculation unit segment.

[0028] Computational information from the basic feature data is obtained, and the architectural correlation strength of each computation is calculated using a bidirectional correlation network model. This bidirectional correlation network model consists of two processes: forward propagation and backward propagation. Forward propagation starts from the computational operation and analyzes its requirements for architectural components such as storage units and control units. Backward propagation starts from the architectural components and evaluates their support for the computational operation. Through the interaction between the two processes, specific values ​​such as the correlation strength between convolution operations and storage units are obtained, generating computation-level task correlation data. Based on the computation-level task correlation data, a task similarity matrix is ​​constructed, and the cosine similarity algorithm is used to calculate the feature similarity between different tasks. Each element in the matrix represents the degree of similarity between the corresponding two tasks. Key task units are identified by setting similarity thresholds, generating task unit data. The task unit data is mapped to a predefined requirement space consisting of three dimensions: performance, power consumption, and accuracy. Each dimension is divided into multiple levels. A pretrained neural network model is used for nonlinear mapping to calculate the probability distribution of requirements and generate task feature data.

[0029] At the same time, the operation sequence in the NPU history is obtained. The operation sequence contains all the computing operations performed by the NPU in the past period of time and their order. A time series feature vector is constructed, and the operation type is converted into the corresponding vector representation. It is arranged in chronological order to generate historical computing data. The current architecture status is analyzed, including the architecture usage time, that is, the cumulative working time of the NPU from startup to the current moment; the operation rounds, which record the number of complete computing cycles completed by the NPU; and the computing coherence, which is measured by calculating the duration of continuous operations and the number of interruptions to form architecture status data. The task feature data, historical computing data, and architecture status data are fused and a weighted summation method is used. The weights are dynamically adjusted according to the real-time nature and importance of the data to output the final performance feature vector, which comprehensively reflects the characteristics of the current computing task, historical execution status, and architecture status.

[0030] In step S13, based on the performance feature vector, the preset feature-requirement mapping rules are used to identify the type of inference performance requirement. The preset feature-requirement mapping rules are constructed using a decision tree model. The nodes of the decision tree are divided based on the various dimensions of the performance feature vector, and the leaf nodes correspond to different requirement types, including computing speed, storage bandwidth, power consumption limit, etc. When calculating the requirement priority score, weights are assigned based on the real-time requirements and importance of the task. Tasks with high real-time requirements are given higher weights. The importance is measured by the criticality of the task in the entire inference process to generate the requirement feature data.

[0031] Analyze the demand feature data and estimate the computing resource requirements and time window requirements. Computing resource requirements include the number of computing cores and storage capacity required. For example, based on the computing speed requirement, it is estimated that 8 convolutional computing cores and 2GB of storage capacity are required. The time window requirement is calculated based on the task deadline and current time. For example, if a task needs to be completed within 100 milliseconds, the resource requirement data is obtained. Based on the demand feature data and resource requirement data, the dependencies between requirements are identified. By constructing a directed graph, nodes represent requirements and edges represent dependencies between requirements. For example, storage requirements must be met before computing requirements. A dependency graph is constructed. Finally, based on the inference performance requirement type, priority score, resource requirement data, and dependency graph, preliminary analysis result data is generated. This data is presented in a structured manner and contains all identified requirements and their interrelationships, resource requirements, time requirements, and other information, providing a foundation for subsequent analysis.

[0032] Example 2: This embodiment refines step S2: In step S21, when converting the preliminary analysis result data into a feature matrix, it is necessary to structure the information such as the demand type, priority score, resource demand value, etc. in the preliminary analysis result. For example, a two-dimensional matrix is ​​constructed with behavioral demand items and columns as feature dimensions, and each cell corresponds to the characteristic value of a specific demand, such as the priority score corresponding to the computing speed requirement, the required number of computing cores, etc. When converting the standardized computing data into a vector representation, it is necessary to integrate the computing unit parameters, task characteristics, architecture status and other data, and arrange them in a preset order to form a one-dimensional vector. The vector elements contain specific values ​​such as the computing unit working frequency, task type code, and architecture usage time.

[0033] After combining the feature matrix and vector representation to form the initial analysis matrix, relevant historical computation records are extracted. These records contain performance data, resource usage, and analysis results of similar past tasks. The historical information weights are calculated using the TF-IDF algorithm, which assigns weights based on the frequency and importance of each feature in the historical records. This algorithm prioritizes frequently occurring features that have a greater impact on performance analysis, giving them higher weights. When integrating the historical information weights with the initial analysis matrix, a weighted overlay approach is employed. This enhanced analysis matrix incorporates both the real-time features of the current data and reference information from historical experience, improving the matrix's information richness and analytical reliability.

[0034] In step S22, the performance description information in the enhanced analysis matrix is ​​read and an architectural dependency tree is constructed using a graph neural network model. This model uses the features in the matrix as nodes and the relationships between features as edges. Through information transfer through a multi-layer neural network, it highlights core analysis nodes that have a significant impact on performance, such as high-load computing units or critical task features. After extracting the core analysis nodes, they are sorted by importance to generate analysis sequence data, which reflects the critical path of performance analysis.

[0035] When identifying key performance objectives based on sequence analysis data, business requirements and architectural constraints are combined to identify core objectives such as "improving inference speed" and "reducing power consumption." Objective-related data is generated by calculating the strength of logical relationships between objectives, such as the degree of conflict between speed and power consumption (using a scale of 0 to 100 to represent conflict strength). The sequence analysis data and objective-related data are combined to output performance objective data that includes objective priorities and relationships.

[0036] Based on performance target data, pattern matching is performed using a predefined model goal decomposition template library. The template library includes various templates, such as hierarchical decomposition and functional decomposition. For example, the goal of "improving inference speed" can be decomposed into sub-goal units such as "optimizing computing unit efficiency" and "reducing data access latency" to generate an initial set of sub-goals. The execution conditions and completion criteria of each sub-goal in the initial set are analyzed, and a sub-goal constraint relationship diagram is constructed. Directed edges are used to represent the sequence and dependencies between sub-goals. The initial set of sub-goals is optimized and reorganized, and sub-goal sequence data is output in a logically ordered sequence.

[0037] Based on the sub-goal sequence data, the input-output dependencies of each sub-goal are extracted, and a data flow graph is constructed to clarify the data transfer paths and processing flows between sub-goals, generating data dependency data. Execution order constraints are analyzed to identify sub-goals that can be executed in parallel. A target execution network is constructed, with nodes representing sub-goals and edges representing execution order or parallelism, generating execution dependency data. The data dependency and execution dependency data are integrated to construct a complete target dependency graph, which comprehensively reflects the logical relationships and execution paths between sub-goals.

[0038] Obtain historical execution records and, based on the subgoal sequence data and historical records, calculate the processing complexity and resource consumption characteristics of each subgoal to generate resource characteristic data. Analyze the time sensitivity and priority factors of the subgoals and construct a target scheduling weight matrix. The matrix elements represent the weight of each subgoal in the scheduling process, forming scheduling characteristic data. Combine the resource characteristic data and scheduling characteristic data with the target dependency graph to output the final performance analysis results, including resource requirements, execution order, and scheduling priority.

[0039] In step S23, the resource demand information in the performance analysis result data is read, and the resource utilization threshold is calculated using a statistical method based on the pre-stored historical analysis records. For example, the CPU utilization data in the historical records is sorted, and the 90th percentile is taken as the utilization threshold to generate resource evaluation data. The target completion time requirements are analyzed, and the time pressure coefficient is calculated in combination with the current load status of the system (such as the current CPU utilization and the remaining memory capacity). This coefficient is obtained by adjusting the system load through the ratio of the target deadline to the current remaining time to form the time evaluation data. Based on the resource evaluation data and the time evaluation data, the reasoning performance analysis necessity score matrix is ​​calculated. Each element in the matrix corresponds to the necessity score under different resource and time combinations, and the analysis necessity data is output.

[0040] Based on the analysis necessity data, characteristic patterns of historical analysis failure cases are extracted. For example, in cases of failure due to insufficient resources, the frequency of memory capacity falling below a threshold and the number of insufficient computing cores are observed. A risk feature vector is constructed, which contains multiple failure-related characteristic dimensions and generates risk pattern data. The similarity between the current target and pre-stored historical high-risk scenarios is analyzed. By calculating the Euclidean distance between the feature vectors, this is converted into a multi-dimensional risk coefficient to form risk assessment data. Based on the risk pattern data and risk assessment data, a risk-benefit assessment matrix is ​​constructed, with rows and columns representing different risk levels and expected returns. The analysis risk data is then output.

[0041] Based on analysis necessity data and analysis risk data, a time-series network for reasoning performance analysis is constructed. This network is time-based, with nodes representing analysis tasks and edges representing temporal dependencies between tasks. A time-series planning algorithm is used to calculate the optimal analysis time window, taking into account factors such as system load troughs and task deadlines, to generate analysis time-series data. Priority dependencies between analysis computing modes are analyzed, and an analysis priority queue is established, sorting by priority to generate priority data. Combining the analysis time-series data with the priority data, an analysis execution plan is constructed that includes time scheduling and task sequencing, and the analysis strategy data is output.

[0042] Based on analysis policy data and pre-stored historical failure handling records, a fault handling decision tree is constructed. The root node of the decision tree is the fault phenomenon, the branches are different handling paths, and the leaf nodes are specific solutions. Fault recovery data is generated. A multi-level adjustment plan is constructed, including alternative calculation chains and simplified strategies, to form adjustment strategy data. The analysis policy data, fault recovery data, and adjustment policy data are integrated. The reliability score of each strategy is calculated by evaluating the reliability indicators of each strategy. A complete analysis path diagram is constructed, and the final output is analysis decision data including strategy, fault handling, and adjustment plan.

[0043] In step S24, the analysis and decision data is internally verified for consistency, checking for logical contradictions between objectives and strategies, such as whether resource requirements exceed system capabilities or whether there are scheduling conflicts. This generates consistency verification data. Based on this consistency verification data, resource availability is verified, such as whether the current number of computing cores meets analysis requirements and whether memory capacity is sufficient. Computational constraints and timing restrictions are also checked, such as whether the time window for task execution complies with system scheduling rules, to generate feasibility assessment data.

[0044] Based on consistency verification data and feasibility assessment data, a verification score is calculated, which comprehensively considers the degree of consistency and feasibility. Risk points, such as insufficient resources and tight schedules, are identified. Optimization suggestions are generated for each risk point, such as requesting additional computing resources or adjusting the order of task execution. This generates post-verification decision data. This data undergoes multiple verification and optimization processes, providing a reliable basis for subsequent computational feature space construction and inference solution optimization.

[0045] Example 3: This embodiment refines step S3: In step S31, based on the pre-stored original information of the NPU computing feature library and the verified decision data, the functional feature vector, attribute index vector and resource requirement vector of each computing mode are extracted. The functional feature vector covers the input and output dimensions, operation type, etc. of the computing mode, such as the input feature map size, convolution kernel size, number of output channels and other parameters of the convolution computing mode; the attribute index vector includes computing accuracy, power consumption parameters, such as floating-point operation accuracy and power consumption per unit time; the resource requirement vector includes the required memory capacity, number of computing cores, etc., such as the minimum memory space required to complete the computing mode and the recommended number of computing cores to be allocated, to generate static feature data.

[0046] Calculate the historical performance matrix, average inference time vector, and resource consumption distribution of the computing mode. The historical performance matrix records metrics such as inference accuracy under different datasets, such as classification accuracy on the ImageNet and CIFAR datasets. The average inference time vector is the average time taken to run the computing mode multiple times, accurate to the millisecond level. The resource consumption distribution is formed by statistically analyzing the probability distribution of data such as memory usage and computing core utilization, such as the frequency of memory usage in different intervals, to form dynamic feature data.

[0047] A computational dependency graph is constructed, with nodes representing computational modes and edges representing sequential dependencies between modes. A compatibility matrix is ​​calculated, assessing the degree of resource conflict when two computational modes are running simultaneously to obtain a compatibility score, ranging from 0 to 100, with higher values ​​indicating better compatibility. A combination effect tensor is calculated to quantify the synergistic effect of combining three or more computational modes, generating associated feature data. Static, dynamic, and associated feature data are integrated to form an enhanced computational feature space that encompasses the computational mode functions, performance, resource requirements, and interrelationships.

[0048] In step S32, based on the functional feature vectors and decision requirements in the enhanced computing feature space, the functional applicability score of each computing mode is calculated through the support vector machine model. The model takes the functional feature vector as input and the decision requirements as constraints, and outputs a score between 0 and 1. The higher the score, the more it meets the functional requirements, and generates functional applicability data. Based on the functional applicability data, the attribute indicator vector is used to filter the constraints, set the thresholds of attributes such as power consumption and accuracy, and filter out the calculation sets that meet the attribute requirements. Combined with the functional applicability data and attribute filtering data, a preliminary calculation list and its scoring matrix are constructed. The scoring matrix contains information such as functional applicability scores and attribute compliance, and outputs initial candidate data.

[0049] Obtain the contextual features of the current NPU, including the timing window, resource status, and target priority. The timing window records the time limit of the current task, such as the inference must be completed within a specified time; the resource status includes real-time data such as the current available memory and computing core load; the target priority is divided into levels according to the importance of the task to generate contextual feature data. Based on the contextual feature data, analyze the computing usage effects in similar historical scenarios, construct a scenario correlation matrix, and the matrix elements represent the similarity between the current scenario and the historical scenario to form scenario matching data. Through similarity calculation and weight adjustment, calculate the context adjustment coefficient and output the context score data, which reflects the adaptability of the current context to each computing mode.

[0050] Based on the initial candidate data and contextual scoring data, a dynamic weighting algorithm is used to adjust the calculation score. Weights are dynamically assigned based on factors such as the context's urgency and resource availability. For example, when resources are limited, the weight of the low-power computing mode is increased to generate adjusted weight data. The historical success rate and stability indicators of the computing mode are obtained. The historical success rate is the percentage of successful runs of the mode in the past. The stability indicator is measured using parameters such as the variance of the running time. The calculation credibility score is updated to generate credibility data. Based on the adjusted weighting data and credibility data, the preliminary calculation list is reordered, prioritizing computing modes with high scores and high credibility, and outputting optimized candidate data.

[0051] Based on the computational combination feature tensors in the optimization candidate data, a feasible computational combination solution set is constructed. The combination solution set includes multiple possible combinations of computational modes to generate combination solution data. The synergy effect scores of different combination solutions are calculated, including functional complementarity and attribute gains, such as the acceleration effect when combining convolutional and pooling calculations. The scores are obtained by evaluating the improvement in computational efficiency before and after the combination to form synergy evaluation data. The complexity and risk factors of the combination solution data are analyzed. Complexity is measured by the number of steps in the computational process and the complexity of data dependencies. Risk factors include the possibility of resource conflicts and the stability of the combination solution. A comprehensive evaluation matrix is ​​constructed to obtain solution evaluation data.

[0052] Based on the combination scheme data, collaborative evaluation data, and scheme evaluation data, a multi-objective optimization algorithm is used to comprehensively consider multiple objectives such as functional applicability, synergy, complexity, and risk. A comprehensive score is calculated for each combination scheme, generating optimized scoring data. The optimal combination scheme is selected based on the comprehensive score, and a detailed computational reasoning sequence is constructed to clarify the execution order and data flow of each computational mode, generating reasoning sequence data. The optimized scoring data and reasoning sequence data are integrated to output an optimized computational selection scheme that includes the optimal combination scheme and a detailed execution sequence.

[0053] In step S33, based on the optimized computational selection scheme, an inference dependency graph is constructed, with nodes representing computational steps and edges representing dependencies between steps, such as data input / output dependencies and timing dependencies. A graph theory algorithm is used to calculate the critical path (the longest path in the entire inference process). A parallel inference scheme is generated, executing independent computational steps in parallel to shorten the overall execution time. Based on the parallel inference scheme, inference sequence data is generated, which contains information such as the execution order of each computational step and the parallel grouping.

[0054] A resource allocation matrix is ​​constructed to allocate computing cores, memory, and other resources to each computational step. Resource usage periods and allocation amounts are clearly defined, and inference timing is optimized. Scheduling algorithms are used to determine the start and end times of each step, ensuring resource utilization while meeting time constraints. A caching strategy is constructed based on the optimized inference timing, analyzing data access patterns to preload frequently used data into the cache, reducing data access latency and generating resource-optimized data.

[0055] Build a failure handling strategy, alternative solutions, and monitoring point sets. The failure handling strategy defines the recovery process when a computational step fails, such as a retry mechanism and error logging. The alternative solution is an alternative computational process when the primary computational path fails. The monitoring point set includes performance monitoring metrics for key computational steps and generates fault-tolerance mechanism data. Inference sequence data, resource optimization data, and fault-tolerance mechanism data are integrated to form an optimized inference solution that includes execution flow, resource allocation, and fault-tolerance measures.

[0056] In step S34, an online status check is performed on the computing model in the optimized inference scheme to verify resource adequacy, such as checking whether the current available memory meets the requirements of the computing model and whether the computing core load is within a reasonable range; test interface responses, such as the delay time of the input and output interfaces, data transmission rate, etc., to generate availability verification data to ensure that the computing model can operate normally in the current environment.

[0057] Based on availability verification data, we conduct permission checks, risk assessments, and compliance verification. Permission checks confirm access to required resources; risk assessments analyze the probability and impact of potential failures during the computational model's operation; and compliance verification ensures that the computational process complies with relevant regulations for data security and privacy protection, generating security assessment data.

[0058] Response time is estimated by using historical execution data and current system load to predict the time required to complete inference. Resource consumption is predicted, estimating memory usage, power consumption, and other resource usage. Success probability is calculated, assessing the likelihood of inference success based on historical success rates and current environmental factors, generating performance prediction data. Availability verification data, security assessment data, and performance prediction data are integrated to generate a pre-verification report containing availability, security, and performance predictions, providing a comprehensive evaluation basis for the implementation of the inference solution.

[0059] Example 4: This embodiment further refines steps S12, S22, S23, and S32: In step S12, when reading the computing unit fragments in the standardized computing data, the sliding window technology is used to segment the data, and the window size is set to 200 clock cycles to ensure that each fragment contains a complete sequence of basic computing operations such as convolution or pooling. Taking the convolution layer calculation in the image classification task as an example, the fragments in each window contain data such as the size of the input feature map, the convolution kernel parameters, and the calculation process of the output feature map. When calculating the parameter value, the real-time data of the hardware performance counter, such as the cache hit rate and the number of instruction cycles, is combined. For example, when calculating the load value of the convolution unit, not only the number of operations is considered, but also the L1 cache hit rate (obtained by real-time monitoring) is introduced as a correction factor to avoid load calculation deviations caused by cache failures and improve the accuracy of parameter calculations.

[0060] When constructing the task similarity matrix, the cosine similarity algorithm is used to calculate the feature similarity between tasks. Taking image classification tasks and target detection tasks as examples, the feature vectors such as the usage pattern of computing units (such as the ratio of convolutional layers to fully connected layers) and data throughput are extracted from the two tasks. The similarity value is calculated using the cosine similarity formula. The threshold is set to 0.6, and tasks with similarity above this threshold are classified into the same category to ensure the recognition accuracy of key task units. When mapping task unit data to the demand space, a pre-trained neural network model is used for nonlinear mapping. For example, the high real-time requirements of target detection tasks are mapped to the high-priority area of ​​the performance dimension, so that the demand probability distribution is more in line with actual business needs.

[0061] In step S22, when constructing the architecture dependency tree, the attention mechanism model is used to highlight the influence weights of key analysis nodes. Taking a certain NPU architecture as an example, when analyzing the performance bottleneck of matrix multiplication operations, the model will automatically give higher attention weights to nodes related to storage access, because the performance of matrix multiplication is often limited by memory bandwidth, ensuring the accurate extraction of core nodes. In the process of goal decomposition, a rule base constructed by expert knowledge is introduced. For example, when the goal of "improving inference speed" is decomposed into "optimizing computing unit parallelism" and "reducing data movement overhead", the rule base will verify whether the decomposed sub-goals cover all possible optimization paths to avoid logical contradictions in the decomposed sub-goals.

[0062] When building the target execution network, the Petri net model is incorporated to formally describe parallel execution opportunities. For example, in an image processing task, if feature extraction and classifier computation lack data dependencies, the Petri net model explicitly indicates that these two sub-goals can be executed in parallel, ensuring accurate representation of execution dependencies. When calculating resource feature data, real-time monitoring data and historical statistical data are integrated, and a Kalman filter algorithm is used to dynamically update resource consumption features. For example, the memory usage prediction of the convolutional layer is updated every 10 milliseconds, improving data timeliness.

[0063] In step S23, when calculating the resource utilization threshold, a quantile regression method is used to address outliers in historical data. Taking CPU utilization as an example, the 95th percentile of the historical data is selected as the threshold to avoid interference from individual peak data on the threshold setting, making the threshold more robust. When constructing the risk feature vector, multi-source risk data is integrated, including hardware failure history (such as computing core overheating records) and software error logs (such as memory access out-of-bounds errors). Principal component analysis is used to reduce feature dimensionality, reducing the original 20-dimensional features to 8-dimensional key features, thereby improving risk assessment efficiency.

[0064] When building analysis and execution plans, a genetic algorithm is introduced to jointly optimize time windows and resource allocation. For example, in a certain inference task, the genetic algorithm uses iterative optimization to allocate more computing resources to perform high-power matrix operations during periods of low system load (such as 2 a.m.), ensuring the feasibility and optimality of the plan. When building a fault handling decision tree, a reinforcement learning model is used to dynamically adjust the decision path. For example, if out-of-memory faults occur repeatedly, the model automatically increases the priority of the "increase virtual memory" strategy, optimizing the decision logic based on historical fault handling results.

[0065] In step S32, when calculating the functional applicability score, an integrated learning model is used to combine the results of multiple classifiers. Taking the determination of whether a certain computing mode is suitable for image segmentation tasks as an example, the prediction results of random forest, support vector machine and neural network models are integrated, and the final score is determined through a voting mechanism to improve the reliability of the score. In the attribute constraint filtering process, fuzzy logic is introduced to deal with unclear constraints. For example, "higher precision" is converted into a fuzzy set, and the computing mode with an accuracy threshold between 0.9 and 0.95 is set as acceptable, so that the filtering results are more in line with actual needs.

[0066] When calculating context adjustment coefficients, the dynamic changes in the current system load are taken into account, and a real-time feedback mechanism is used to update the adjustment coefficients. For example, when the GPU load is detected to exceed 80%, the weight of computing modes that require a large amount of GPU resources is automatically reduced to ensure that the priority ranking of candidate computing sets adapts to real-time scenarios. When constructing the comprehensive evaluation matrix, the entropy weight method is introduced to objectively calculate the weights of each evaluation indicator. Taking the three indicators of functional applicability, synergy effect, and complexity as examples, the entropy weight method calculates their weights as 0.4, 0.35, and 0.25, respectively, to avoid the influence of subjective factors on solution evaluation.

[0067] Taking a medical image segmentation task handled by an NPU as an example, in the optimization candidate calculation set in step S32, the initial candidates included 3D convolution calculation mode A and hybrid convolution calculation mode B. Through contextual features (the current GPU has 2GB of remaining video memory, so the task has a high priority) and historical scene matching (mode B has a higher success rate under similar video memory conditions), the adjusted score of mode B increased from 0.7 to 0.85, ultimately selecting it as the core calculation mode of the optimal combination solution. The synergy effect of its combination with the pooling mode was evaluated to be 0.72, higher than other combination solutions, ensuring efficient completion of the segmentation task under limited video memory conditions.

[0068] Example 5: This embodiment adds step S4, which specifically includes collecting, analyzing, and optimizing the execution data flow and historical monitoring data during the neural network inference process to generate a dynamic optimization strategy. The following is a detailed description with reference to specific examples: In step S41, taking an NPU processing an object detection task in an autonomous driving scenario as an example, real-time computation parameters are read from the execution data stream, including the timestamp (accurate to the millisecond) of each frame's image inference, computational load, and resource usage information (such as GPU memory usage). For example, when a vehicle approaches an intersection, the NPU needs to process image data from multiple cars ahead. Real-time monitoring data records the timestamp of each frame's convolutional layer processing, the number of floating-point operations per layer, and the specific value of GPU memory usage (e.g., 1.2GB). Simultaneously, a library of anomaly patterns from historical monitoring data is read. This library stores anomaly features from similar scenarios in the past. For example, a memory overflow anomaly corresponds to the feature "memory usage exceeds 1.5GB for five consecutive frames and continues to increase." An anomaly feature dictionary is constructed to generate historical monitoring data. Feature fusion is performed on the real-time monitoring data with the historical monitoring data. For example, the memory usage data for the current frame is combined with the memory features from historical anomaly patterns. A comprehensive monitoring data package containing both real-time status and historical patterns is output, providing multi-dimensional data support for subsequent anomaly detection.

[0069] In step S42, a pre-trained anomaly detection model (such as the Isolation Forest Model) is used to calculate data deviation based on the comprehensive monitoring data packets. Taking video memory usage as an example, the model calculates the deviation of the current video memory usage value based on the distribution of historical normal data. If the video memory usage of a certain frame is 1.6GB, while the historical normal data has a mean of 1GB and a standard deviation of 0.2GB, the model calculates the degree of deviation from the normal distribution and generates deviation data. Based on the deviation data, the anomaly type and severity level are identified. For example, if the deviation exceeds three times the standard deviation, it is identified as a "video memory overflow risk" anomaly and the severity level is set to high. An anomaly classification tree is constructed (e.g., with a root node of "Resource Anomaly" and child nodes such as "Video Memory Anomaly" and "Compute Core Anomaly") to generate anomaly classification data. The deviation data and anomaly classification data are integrated to output anomaly analysis result data. For example, in the above scenario, the anomaly analysis result would clearly indicate that "the video memory usage during inference for the 15th frame deviates from the normal range, presents an overflow risk, and has a high severity level."

[0070] In step S43, based on the abnormality analysis result data, the impact range and duration of the abnormality are extracted. For example, the impact range of the "video memory overflow risk" abnormality is the image frame currently being processed and subsequent frames. If it is not handled in time, the entire inference task may be interrupted; the duration is estimated to be that if the video memory usage of the current frame continues to increase, it may trigger an overflow after 3 frames. After generating the impact assessment data, the pre-stored policy template library is called to match the corresponding adjustment rules. The policy template library has two preset policies for "video memory overflow risk": "dynamically adjust image resolution" and "enable paged video memory". According to the current scenario, the "dynamically adjust image resolution" policy is matched to generate the initial optimization policy. Analyze the feasibility of the policy and the cost of adjustment. For example, when adjusting the resolution to 720p, the inference accuracy may drop by 5%, but the video memory usage can be reduced by 30%. Construct a policy evaluation matrix to form the optimization policy set data, and finally determine to adopt the combined policy of "reducing the resolution to 720p and enabling asynchronous data loading".

[0071] In step S44, based on the optimization strategy set data, the historical application effect records of each strategy are extracted. For example, in the past 100 applications of the "dynamically adjust image resolution" strategy, the number of times it successfully alleviated the video memory pressure was 85 times, with an average accuracy loss of 4.2%. The strategy effectiveness score is calculated and the effect evaluation data is generated. An adaptive learning algorithm is used to update the strategy weights, and the priority of each strategy is adjusted according to the characteristics of the current scenario. For example, when it is detected that there are dense vehicles ahead, the weight of the "maintain resolution but enable video memory compression" strategy is automatically increased, the learning model parameters are constructed, and the learning update data is formed. The effect evaluation data and the learning update data are integrated to output the optimization update data packet, which will be stored in the historical optimization effect database for subsequent strategy optimization in similar scenarios.

[0072] Taking another example, when an NPU was processing a medical imaging diagnosis task, real-time computational parameters in the data stream indicated a sudden increase in computational time for a certain neural network layer. The timestamp indicated that the frame processing time was 50% longer than the average. The comprehensive monitoring data packets collected in step S41, combined with historical monitoring data, revealed that this anomaly was similar to past anomalies that resulted in compute core frequency reduction due to insufficient cooling. Step S42 identified a "computing core frequency reduction" anomaly with a medium severity rating. Step S43 invoked the policy template library and matched the "reduce computational precision to 16-bit floating point" and "enable full fan speed" policies. Evaluation revealed that while reducing precision had a minimal impact on diagnostic results (a loss of approximately 2%), it significantly reduced computational time. Furthermore, full fan speed could reduce core temperature by 10°C within 5 seconds. Step S44, based on historical data, determined that the "reduce precision" policy had an effectiveness score of 0.78 in similar scenarios. Taking into account the current cooling status, the policy weights were updated, prioritizing execution of this policy. Fans were also enabled to assist with cooling, ultimately restoring computational efficiency without significantly impacting diagnostic accuracy.

[0073] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.

[0074] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A neural network reasoning performance analysis method for NPU computing architecture, characterized by: The steps include: S1. Obtain the original data of the NPU computing architecture and standardize it to obtain standardized computing data. Based on the standardized computing data, extract the computing unit characteristics, computing reasoning task characteristics, and architecture load characteristics to generate a performance feature vector. Based on the performance feature vector, identify the reasoning performance requirements and determine the analysis dimensions to obtain preliminary analysis result data. S2, integrating the preliminary analysis result data and the standardized calculation data to generate an enhanced analysis matrix; Based on the enhanced analysis matrix, multi-dimensional performance analysis is performed to obtain performance analysis result data; Analyze the performance analysis result data, calculate the reasoning performance analysis necessity score and architecture impact value, and form analysis decision data; Conduct multiple verifications on analytical decision data to generate verified decision data; S3: Based on the pre-stored original information of the NPU computing feature library and the verified decision data, an enhanced computing feature space is constructed; Based on the enhanced computing feature space, computing matching and parameter optimization are performed to generate an optimized computing selection plan; the reasoning path of the optimized computing selection plan is optimized to form an optimized reasoning plan; Based on the optimized reasoning scheme, a multi-dimensional pre-inspection is performed before analysis, and the pre-inspection report data is finally output.

2. The neural network reasoning performance analysis method for NPU computing architecture according to claim 1, characterized in that: Step S1 is further as follows: S11. Obtaining NPU computing architecture raw data including computing unit configuration, timestamp, core identifier, and task identifier; converting the computing unit configuration in the NPU computing architecture raw data into a unified coding format, removing outliers and redundant information, and obtaining processed input data; performing data length normalization processing based on the processed input data to generate standardized computing data; S12. Based on the standardized computing data, calculate the computing unit parameter values, extract the key computing feature set, identify the reasoning task type, and generate basic feature data; based on the basic feature data, use a preconfigured task analysis model to calculate the performance requirement probability distribution and task vector of the reasoning task to obtain task feature data; obtain the NPU historical records, and extract the architecture status information based on the historical records; combine the architecture status information with the task feature data to form a performance feature vector; S13. Based on the performance feature vector, use the preset feature-requirement mapping rules to identify the reasoning performance requirement type, calculate the requirement priority score, and generate the requirement feature data; analyze the requirement feature data, estimate the computing resource requirement value and the time window requirement value, and obtain the resource requirement data; based on the requirement feature data and resource requirement data, identify the dependency relationship between the requirements and construct a dependency relationship graph; based on the reasoning performance requirement type, priority score, resource requirement data and dependency relationship graph, generate preliminary analysis result data.

3. The neural network reasoning performance analysis method for NPU computing architecture according to claim 2, characterized in that: Step S2 is further as follows: S21. Converting the preliminary analysis result data into a feature matrix, converting the standardized calculation data into a vector representation, and combining the feature matrix and the vector representation to form an initial analysis matrix; extracting relevant historical calculation records and calculating historical information weights; fusing the historical information weights with the initial analysis matrix to generate an enhanced analysis matrix; S22. Analyze the enhanced analysis matrix to identify key performance objectives; Decompose the main performance goal into a set of sub-goals, build a goal dependency graph, and obtain the goal structure data; Calculate the resource requirement vector and target priority matrix of each sub-goal to generate target resource data; Analyze computing feature requirements and calculate computing importance weights based on key performance goals and historical computing records; Based on the calculated importance weights, the target structure data and target resource data are integrated into performance analysis result data; S23. Based on the performance analysis result data, calculate the reasoning performance analysis necessity score, evaluate the analysis risk value, and obtain analysis evaluation data; Based on the analysis and evaluation data, determine the analysis timing and generate an analysis priority list; Based on the analysis priority list, formulate adjustment strategies and form analysis strategy data; Integrate analytical assessment data and analytical strategy data to generate an analytical path map; calculate confidence scores based on the analytical path map to ultimately form analytical decision data; S24. Perform internal consistency verification on the analysis and decision-making data to generate consistency verification data; Based on consistency verification data, verify resource availability, check computing constraints and timing restrictions, and obtain feasibility assessment data. Based on consistency verification data and feasibility assessment data, calculate verification scores, mark risk points, and generate optimization suggestions. Based on the optimization suggestions, verified decision data is formed.

4. The neural network reasoning performance analysis method for NPU computing architecture according to claim 3, characterized in that: Step S3 is further as follows: S31, based on the pre-stored original information of the NPU computing feature library and the verified decision data, extract the functional feature vector, attribute index vector and resource requirement vector of each computing mode to generate static feature data; Calculate the historical effect matrix, average inference time vector, and resource consumption distribution of the computing mode to form dynamic feature data; construct a computing dependency graph, calculate the compatibility matrix of the computing mode and the computing combination effect tensor to obtain associated feature data; Integrate static feature data, dynamic feature data and associated feature data into an enhanced computing feature space; S32. Calculate the functional applicability based on the enhanced computational feature space; Based on the functional applicability, attribute constraint filtering is performed to generate an initial candidate calculation set; the context features of the current NPU are obtained and the context relevance score is calculated; Based on the context relevance score, the candidate calculation weights are adjusted and the initial candidate calculation set is reordered to obtain the optimized candidate calculation set; a set of feasible calculation combinations is constructed and the combination synergy score is calculated; Based on the combined synergy score, the optimal combination scheme is selected from the optimization candidate calculation set to form an optimized calculation selection scheme; S33. Based on the optimized calculation selection scheme, construct a reasoning dependency graph, calculate the critical path, and generate a parallel reasoning scheme; Based on the parallel reasoning scheme, the reasoning sequence data is formed; Based on the inference sequence data, a resource allocation matrix is ​​constructed to optimize the inference timing. Based on the optimized inference timing, a cache strategy is constructed to obtain resource optimization data. Based on resource optimization data, failure handling strategies, alternative solutions, and monitoring point sets are constructed to generate fault-tolerant mechanism data. Inference sequence data, resource optimization data, and fault-tolerant mechanism data are integrated into an optimized inference solution. S34. Perform an online status check on the computing mode in the optimized inference solution to verify resource adequacy, test interface response, and generate availability verification data. Based on the availability verification data, we conduct permission checks, risk assessments, and compliance verifications to generate security assessment data. Based on the security assessment data, we estimate response time, predict resource consumption, and calculate success probability to generate performance prediction data. Integrate usability verification data, safety assessment data, and performance prediction data into pre-inspection report data.

5. The neural network reasoning performance analysis method for NPU computing architecture according to claim 4, characterized in that: Step S12 is further as follows: S121. Reading computing unit segments from the standardized computing data, calculating parameter values, calculation values, and load values ​​for each computing unit segment, and generating computing statistical data; based on the computing statistical data, using an analysis tool to segment the computing unit segments, obtaining computing statistics, and generating computing feature data; Combine statistical data and characteristic data to build complete basic characteristic data; S122. Obtain computation information from the basic feature data, and based on the computation information, calculate the architectural association strength of each computation through the bidirectional association network of the model to generate computation-level task association data. Based on the computational-level task association data, a task similarity matrix is ​​constructed, key task units are calculated, and task unit data is formed; Map the task unit data to the predefined demand space, calculate the demand probability distribution, and obtain the task feature data; S123. Obtain the operation sequence in the pre-stored NPU historical records, construct a timing feature vector, and generate historical calculation data; obtain and analyze the current architecture state, including architecture usage time, operation rounds, and calculation continuity, to form architecture state data; perform feature fusion on the task feature data, historical calculation data, and architecture state data, and output the final performance feature vector.

6. The neural network reasoning performance analysis method for NPU computing architecture according to claim 4, characterized in that: Step S22 is further as follows: S221. Read the performance description information in the enhanced analysis matrix, build an architecture dependency tree using a model based on the performance description information, extract core analysis nodes, and generate analysis sequence data; Based on analyzing sequence data, identify key performance targets, calculate the strength of logical relationships between targets, and form target association data; Combine the analysis sequence data and target association data to output performance target data; S222. Based on the performance target data, a predefined model target decomposition template library is used to perform pattern matching, identify decomposable sub-target units, and generate an initial sub-target set; the execution conditions and completion criteria of each sub-target in the initial sub-target set are analyzed, a sub-target constraint relationship diagram is constructed, and target constraint data is obtained; Based on the target constraint data, the initial sub-target set is optimized and reorganized, and the sub-target sequence data is output; S223. Based on the sub-goal sequence data, extract the input-output dependency relationship of each sub-goal, construct a data flow graph, and generate data dependency data; Based on data dependency data, we analyze execution order constraints, identify parallel execution opportunities, build a target execution network, and generate execution dependency data. We also integrate data dependency data and execution dependency data to build a complete target dependency graph. S224. Obtain historical execution records, calculate the processing complexity and resource consumption characteristics of each sub-goal based on the sub-goal sequence data and historical execution records, and generate resource characteristic data; Based on resource feature data, the time sensitivity and priority factors of sub-goals are analyzed, and a target scheduling weight matrix is ​​constructed to form scheduling feature data. The resource feature data, scheduling feature data and target dependency graph are combined to output the final performance analysis result data.

7. The neural network reasoning performance analysis method for NPU computing architecture according to claim 4, characterized in that: Step S23 is further as follows: S231, reading resource demand information in the performance analysis result data, calculating the resource utilization threshold based on the resource demand information and pre-stored historical analysis records, and generating resource evaluation data; Based on resource assessment data, analyze the target completion time requirements, combine the current system load status, calculate the time pressure coefficient, and form time assessment data; Based on the resource evaluation data and time evaluation data, calculate the reasoning performance analysis necessity score matrix and output the analysis necessity data; S232. Based on the analysis necessity data, extract characteristic patterns of historical analysis failure cases, construct risk characteristic vectors, and generate risk pattern data; Based on risk pattern data, analyze the similarity between the current target and pre-stored historical high-risk scenarios, calculate multi-dimensional risk coefficients, and form risk assessment data; Based on risk model data and risk assessment data, build a risk-benefit assessment matrix and output analytical risk data; S233. Based on the analysis necessity data and the analysis risk data, construct a reasoning performance analysis time series network, calculate the optimal analysis time window, and generate analysis time series data; Based on analyzing time series data, analyze the priority dependencies between computing modes, establish analysis priority queues, and form priority data; Combine analysis time series data and priority data to build an analysis execution plan and output analysis strategy data; S234. Based on the analysis strategy data and pre-stored historical failure processing records, a fault processing decision tree is constructed to generate fault recovery data. Based on the fault recovery data, a multi-level adjustment plan is constructed, including alternative calculation chains and simplification strategies, to form adjustment strategy data; Integrate analysis strategy data, fault recovery data, and adjustment strategy data, calculate strategy reliability scores, build a complete analysis path map, and finally output analysis decision data.

8. The neural network reasoning performance analysis method for NPU computing architecture according to claim 4, characterized in that: Step S32 is further as follows: S321. Based on the functional feature vectors and decision requirements in the enhanced computing feature space, the model is used to calculate the functional applicability score of each computing mode to generate functional applicability data. Based on the functional applicability data, the attribute indicator vector is used to perform constraint filtering to select a computing set that meets the attribute requirements to form attribute filtering data. Combine functional applicability data and attribute filtering data to build a preliminary calculation list and its scoring matrix, and output initial candidate data; S322. Obtain context features of the current NPU, including timing window, resource status, and target priority, and generate context feature data; Based on contextual feature data, we analyze the computing effects in historically similar scenarios, build a scenario correlation matrix, and generate scenario matching data. Based on the context feature data and scene matching data, the context adjustment coefficient is calculated and the context score data is output; S323. Based on the initial candidate data and contextual scoring data, a dynamic weighting algorithm is used to adjust the calculated scores and generate adjusted weighting data. Obtain and update the calculation credibility score based on the historical success rate and stability indicators of the calculation model to form credibility data; Based on the adjusted weight data and credibility data, the preliminary calculation list is re-sorted and the optimized candidate data is output; S324. Based on the calculation combination feature tensor in the optimization candidate data, a feasible calculation combination solution set is constructed to generate combination solution data; Based on the combination scheme data, the synergistic effect scores of different combination schemes are calculated, including functional complementarity and attribute gains, to form synergistic evaluation data; Based on collaborative evaluation data, analyze the complexity and risk factors of the combined solution data, build a comprehensive evaluation matrix, and obtain solution evaluation data; S325. Based on the combination scheme data, collaborative evaluation data, and scheme evaluation data, a multi-objective optimization algorithm is used to calculate the comprehensive score of each combination scheme and generate optimized scoring data; Based on the optimized scoring data, the optimal combination scheme is selected, a detailed calculation and reasoning sequence is constructed, and reasoning sequence data is formed; the optimized scoring data and reasoning sequence data are integrated, and finally the optimized calculation selection scheme is output.

9. The neural network reasoning performance analysis method for NPU computing architecture according to claim 8, characterized in that: Also includes: S4, collecting the execution data flow and historical monitoring data during the neural network inference process to generate a comprehensive monitoring data packet; Based on the comprehensive monitoring data package, perform anomaly detection and early warning analysis, and output anomaly analysis result data; Generate dynamic optimization strategies based on abnormal analysis result data and comprehensive monitoring data packets to form optimization strategy set data; Based on the optimization strategy set data and pre-stored historical optimization effect data, adaptive learning is performed and the optimization update data package is finally output.

10. The neural network reasoning performance analysis method for NPU computing architecture according to claim 9, characterized in that: Step S4 is further as follows: S41. Read the real-time computing parameters in the execution data stream, extract the timestamp, computing load, and resource usage information, and generate real-time monitoring data; read the abnormal pattern library in the historical monitoring data, build an abnormal feature dictionary, and form historical monitoring data; perform feature fusion on the real-time monitoring data and the historical monitoring data, and output a comprehensive monitoring data package; S42. Based on the comprehensive monitoring data packet, use the pre-trained anomaly detection model to calculate the data deviation and generate deviation data; Based on the deviation data, identify the abnormal type and severity level, build an abnormal classification tree, and form abnormal classification data; Integrate the deviation data and abnormal classification data, and output abnormal analysis result data; S43. Based on the abnormality analysis result data, extract the abnormality impact range and duration, and generate impact assessment data; Based on the impact assessment data, the pre-stored policy template library is called, the corresponding adjustment rules are matched, and the initial optimization strategy is generated; Based on the initial optimization strategy, analyze the feasibility and adjustment cost of the strategy, build a strategy evaluation matrix, and form the optimization strategy set data; S44. Based on the optimized strategy set data, extract the historical application effect records of each strategy, calculate the strategy effectiveness score, and generate effect evaluation data; Based on the effect evaluation data, an adaptive learning algorithm is used to update the strategy weights, build the learning model parameters, and form learning update data; Integrate the effect evaluation data and learning update data, and finally output the optimized update data package.

Citation Information

Patent Citations

  • NPU power consumption optimization system and method based on neural network structure

    CN114217688A

  • Method and system for quantitatively deploying LSTM operator in embedded neural network accelerator

    CN117035011A

  • Neural network reasoning performance analysis method for NPU computing architecture

    CN117851207A

  • AI algorithm acceleration method, device and equipment and readable storage medium

    CN118536565A

  • AI reasoning optimization method and system of dynamic model switching framework for edge device

    CN120354954A