A Trustworthiness Evaluation Method for Embedded Intelligent Computing Systems

By generating and evaluating generalization and robustness test cases, the problem of credibility evaluation of embedded intelligent computing systems is solved, and the reliability and adaptability of the system are improved.

CN115269369BActive Publication Date: 2025-06-03SHENYANG AIRCRAFT DESIGN INST AVIATION IND CORP OF CHINA
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210399418.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-15
Publication Date
2025-06-03
Estimated Expiration
2042-04-15

AI Technical Summary

Technical Problem

It is difficult for the prior art to effectively evaluate the credibility of embedded intelligent computing systems, especially in generalization and robust capabilities.

Method used

By generating generalization and robustness test cases, grouping them to form multiple sets of test sample sets, and testing them in PC-side and embedded-side algorithms respectively, evaluating them using a variety of generalization and robustness indicators, and finally fusing the evaluation values ​​to obtain the credibility evaluation values.

Benefits of technology

The reliability, availability and scenario adaptability of embedded intelligent computing systems have been improved, and the improvement of algorithms is guided through the system's credibility evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115269369B_ABST
    Figure CN115269369B_ABST
Patent Text Reader

Abstract

This application belongs to the technical field of test and evaluation, and particularly relates to a credibility evaluation method for an embedded intelligent computing system. The method includes: Step S1, generating generalization test cases and grouping the test cases; Step S2, respectively obtaining a first type of standard solution and a first type of test response through calculation of the generalization test sample set; Step S3, respectively solving using multiple generalization indicators, and fusing the solved results to obtain a generalization evaluation value; Step S4, generating robustness test cases and grouping the robustness test cases; Step S5, for each robustness test sample set, adding perturbations with multiple amplitudes, and respectively obtaining a second type of standard solution and a second type of test response through calculation of the perturbed data; Step S6, respectively solving using multiple robustness indicators, and fusing the solved results to obtain a robustness evaluation value; Step S7, using the generalization evaluation value and the robustness evaluation value as the credibility evaluation value for the embedded intelligent computing system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of test evaluation, and particularly relates to a credibility evaluation method for an embedded intelligent computing system. Background Art

[0002] Currently, the research on the credibility of intelligent systems is relatively scattered and no standards have been formed yet. From the perspective of single technology research, generalization ability and robustness are the most critical technical characteristics. To ensure the security and reliability of an embedded intelligent computing system, a credibility evaluation method for an embedded intelligent computing system needs to be designed to improve the reliability analysis ability of the embedded intelligent computing system. Summary of the Invention

[0003] To solve the above problems, this application provides a credibility evaluation method for an embedded intelligent computing system, which mainly includes:

[0004] Step S1: Generate generalization test cases, group the test cases, and obtain multiple groups of generalization test sample sets;

[0005] Step S2: Input the generalization test sample sets into the PC - side and embedded - side algorithms respectively, and obtain the first - type standard solutions and the first - type test responses respectively;

[0006] Step S3: Solve the first - type standard solutions and the first - type test responses respectively using multiple generalization metrics, and fuse the solved results to obtain a generalization evaluation value;

[0007] Step S4: Generate robustness test cases, group the robustness test cases, and obtain multiple groups of robustness test sample sets;

[0008] Step S5: For each robustness test sample set, add perturbations of multiple amplitudes, input the perturbed data into the PC - side and embedded - side algorithms respectively, and obtain the second - type standard solutions and the second - type test responses respectively;

[0009] Step S6: Solve the second - type standard solutions and the second - type test responses respectively using multiple robustness metrics, and fuse the solved results to obtain a robustness evaluation value;

[0010] Step S7: Use the generalization evaluation value and the robustness evaluation value as the credibility evaluation value for the embedded intelligent computing system.

[0011] Preferably, in step S1, when grouping the test cases, it includes obtaining multiple groups of test sample sets with different balances.

[0012] Preferably, the balance of the test sample set is calculated through a positive - negative sample balance, a class - sample balance, or a scenario - target balance distribution evaluation method.

[0013] Preferably, the number of the test sample sets is not less than 10 groups.

[0014] Preferably, the generalization indexes include accuracy, precision, recall rate, F1 value, confusion matrix, ROC curve, AUC area, Kappa coefficient, Hamming distance, and Jaccard similarity coefficient.

[0015] Preferably, in step S5, the perturbations include noise interference, geometric distortion interference, illumination change interference, cloud interference, or scale change interference.

[0016] Preferably, in step S5, the multi - amplitude perturbations include perturbations of 20 consecutive amplitudes.

[0017] Preferably, the robustness indexes include mean structural similarity, average confidence ACAC, average confidence of correct label ACTC when the adversarial attack is successful, maximum margin distance, strong neuron activation coverage SNAC, average Lp distortion, K - multi - segment neuron coverage, ENI, neuron boundary coverage NBC, noise capacity estimation, neuron coverage NC, perturbation - sensitive distance, Top - k neuron coverage, robustness to Gaussian blur, Top - k neuron pattern, and robustness to image compression.

[0018] This application evaluates the credibility of the embedded intelligent computer system through generalization and robustness, and can improve the reliability, availability, and scenario adaptation ability of the embedded intelligent computing system. Description of the Drawings

[0019] Figure 1 is a flowchart of a preferred embodiment of the credibility evaluation method for the embedded intelligent computing system of this application. Detailed Embodiments

[0020] To make the purpose, technical solutions, and advantages of the implementation of this application clearer, the technical solutions in the embodiments of this application will be described in more detail below with reference to the drawings in the embodiments of this application. In the drawings, the same or similar reference numerals denote the same or similar elements or elements with the same or similar functions from beginning to end. The described embodiments are some, but not all, of the embodiments of this application. The embodiments described below with reference to the drawings are exemplary and are intended to explain this application and should not be construed as limiting this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of this application without creative efforts shall fall within the scope of protection of this application. The embodiments of this application will be described in detail below with reference to the drawings.

[0021] The present application provides a credibility evaluation method for an embedded intelligent computing system, as follows Figure 1 shown, mainly including:

[0022] Step S1: Generate generalization test cases, group the test cases, and obtain multiple groups of generalization test sample sets;

[0023] Step S2: Input the generalization test sample sets into the PC - side and embedded - side algorithms respectively, and obtain the first - type standard solutions and the first - type test responses respectively;

[0024] Step S3: Solve the first - type standard solutions and the first - type test responses respectively using multiple generalization metrics, and fuse the solved results to obtain a generalization evaluation value;

[0025] Step S4: Generate robustness test cases, group the robustness test cases, and obtain multiple groups of robustness test sample sets;

[0026] Step S5: For each robustness test sample set, add perturbations with multiple amplitudes, input the perturbed data into the PC - side and embedded - side algorithms respectively, and obtain the second - type standard solutions and the second - type test responses respectively;

[0027] Step S6: Solve the second - type standard solutions and the second - type test responses respectively using multiple robustness metrics, and fuse the solved results to obtain a robustness evaluation value;

[0028] Step S7: Use the generalization evaluation value and the robustness evaluation value as the credibility evaluation value for the embedded intelligent computing system.

[0029] The present application constructs a generalization ability evaluation index system and a robustness ability evaluation index system, thus forming a method for evaluating the credibility of an embedded intelligent computing system. Among them, for the generalization ability evaluation of the embedded intelligent computing system, research on the construction method of the generalization ability evaluation sample space is carried out to generate a sample set for the generalization evaluation of the intelligent computing system; research on the calculation method of the generalization ability evaluation index is carried out to break through key technologies such as the sample space description model applicable to the generalization ability evaluation, the generalization ability evaluation method, the multi - index fusion method for ability description, and the evaluation comprehensive index calculation method; and a method, index, and algorithm module for the generalization evaluation of the embedded intelligent computing system are formed. For the robustness ability evaluation of the embedded intelligent computing system, research on the construction method of the anti - interference ability evaluation sample space, the anti - active interference ability evaluation model, the anti - passive interference ability evaluation model, and the construction and calculation method of the comprehensive index of the main and passive interference abilities is carried out; key technologies such as the sample space description method for anti - interference ability evaluation, the modeling method based on active and passive interference, and the comprehensive index calculation method are broken through, and a method, index, and algorithm module for the robustness evaluation of the embedded intelligent computing system are formed.

[0030] A method for evaluating the generalization ability of an embedded intelligent computing system constructs a training set and a test set for the target data generated by a data generation system, and realizes the evaluation of the generalization ability from the aspects of sample sparsity and balance. For sparse samples, methods such as equivalent class partitioning, centroid positioning method for pairwise boundary partitioning, and sample boundary evaluation method are used to define sparsity; methods for evaluating the balance of positive and negative samples, class samples, and scene / target balance distribution are used to evaluate balance. Different test sample sets with different balance and equilibrium are used to calculate 10 characterization indicators such as the target accuracy rate of the recognition decision system, and finally the performance indicators are fused to generate the final evaluation of the generalization ability. The specific ten indicators are as follows:

[0031] (1) Accuracy: Accuracy refers to the ratio of the number of correctly predicted samples to the total number of samples. Generally speaking, the higher the accuracy, the better the classifier. However, for the case of unbalanced data distribution, accuracy is not balanced and comprehensive enough.

[0032]

[0033] (2) Precision: Precision is the proportion of truly positive samples among all samples judged to be positive.

[0034]

[0035] (3) Recall: Recall is a measure of coverage, measuring how many positive examples are classified as positive, that is, the proportion of samples correctly predicted as positive among all actually positive samples.

[0036]

[0037] (4) F1-score: The F1-score is an index used in statistics to measure the accuracy of a binary classification model and is used to measure the accuracy of unbalanced data. It takes into account both the precision and recall of the classification model. The F1-score can be regarded as a weighted average of the model's precision and recall, with a maximum value of 1 and a minimum value of 0.

[0038]

[0039] (5) Confusion matrix: If for each class, you want to know the situation of misclassification between classes and check whether there is confusion between specific classes, you can use the confusion matrix to draw the detailed prediction results of the classification. For tasks with multiple classes, the confusion matrix clearly reflects the misclassification probability between classes.

[0040] (6)ROC curve: ROC is a comprehensive index reflecting the sensitivity and specificity of continuous variables. It reveals the mutual relationship between sensitivity and specificity by means of a graphical method. By setting multiple different critical values for continuous variables, a series of sensitivities and specificities are calculated, and then a curve is plotted with sensitivity as the vertical axis and (1 - specificity) as the horizontal axis.

[0041] (7)AUC area: AUC is the abbreviation of the area under the ROC curve. The value of AUC is the size of the area under the ROC curve. Usually, the value of AUC ranges from 0.5 to 1.0. The larger the AUC of a classifier, the higher the judgment accuracy.

[0042] (8)Kappa coefficient: The Kappa coefficient is a statistic for measuring the consistency of classification results and is the basis for measuring the stability of classifier performance. The larger the Kappa coefficient value, the more stable the classifier performance.

[0043]

[0044] p o is the sum of the number of correctly classified samples in each category divided by the total number of samples, that is, the overall classification accuracy.

[0045] Assume that the number of true samples in each category is a 1 , a 2 ,,..., a c

[0046] And the number of samples predicted for each category is b 1 , b 2 ,,..., b c

[0047] The total number of samples is n

[0048] Then there is

[0049] The Kappa coefficient is used for consistency testing and can also be used to measure classification accuracy. The calculation result ranges from -1 to 1, but usually kappa falls between 0 and 1 and can be divided into five groups to represent different levels of consistency: 0.0 - 0.20: very low consistency; 0.21 - 0.40: general consistency; 0.41 - 0.60: medium consistency; 0.61 - 0.80: high consistency; 0.81 - 1: almost perfect consistency.

[0050] (9)Hamming distance is used in scenarios where multiple labels of samples need to be classified. For a given sample i, is the prediction result for the jth label, is the true result of the jth label, and L is the number of labels, then and yi The Hamming distance between

[0051]

[0052] where 1(x) is the indicator function. When the prediction result is exactly the same as the actual situation, the distance is 0; when the prediction result is completely inconsistent with the actual situation, the distance is 1; when the prediction result is a proper subset or proper superset of the actual situation, the distance is between 0 and 1. The overall performance of the algorithm on the test set can be obtained by averaging the prediction situations of all samples. When the number of labels L is 1, it is equal to 1 - accuracy rate, and when L > 1, it also has good discrimination.

[0053] (10) Jaccard similarity coefficient: The Jaccard similarity coefficient is also used in scenarios where multiple labels of samples need to be classified. For a given sample i, is the prediction result, y i is the true result, and L is the number of labels. Then the Jaccard similarity coefficient of the i-th sample is

[0054]

[0055] The difference between it and the Hamming distance lies in the denominator. When the prediction result is exactly the same as the actual situation, the coefficient is 1; when the prediction result is completely inconsistent with the actual situation, the coefficient is 0; when the prediction result is a proper subset or proper superset of the actual situation, the distance is between 0 and 1. The overall performance of the algorithm on the test set can be obtained by averaging the prediction situations of all samples. When the number of labels L is 1, it is equal to the accuracy rate.

[0056] The technical route of the robust ability evaluation method for the embedded intelligent computing system constructs training samples through a data generation system, and realizes 1 - 20% interference control through small perturbation amplitude control. The control includes noise interference, geometric distortion interference, illumination change interference, cloud interference and scale change interference. For different interference types and amplitudes, analysis methods such as neuron coverage rate, neuron boundary coverage, noise capacity estimation, robustness to image compression, robustness to Gaussian blur, and maximum boundary distance are used to realize the estimation of robustness. The specific sixteen indicators are as follows:

[0057] (1) Average structural similarity: SSIM, as one of the commonly used indicators for quantifying the similarity between two images, is considered to be more in line with human visual perception than the Lp similarity. To evaluate the imperceptibility of adversarial samples, ASS is defined as the average similarity between all successfully attacked adversarial samples and their original samples, that is

[0058]

[0059] Among them, n represents the number of all successfully attacked adversarial samples. The larger the ASS value, the stronger the imperceptibility of the adversarial samples.

[0060] (2) Average Confidence of Adversarial Classification (ACAC): The average predicted confidence for the wrong class. It is defined as the average probability of all misclassified classes for all successfully attacked adversarial samples after adversarial attacks. The formula is defined as follows:

[0061]

[0062] Among them, n represents the number of all successfully attacked adversarial samples.

[0063] (3) Average Confidence of True Class when Adversarial Attack Succeeds (ACTC): The average confidence of the true class. Calculate the average of the predicted confidence levels for the true classes of the adversarial attack samples. It is used to evaluate to what extent the attack deviates from the true value.

[0064]

[0065] (4) Maximum Margin Distance: The distance between data points to the decision boundary measures the stability and robustness of the model in the worst case. Randomly find N orthogonal directions, and then calculate how much each direction needs to be moved for a model to change the classification label of the test set, to measure the maximum distance of the model's decision boundary.

[0066]

[0067] Among them, V represents a randomly generated set, φi(V) is the RMS distance to the model's decision boundary, and di is the maximum value of the distance to the decision boundary.

[0068] (5) Strong Neuron Activation Coverage (SNAC): Strong Neuron Activation Coverage measures how many corner cases are covered by a given test input T.

[0069] (6) Average Lp Distortion. Almost all attacks use the Lp norm distance (p = 0, 2, ∞) as the distortion metric for evaluation. Specifically, L0 calculates the number of pixels that change after perturbation; L2 calculates the Euclidean distance between the original example and the perturbed example; L∞ measures the maximum change in the full dimension of the adversarial sample. ALDp is the average normalized Lp distortion of all successfully attacked adversarial samples. The smaller the ALDp, the stronger the imperceptibility of the adversarial samples.

[0070]

[0071] Among them, n is the number of all successfully attacked adversarial samples.

[0072] (7) K - Multi - segment Neuron Coverage: Given a neuron n, the k - multi - segment neuron coverage measures the thoroughness of covering the given test input set T within the coverage range [lown, highn].

[0073] (8) ENI is a test set that combines adversarial attacks and natural noise. The lower the ENI value, the stronger the imperceptibility of the attack.

[0074] (9) Neuron Boundary Coverage NBC: Neuron Boundary Coverage measures how many corner regions (including upper and lower boundary values) are covered by the given test input set T.

[0075] (10) Noise Capacity Estimation. The robustness of adversarial samples can be estimated by the noise tolerance, which reflects the amount of noise that an adversarial sample can tolerate while maintaining the classification category unchanged. Specifically, NTE calculates the difference between the misclassification probability and the maximum probability of other classes.

[0076]

[0077] Where m is the total number of pixel points, δ{i,j} represents the j - th pixel point of the i - th example, R(x{i,j}) represents the square region near x{i,j}, and std represents the standard deviation function. The smaller the value of PSD, the stronger the imperceptibility of the adversarial sample.

[0078] (11) Neuron Coverage NC: Neuron Coverage is the ratio of the number of uniquely activated neurons in all test inputs to the total number of neurons in the DNN.

[0079] (12) Perturbation - Sensitive Distance: Perturbation - Sensitive Distance is used to evaluate the human's perception ability of perturbations.

[0080]

[0081] Where m is the total number of pixel points, δ{i,j} represents the j - th pixel point of the i - th example, R(x{i,j}) represents the square region near x{i,j}, and std represents the standard deviation function. The smaller the value of PSD, the stronger the imperceptibility of the adversarial sample.

[0082] (13) Top - k Neuron Coverage: The coverage of the top - k neurons measures the number of the k most active neurons on each layer. It is defined as the ratio of the total number of top - k neurons in each layer to the total number of neurons in the DNN.

[0083] (14) Robustness to Gaussian Blur: Gaussian blur is often used for image denoising in computer vision algorithms. Normally, a highly robust adversarial sample should maintain its misclassification effect after Gaussian blur. It can be defined as:

[0084]

[0085]

[0086] UA represents non - directed attack, TA represents directed attack, and the GB function represents Gaussian blur processing. The higher the RGB result, the stronger the robustness of the adversarial example.

[0087] (15) Top - k neuron pattern: Intuitively, the Top - k neuron pattern represents different activation scenarios of the top over - active neurons in each layer.

[0088] (16) Robustness to image compression: Image compression is often used for image denoising in computer vision algorithms. Normally, a highly robust adversarial example should maintain its misclassification effect after image compression. It can be defined as:

[0089]

[0090]

[0091] UA represents non - directed attack, TA represents directed attack, and the IC function represents image compression processing. The higher the RIC result, the stronger the robustness of the adversarial example.

[0092] The fusion of the results in step S3 and step S6 includes calculating the average value of all indicators as the fusion result.

[0093] When conducting the generalization index evaluation, since the test cases are not general and universal, to ensure the test effect and the effectiveness of the algorithm, the test cases for the generalization of the decision - making algorithm are 10 groups of data. After generating the test cases, they are sent to the PC - side algorithm to generate the standard solution, and sent to the embedded - side to generate the test response. The standard solution and the test response are evaluated.

[0094] When conducting the robustness index evaluation, since the test cases are not general and universal, to ensure the test effect and the effectiveness of the algorithm, the test cases for the robustness of the decision - making algorithm are 3 groups of data, and each group of data is added with 20 intensities of noise. After generating the test cases, they are sent to the PC - side algorithm to generate the standard solution, and sent to the embedded - side to generate the test response. The standard solution and the test response are evaluated.

[0095] The standard solution of the embedded intelligent system is generated by directly running the.exe file encapsulated by the PC algorithm. Then, the standard solution of the system can be obtained. Correspondingly, when the algorithm running on the embedded - side gets the test response, the test solution of the system can be obtained. When the data evaluation system receives the standard solution and the test response respectively, the test results can be seen on the data evaluation interface.

[0096] Table 1 presents the results of the generalization test.

[0097] Table 1 Generalization Test Results of the Intelligent Computing System

[0098]

[0099] As can be seen from the above table, this application can evaluate the credibility of the embedded intelligent computing system well. The evaluation results can show the advantages and disadvantages of the generalization and robustness of the algorithms of the embedded intelligent computing system, and thus better guide the improvement of the algorithms of the embedded intelligent computing system.

[0100] The above are only specific embodiments of this application, but the protection scope of this application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in this application should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claimed rights.

Claims

1. A credibility evaluation method for an embedded intelligent computing system, characterized in that, it includes: Step S1, generate generalization test cases, group the test cases, and obtain multiple groups of generalization test sample sets; Step S2, input the generalization test sample sets into the PC - side and embedded - side algorithms respectively, and obtain the first - type standard solutions and the first - type test responses respectively; Step S3, solve the first - type standard solutions and the first - type test responses respectively using multiple generalization metrics, and fuse the solved results to obtain a generalization evaluation value. The generalization metrics include accuracy, precision, recall, F1 - value, confusion matrix, ROC curve, AUC area, Kappa coefficient, Hamming distance, and Jaccard similarity coefficient; Step S4, generate robustness test cases, group the robustness test cases, and obtain multiple groups of robustness test sample sets; Step S5, for each robustness test sample set, add perturbations with multiple amplitudes, input the perturbed data into the PC - side and embedded - side algorithms respectively, and obtain the second - type standard solutions and the second - type test responses respectively; Step S6, solve the second - type standard solutions and the second - type test responses respectively using multiple robustness metrics, and fuse the solved results to obtain a robustness evaluation value. The robustness metrics include mean structural similarity, average confidence ACAC, average confidence of correct labels when adversarial attack is successful ACTC, maximum margin distance, strong neuron activation coverage SNAC, average Lp distortion, K - multi - segment neuron coverage, ENI, neuron boundary coverage NBC, noise capacity estimation, neuron coverage rate NC, perturbation - sensitive distance, Top - k neuron coverage, robustness to Gaussian blur, Top - k neuron pattern, and robustness to image compression; Step S7, use the generalization evaluation value and the robustness evaluation value as the credibility evaluation value for the embedded intelligent computing system.

2. The credibility evaluation method for an embedded intelligent computing system according to claim 1, characterized in that, in Step S1, when grouping the test cases, it includes obtaining multiple groups of test sample sets with different degrees of balance.

3. The credibility evaluation method for an embedded intelligent computing system according to claim 2, characterized in that, the balance of the test sample set is calculated by a positive - negative sample balance, class - sample balance, or scene - target balance distribution evaluation method.

4. The credibility evaluation method for an embedded intelligent computing system according to claim 1, characterized in that, the number of the test sample sets is not less than 10 groups.

5. The credibility evaluation method for an embedded intelligent computing system according to claim 1, characterized in that, in Step S5, the perturbations include noise interference, geometric distortion interference, illumination change interference, cloud interference, or scale change interference.

6. The credibility evaluation method for an embedded intelligent computing system according to claim 1, characterized in that, in Step S5, the perturbations with multiple amplitudes include 20 consecutive amplitudes of perturbations.

Citation Information

Patent Citations

  • System and methods for training and validation of an end-to-end artificially intelligent neural network for autonomous driving at scale

    US20250005378A1