Reliability evaluation method and system for satellite-borne intelligent software neural network model
By constructing a benchmark neural network model and conducting metamorphosis testing, mutation testing and adversarial testing, combined with the confidence interval method, the problem of imperfect reliability evaluation system of on-board intelligent software was solved, and a comprehensive reliability evaluation and stability analysis of the neural network model of on-board intelligent software was achieved.
Patent Information
- Application Number
- CN202510576141.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-09-19
AI Technical Summary
The existing intelligent software reliability evaluation system cannot be directly applied to the field of spaceborne software. There is a lack of system documentation to guide its reliability testing process. The evaluation indicators are complicated and lack unified comprehensive measurement indicators, which makes it difficult to ensure the reliability, robustness and security of spaceborne intelligent software.
A benchmark neural network model is constructed, and the neuron coverage index is calculated. Through metamorphic testing, mutation testing, and adversarial testing, combined with the confidence interval method, a comprehensive reliability measurement index based on confidence interval is proposed. The index includes neuron coverage, metamorphic testing factor, mutation testing factor, and adversarial attack success rate, which is used to evaluate the reliability of the neural network model of onboard intelligent software.
A comprehensive reliability assessment of the onboard intelligent software neural network model was achieved, which comprehensively considered the adequacy of the test and the robustness and stability of the neural network model, and provided a probabilistic measurement method to guide the retraining of the neural network model and improve the stability of the confidence interval and center.
Smart Images

Figure CN120670188A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of aerospace software reliability testing, and in particular to a reliability testing and evaluation method and system for a spaceborne intelligent software neural network model. Background Art
[0002] In the aerospace sector, satellites operate in an extremely unique space environment. Not only must they withstand extreme temperatures, radiation, and other environmental factors, they also face increasingly severe electronic countermeasures. With the gradual establishment of satellite internet constellations, the number of satellites in orbit continues to climb, and satellite remote sensing data acquisition is evolving towards multi-platform, multi-angle, and multi-sensor capabilities. This trend has significantly increased the demand for on-orbit processing of multi-source information, and the volume of onboard intelligent application software has also continued to grow. In this context, ensuring the reliability of onboard software has become increasingly critical.
[0003] Traditional software quality assessments primarily focus on general metrics such as functional correctness, efficiency, and compatibility. However, these metrics don't fully guarantee practical application in the unique context of spaceflight. Most onboard intelligent software is developed using open-source deep learning frameworks such as PyTorch and TensorFlow, and algorithm performance is currently evaluated based on accuracy. However, in the complex spaceflight environment, accuracy alone is far from sufficient to ensure software reliability in actual operation.
[0004] In particular, intelligent software for target detection and recognition in the satellite sector faces significant challenges to its security and reliability due to the uncertainties of the on-orbit space environment, the limitations of remote sensing imaging detector performance, and the potential for adversarial attacks against detected targets. Currently, reliability testing of onboard intelligent software faces a series of prominent issues, including a lack of research on its underlying mechanisms, a lack of systematic process methods, and a lack of clear evaluation criteria.
[0005] "24028:2020 AI Trustworthiness Overview," published by the International Organization for Standardization (ISO / IEC), defines the reliability of AI algorithms as "the ability to consistently deliver expected behavior and results." However, in the software field, existing standards primarily focus on software engineering, software process management, and software quality, and standards addressing software issues from a reliability perspective are extremely scarce. In 2020, Xu et al. pointed out that evaluating algorithm performance solely based on accuracy cannot meet practical application needs. They constructed a deep learning algorithm reliability assessment system based on three aspects: data, model, and runtime framework, and elaborated on the evaluation methods for each metric. Niu et al. explored various defect localization techniques and algorithm evaluation methods, and proposed a deep learning algorithm model testing strategy based on convolutional neural networks. In 2022, Zhang et al. constructed an AI software reliability indicator system through theoretical research, identifying no fewer than seven factors influencing software reliability, including software composition, architecture, algorithm performance, training data, adversarial examples, application capabilities, and mission requirements, as well as a reliability indicator system. However, these studies failed to provide detailed quantitative analysis of each metric for the specific application scenario of spaceborne software.
[0006] Through searching patent documents, we found an invention patent with publication number CN114925613A, which discloses a method for evaluating software reliability equivalent in a space radiation environment based on deep learning. First, different combinations of space radiation environment conditions and shielding methods are set to simulate the single-particle flip rate. Then, the software system is modeled based on a continuous-time Markov model, reliability is defined based on continuous random logic, and a large amount of reliability simulation data is generated using probabilistic model detection technology. Next, a multi-channel-multivariable convolutional neural network is established, and the network is trained and tested with pre-processed simulation data to obtain a reliability equivalent evaluation model. This patent only focuses on the reliability evaluation of SRAM-based FPGA system software in a space radiation environment, and its application scenarios are limited.
[0007] In summary, existing intelligent software reliability evaluation systems cannot be directly applied to the field of onboard software. Onboard intelligent software lacks system documentation to guide its reliability testing process, and the evaluation indicators are complex and lack unified comprehensive metrics. This makes it difficult to ensure the reliability, robustness, and security of onboard intelligent software. Therefore, the research of a reliability testing and evaluation method and system for neural network models of onboard intelligent software has become a key task that needs to be solved urgently. Summary of the Invention
[0008] In view of the defects in the prior art, the purpose of the present invention is to provide a reliability testing and evaluation method and system for a space-borne intelligent software neural network model.
[0009] According to the present invention, a reliability testing and evaluation method for a spaceborne intelligent software neural network model is provided, comprising the following steps:
[0010] Step S1, constructing a benchmark neural network model and calculating the neuron coverage index;
[0011] Step S2, generating a test data set based on the benchmark neural network model and calculating test indicators;
[0012] Step S3: Calculate the comprehensive metric based on the test metric and the neuron coverage metric.
[0013] Preferably, step S1 includes the following sub-steps:
[0014] Step S1.1: Load an open-source ship remote sensing image dataset for the neural network model of the onboard intelligent software to be tested. The open-source ship remote sensing image dataset includes a test set, a training set, and a validation set. The training set, validation set, and test set all contain corresponding labels and classification information.
[0015] In step S1.2, a neural network model is trained using an open-source ship remote sensing image dataset to obtain a baseline neural network model and weight file. The model performance is then evaluated on the test set to obtain evaluation metrics including precision, recall, and mAP50. mAP50 refers to the mean average precision at an intersection-over-union (IoU) of 0.5 in the object detection task.
[0016] In step S1.3, add a hook function to the inference code of the baseline neural network model to capture the activation data before and after each neuron during the inference process of each image;
[0017] In step S1.4, based on the activation data, the neuron coverage index of the benchmark neural network model is calculated. The calculation formula is as follows: Assuming that the given benchmark neural network model N = {N1, N2, ..., N K}, where N k ={n k,1 ,n k,2 ,…},k=1,2,…,K, the benchmark neural network model has K layers, N k Represents the set of all neurons in the kth layer; the training set sample of the benchmark neural network model N is X train ={xt1,xt2,…}, the test set sample is X test ={x1,x2,…}, the activation value of neuron n when the input sample is X is f(n,X); the neuron coverage index Neuron Cov is:
[0018]
[0019] Among them, t represents the threshold value of whether the neuron is activated, and the average value of the neuron coverage index with different thresholds t in the range of 0.5-0.9 is recorded as NC 0.5:0.9 , n k,j represents the jth neuron in the kth layer, f(n k,j ,x) represents the activation value of the jth neuron in the kth layer when the input sample is x.
[0020] Preferably, the test indicators include a degradation test factor, a mutation test factor, and a counter-attack success rate, and step S2 includes the following sub-steps:
[0021] Step S2.1, calculating a degradation test factor through degradation testing based on a ship remote sensing image dataset and a benchmark neural network model;
[0022] Step S2.2, calculating the mutation test factor through mutation testing based on the ship remote sensing image dataset and the benchmark neural network model;
[0023] In step S2.3, based on the ship remote sensing image dataset and the benchmark neural network model, the success rate of the adversarial attack is calculated through adversarial testing.
[0024] Preferably, step S2.1 includes the following sub-steps:
[0025] Step S2.1.1, performing image transformation on the test set of the open source ship remote sensing image dataset to generate a metamorphic test dataset;
[0026] Step S2.1.2: Input the metamorphic test data set into the benchmark neural network model and calculate the metamorphic test index mAP50 mate ;
[0027] Step S2.1.3, based on the degradation test index, define the degradation test factor r mate , the formula is as follows:
[0028]
[0029] Where, eps = 1.0 × 10 -5 , used to prevent the denominator from being 0 or very close to 0.
[0030] Preferably, step S2.2 includes the following sub-steps:
[0031] Step S2.2.1, performing a mutation operation on the test set of the open source ship remote sensing image dataset to generate a mutated test dataset;
[0032] Step S2.2.2: Input the mutation test dataset into the benchmark neural network model and calculate the mutation test index mAP50 muta ;
[0033] Step S2.2.3, based on the mutation test index, define the mutation test factor r muta , the formula is as follows:
[0034]
[0035] Where, eps = 1.0 × 10 -5 , used to prevent the denominator from being 0 or very close to 0.
[0036] Preferably, step S2.3 includes the following sub-steps:
[0037] Step S2.3.1, constructing a generative adversarial network using the UEA method;
[0038] In step S2.3.2, the baseline neural network model is introduced into the generative adversarial network to generate adversarial samples with slight perturbations to form an adversarial test dataset.
[0039] Step S2.3.3: Use the adversarial test dataset to attack the benchmark neural network model and calculate the success rate r of the targeted adversarial attack with target t. attach , the calculation formula is as follows:
[0040]
[0041] Among them, N is the total number of samples, f() is the benchmark neural network model, x i ′ is an effective adversarial example.
[0042] Preferably, step S3 includes the following sub-steps:
[0043] Step S3.1, determining a baseline confidence interval based on the evaluation index;
[0044] Step S3.2, based on the baseline confidence interval, the test index and the neuron coverage index, correction is performed to obtain a corrected baseline confidence interval.
[0045] Preferably, step S3.1 includes the following sub-steps:
[0046] In step S3.1.1, based on the evaluation index in step S1.2, calculate the reference center of the reliability interval. The reference center is the ratio of the number of correctly predicted samples in the test set to the total number of samples. The calculation formula is as follows:
[0047]
[0048] Where n is the total number of samples and m is defined as:
[0049]
[0050] Among them, [] indicates rounding;
[0051] In step S3.1.2, the Wilson interval method is used to calculate the fiducial confidence interval based on the fiducial center. The formula is as follows:
[0052]
[0053] in, is the sample proportion, z is the standard normal distribution quantile, which depends on the confidence level; is a correction term introduced based on the sample size and confidence level; Expand the standard error to take into account the small sample effect; It is a normalization term that ensures that the upper and lower limits of the interval are within a reasonable range.
[0054] Preferably, step S3.2 includes the following sub-steps:
[0055] Step S3.2.1, based on the neuron coverage index NC 0.5:0.9 , calculate the neuron coverage factor r cover The formula is as follows:
[0056]
[0057] In step S3.2.2, based on the variation test factor and the neuron coverage factor, the interval length is corrected. The corrected interval length is:
[0058] Length new =Length×(1+r muta ×0.15+r cover ×0.2)
[0059] Step S3.2.3: Based on the degradation test factor and the success rate of the counterattack, the interval center is corrected. The corrected interval center is:
[0060]
[0061] Then the revised benchmark confidence interval is obtained.
[0062] The present invention also provides a reliability testing and evaluation system for a spaceborne intelligent software neural network model, comprising:
[0063] Module M1, builds a benchmark neural network model and calculates neuron coverage metrics;
[0064] Module M2 generates a test data set based on the benchmark neural network model and calculates the test indicators;
[0065] Module M3 calculates comprehensive metrics based on test metrics and neuron coverage metrics.
[0066] Compared with the prior art, the present invention has the following beneficial effects:
[0067] 1. The present invention is particularly suitable for ship detection scenarios in remote sensing images. By sequentially performing metamorphosis testing, mutation testing, and adversarial testing on the satellite-borne target detection software, and based on the neuron coverage obtained during the training and verification process of the neural network model, as well as various indicators in the metamorphosis testing, mutation testing, and adversarial testing, an innovative comprehensive reliability measurement indicator based on confidence intervals is proposed. This indicator comprehensively considers the adequacy of the test, the robustness and stability of the neural network model, and can comprehensively evaluate the reliability of the neural network model of the satellite-borne intelligent software. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] Other features, objects and advantages of the present invention will become more apparent upon reading the detailed description of non-limiting embodiments with reference to the following drawings:
[0069] Figure 1 This is an overall framework diagram of a reliability testing and evaluation method for a space-borne intelligent software neural network model in an embodiment of the present invention;
[0070] Figure 2 This is a flow chart for calculating and correcting the confidence interval center and length in an embodiment of the present invention. DETAILED DESCRIPTION
[0071] The present invention will be described in detail below with reference to specific embodiments. The following examples will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those skilled in the art, several changes and improvements can be made without departing from the scope of the present invention. These all fall within the scope of protection of the present invention.
[0072] The present invention discloses a reliability testing and evaluation method and system for a neural network model of onboard intelligent software. The method aims to solve the problem that the reliability evaluation system of existing onboard intelligent software is imperfect and lacks research on the mechanism of neural networks, which makes it difficult to evaluate the reliability of the output results of the neural network model. The method of the present invention includes a reliability testing process and a measurement index calculation, wherein a comprehensive reliability measurement method based on a confidence interval is proposed according to the neuron coverage obtained during the test of the neural network model, as well as various indicators in the decay test, mutation test, and adversarial test. The center of the confidence interval represents the reliability of the model itself, and the length of the interval represents the confidence level of the reliability test. This probabilistic measurement method can guide the retraining of the neural network model, steadily improve the confidence interval and the confidence center, and at the same time reflect the fluctuation range of the neural network performance under different input scenarios, thereby comprehensively evaluating the reliability level of the neural network model.
[0073] Example 1:
[0074] Figure 1 This is an overall framework diagram of a reliability testing and evaluation method for a space-borne intelligent software neural network model in an embodiment of the present invention.
[0075] like Figure 1 As shown, in order to test and evaluate the reliability of onboard intelligent software and quantify the reliability of the neural network model through measurement indicators, this embodiment provides a reliability testing and evaluation method for the neural network model of onboard intelligent software, including the following steps:
[0076] Step S1: construct a benchmark neural network model and calculate the neuron coverage index.
[0077] Specifically, step S1 includes the following sub-steps:
[0078] Step S1.1: Load the open source ship remote sensing image dataset for the neural network model of the spaceborne intelligent software to be tested. The open source ship remote sensing image dataset includes a test set, a training set, and a validation set. The training set, validation set, and test set all contain corresponding labels and classification information.
[0079] In this embodiment, the open-source ship remote sensing image dataset only contains aerial satellite images of ship targets. The test set includes 1573 images, the training set includes 9697 images, and the validation set includes 2165 images.
[0080] In step S1.2, the neural network model is trained using an open-source ship remote sensing image dataset to obtain a baseline neural network model and weight file. The model performance is then evaluated on the test set to obtain evaluation metrics including precision, recall, and mAP50. mAP50 refers to the average precision when the intersection over union (IoU) is 0.5 in the object detection task.
[0081] In this embodiment, the neural network model is trained for no less than 100 rounds.
[0082] In step S1.3, add a hook function to the inference code of the baseline neural network model to capture the activation data before and after each neuron during the inference process of each image;
[0083] In step S1.4, based on the activation data, the neuron coverage index of the benchmark neural network model is calculated. The calculation formula is as follows: Assuming that the given benchmark neural network model N = {N1, N2, ..., N K}, where N k ={n k,1 ,n k,2 ,…},k=1,2,…,K, the benchmark neural network model has K layers, N k Represents the set of all neurons in the kth layer; the training set sample of the benchmark neural network model N is X train ={xt1,xt2,…}, the test set sample is X test ={x1,x2,…}, the activation value of neuron n when the input sample is X is f(n,X); the neuron coverage index Neuron Cov is:
[0084]
[0085] Among them, t represents the threshold value of whether the neuron is activated, and the average value of the neuron coverage index with different thresholds t in the range of 0.5-0.9 is recorded as NC 0.5:0.9 , n k,j represents the jth neuron in the kth layer, f(n k,j ,x) represents the activation value of the jth neuron in the kth layer when the input sample is x.
[0086] Step S2: Generate a test data set based on the benchmark neural network model and calculate test indicators.
[0087] In this embodiment, the test indicators include the degradation test factor, the mutation test factor, and the anti-attack success rate. Step S2 includes the following sub-steps:
[0088] Step S2.1, based on the ship remote sensing image dataset and the benchmark neural network model, calculate the degradation test factor through degradation test.
[0089] Specifically, step S2.1 includes the following sub-steps:
[0090] Step S2.1.1, perform image transformation on the test set of the open source ship remote sensing image dataset to generate a metamorphic test dataset.
[0091] In this embodiment, the image transformation for constructing the transformation relationship includes:
[0092] Translation: shift the image horizontally or vertically by a certain pixel value;
[0093] Flip, flip the image horizontally or vertically;
[0094] Rotate, which rotates the image slightly;
[0095] Cut, crop part of the image;
[0096] Zoom, which slightly scales the image;
[0097] Color inversion, inverts the color of the image (subtracts the color value of each pixel from 255);
[0098] Contrast, which adjusts the difference between light and dark areas in an image by increasing or decreasing the contrast of the image;
[0099] Brightness: adjust the image brightness to increase or decrease the overall brightness;
[0100] Blur, applies a Gaussian blur operation to the image;
[0101] Grayscale, convert the image into grayscale;
[0102] Gaussian Noise, adds Gaussian noise to the image.
[0103] Step S2.1.2: Input the metamorphic test data set into the benchmark neural network model and calculate the metamorphic test index mAP50 mate ;
[0104] Step S2.1.3, based on the degradation test index, define the degradation test factor r mate , the formula is as follows:
[0105]
[0106] Where, eps = 1.0 × 10 -5 , used to prevent the denominator from being 0 or very close to 0.
[0107] In step S2.2, based on the ship remote sensing image dataset and the benchmark neural network model, the mutation test factor is calculated through mutation testing.
[0108] Specifically, step S2.2 includes the following sub-steps:
[0109] In step S2.2.1, a mutation operation is performed on the test set of the open source ship remote sensing image dataset to generate a mutated test dataset.
[0110] In this embodiment, the mutation operations for constructing the mutation relationship include:
[0111] Data replication, duplication of specific types of data;
[0112] The label is wrong, change the data label;
[0113] Data loss, randomly delete training data;
[0114] Data scrambling, disrupting the order of data;
[0115] Noise perturbation, adding noise to the data.
[0116] Step S2.2.2: Input the mutation test dataset into the benchmark neural network model and calculate the mutation test index mAP50 muta .
[0117] Step S2.2.3, based on the mutation test index, define the mutation test factor r muta , the formula is as follows:
[0118]
[0119] Where, eps = 1.0 × 10 -5 , used to prevent the denominator from being 0 or very close to 0.
[0120] In step S2.3, based on the ship remote sensing image dataset and the benchmark neural network model, the success rate of the adversarial attack is calculated through adversarial testing.
[0121] Specifically, step S2.3 includes the following sub-steps:
[0122] Step S2.3.1, using the UEA (unified and efficient adversary) method to construct a generative adversarial network;
[0123] In step S2.3.2, the baseline neural network model is introduced into the generative adversarial network to generate adversarial samples with slight perturbations to form an adversarial test dataset.
[0124] Step S2.3.3: Use the adversarial test dataset to attack the benchmark neural network model and calculate the success rate r of the targeted adversarial attack with target t. attach , the calculation formula is as follows:
[0125]
[0126] Where N is the total number of samples, f(·) is the benchmark neural network model, and x i ′ is an effective adversarial example.
[0127] Step S3: Calculate the comprehensive metric based on the test metric and the neuron coverage metric.
[0128] Specifically, step S3 includes the following sub-steps:
[0129] Step S3.1, determining a baseline confidence interval based on the evaluation index of step S1.2.
[0130] Furthermore, step S3.1 includes the following sub-steps:
[0131] In step S3.1.1, based on the evaluation index, calculate the reference center of the reliability interval. The reference center is the ratio of the number of correctly predicted samples in the test set to the total number of samples. The calculation formula is as follows:
[0132]
[0133] Among them, n is the total number of samples, and for the target detection task, since both precision and recall are important, m is defined as:
[0134]
[0135] Among them, [] indicates rounding;
[0136] In step S3.1.2, the Wilson Score Interval method is used to calculate the benchmark confidence interval based on the benchmark center.
[0137] Specifically, for each sample in the test set of the open source ship remote sensing image dataset, which conforms to the binomial distribution, that is, X~B(n,p), the formula for the Wilson interval is as follows:
[0138]
[0139] Where p∧ is the correct sample proportion, z is the standard normal distribution quantile, which depends on the confidence level; is a correction term introduced based on the sample size and confidence level; Expand the standard error to take into account the small sample effect; It is a normalization term that ensures that the upper and lower limits of the interval are within a reasonable range.
[0140] Step S3.2, based on the baseline confidence interval, the test index and the neuron coverage index, correction is performed to obtain a corrected baseline confidence interval.
[0141] Figure 2 This is a flow chart for calculating and correcting the confidence interval center and length in an embodiment of the present invention.
[0142] like Figure 2 As shown, step S3.2 includes the following sub-steps:
[0143] Step S3.2.1, based on the neuron coverage index NC 0.5:0.9 , calculate the neuron coverage factor r cover The formula is as follows:
[0144]
[0145] In step S3.2.2, based on the variation test factor and neuron coverage factor in step S2.2, the interval length is corrected. The corrected interval length is:
[0146] Length new =Length×(1+r muta ×0.15+r cover ×0.2)
[0147] In this embodiment, the variation test factor and the neuron coverage index mainly detect the adequacy of the test set test, and are therefore used to correct the interval length. The better the test results and the higher the neuron coverage index, the smaller the fluctuation in the reliability level.
[0148] In step S3.2.3, based on the degradation test factor of step S2.1 and the adversarial attack success rate of step S2.3, the interval center is corrected. The corrected interval center is:
[0149]
[0150] Finally, the revised benchmark confidence interval is obtained.
[0151] In this embodiment, the degradation test factor and the anti-attack success rate mainly detect the robustness and stability of the model itself, and are therefore used to correct the interval center. The better the test result, the higher the estimate of the reliability index itself.
[0152] Example 2:
[0153] The present invention also provides a reliability testing and evaluation system for a satellite-borne intelligent software neural network model. The reliability testing and evaluation system for a satellite-borne intelligent software neural network model can be implemented by executing the process steps of a reliability testing and evaluation method for a satellite-borne intelligent software neural network model. That is, those skilled in the art can understand the reliability testing and evaluation method for a satellite-borne intelligent software neural network model as a preferred implementation of the reliability testing and evaluation system for a satellite-borne intelligent software neural network model.
[0154] Specifically, the reliability test and evaluation system for the spaceborne intelligent software neural network model includes:
[0155] Module M1, builds a benchmark neural network model and calculates neuron coverage metrics;
[0156] Module M2 generates a test data set based on the benchmark neural network model and calculates the test indicators;
[0157] Module M3 calculates comprehensive metrics based on test metrics and neuron coverage metrics.
[0158] Those skilled in the art will appreciate that, in addition to implementing the system and its various devices, modules, and units provided by the present invention in purely computer-readable program code, it is entirely possible to implement the same functions of the system and its various devices, modules, and units provided by the present invention in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, the system and its various devices, modules, and units provided by the present invention can be considered a hardware component, and the devices, modules, and units included therein for implementing various functions can also be considered as structures within the hardware component; the devices, modules, and units for implementing various functions can also be considered as both software modules implementing the method and structures within the hardware component.
[0159] The above describes specific embodiments of the present invention. It should be understood that the present invention is not limited to the specific embodiments described above, and those skilled in the art may make various changes or modifications within the scope of the claims, which do not affect the essence of the present invention. The embodiments of this application and the features in the embodiments may be combined with each other in any manner unless there is a conflict.
Claims
1. A reliability testing and evaluation method for a spaceborne intelligent software neural network model, characterized in that: The steps include: Step S1, constructing a benchmark neural network model and calculating the neuron coverage index; Step S2, generating a test data set based on the benchmark neural network model and calculating test indicators; Step S3: Calculate a comprehensive metric based on the test indicator and the neuron coverage indicator.
2. The reliability testing and evaluation method for a spaceborne intelligent software neural network model according to claim 1, characterized in that: The step S1 includes the following sub-steps: Step S1.1, loading an open-source ship remote sensing image dataset for the neural network model of the onboard intelligent software to be tested, wherein the open-source ship remote sensing image dataset includes a test set, a training set, and a validation set, wherein the training set, the validation set, and the test set all contain corresponding labels and classification information; Step S1.2: training the neural network model using the open-source ship remote sensing image dataset to obtain a baseline neural network model and weight file, and then evaluating the model performance on the test set to obtain evaluation metrics including precision, recall, and mAP50, where mAP50 refers to the mean average precision when the intersection-over-union ratio is 0.5 in the object detection task; Step S1.3: Add a hook function to the inference code of the benchmark neural network model to capture the activation data before and after each neuron during the inference process of each image; In step S1.4, based on the activation data, the neuron coverage index of the benchmark neural network model is calculated, and the calculation formula is as follows: Assuming that the given benchmark neural network model N = {N1, N2, ..., N K }, where N k ={n k,1 ,n k,2 ,…},k=1,2,…,K, the benchmark neural network model has K layers, N k Represents the set of all neurons in the kth layer; the training set sample of the benchmark neural network model N is X train ={xt1,xt2,…}, the test set sample is X test ={x1,x2,…}, the activation value of neuron n when the input sample is X is f(n,X); the neuron coverage index Neuron Cov is: Among them, t represents the threshold value of whether the neuron is activated, and the average value of the neuron coverage index with different thresholds t in the range of 0.5-0.9 is recorded as NC 0.5:0.9 , n k,j represents the jth neuron in the kth layer, f(n k,j ,x) represents the activation value of the jth neuron in the kth layer when the input sample is x.
3. The reliability testing and evaluation method for a spaceborne intelligent software neural network model according to claim 2, characterized in that: The test indicators include a degradation test factor, a mutation test factor, and a counterattack success rate. Step S2 includes the following sub-steps: Step S2.1, calculating a degradation test factor through a degradation test based on the ship remote sensing image dataset and the benchmark neural network model; Step S2.2, calculating a mutation test factor through mutation testing based on the ship remote sensing image dataset and the benchmark neural network model; Step S2.3, based on the ship remote sensing image dataset and the benchmark neural network model, calculate the adversarial attack success rate through adversarial testing.
4. The reliability testing and evaluation method for a spaceborne intelligent software neural network model according to claim 3, characterized in that: The step S2.1 includes the following sub-steps: Step S2.1.1, performing image transformation on the test set of the open source ship remote sensing image dataset to generate a metamorphic test dataset; Step S2.1.2, input the degradation test data set into the benchmark neural network model, and calculate the degradation test index mAP50 mate ; Step S2.1.3, based on the degradation test index, define the degradation test factor r mate , the formula is as follows: Where, eps = 1.0 × 10 -5 , used to prevent the denominator from being 0 or very close to 0.
5. The reliability testing and evaluation method for a spaceborne intelligent software neural network model according to claim 4, characterized in that: The step S2.2 includes the following sub-steps: Step S2.2.1, performing a mutation operation on the test set of the open source ship remote sensing image dataset to generate a mutated test dataset; Step S2.2.2: Input the mutation test data set into the benchmark neural network model and calculate the mutation test index mAP50 muta ; Step S2.2.3, based on the mutation test index, define the mutation test factor r muta , the formula is as follows: Where, eps = 1.0 × 10 -5 , used to prevent the denominator from being 0 or very close to 0.
6. The reliability testing and evaluation method for a spaceborne intelligent software neural network model according to claim 5, characterized in that: The step S2.3 includes the following sub-steps: Step S2.3.1, constructing a generative adversarial network using the UEA method; Step S2.3.2: Introduce the baseline neural network model into the generative adversarial network to generate adversarial samples with slight perturbations to form an adversarial test dataset; Step S2.3.3, use the adversarial test data set to attack the benchmark neural network model, and calculate the success rate r of the targeted adversarial attack with target t attach , the calculation formula is as follows: Where N is the total number of samples, f() is the benchmark neural network model, and x′ i is an effective adversarial example.
7. The reliability testing and evaluation method for a spaceborne intelligent software neural network model according to claim 1, characterized in that: The step S3 includes the following sub-steps: Step S3.1, determining a baseline confidence interval based on the evaluation index; Step S3.2: Based on the reference confidence interval, the test index and the neuron coverage index, correction is performed to obtain a corrected reference confidence interval.
8. The reliability testing and evaluation method for a spaceborne intelligent software neural network model according to claim 7, characterized in that: The step S3.1 includes the following sub-steps: In step S3.1.1, based on the evaluation index of step S1.2, the reference center of the reliability interval is calculated. The reference center is the ratio of the number of correctly predicted samples in the test set to the total number of samples. The calculation formula is as follows: Where n is the total number of samples and m is defined as: Among them, [] indicates rounding, Precision is the precision rate, and Recall is the recall rate; In step S3.1.2, the Wilson interval method is used to calculate the reference confidence interval based on the reference center. The formula is as follows: in, is the sample proportion, z is the standard normal distribution quantile, which depends on the confidence level; is a correction term introduced based on the sample size and confidence level; Expand the standard error to take into account the small sample effect; It is a normalization term that ensures that the upper and lower limits of the interval are within a reasonable range.
9. The reliability testing and evaluation method for a spaceborne intelligent software neural network model according to claim 8, characterized in that: The step S3.2 includes the following sub-steps: Step S3.2.1, based on the neuron coverage index NC 0.5:0.9 , calculate the neuron coverage factor r cover The formula is as follows: Step S3.2.2, based on the variation test factor and the neuron coverage factor, the interval length is corrected. The corrected interval length is: Length new =Length×(1+r muta ×0.15+r cover ×0.2) Step S3.2.3, based on the degradation test factor and the counterattack success rate, the interval center is corrected, and the corrected interval center is: Then the revised benchmark confidence interval is obtained.
10. A reliability test and evaluation system for a spaceborne intelligent software neural network model, characterized in that: include: Module M1, builds a benchmark neural network model and calculates neuron coverage metrics; Module M2 generates a test data set based on the benchmark neural network model and calculates test indicators; Module M3 calculates a comprehensive metric based on the test indicator and the neuron coverage indicator.
Citation Information
Patent Citations
Software reliability equivalent evaluation method in space radiation environment based on deep learning
CN114925613A