A fault-tolerant reinforcement method and system for neural networks used in target detection

CN122840281APending Publication Date: 2026-09-29HUAZHONG UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610928781.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-25
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0005]针对现有技术的以上缺陷或改进需求,本发明提供了一种用于目标检测的神经网络容错加固方法及系统,由此解决现有方案难以兼顾高性能与低资源成本的技术问题

Benefits of technology

通过为基准神经网络的各个ReLU激活层增设待优化的上界阈值,得到截断神经网络,并以截断神经网络中各个激活层的上界阈值构成的上界阈值向量为待优化对象,构建优化算法的目标函数、约束条件及搜索空间,以在约束条件下在搜索空间内进行迭代优化以最小化目标函数的目标函数值,得到最优的上界阈值向量,从而为关键层的ReLU激活函数动态增设最优上界阈值,能够有效截断并抑制因参数错误在目标检测推理过程中的异常传播,显著增强了模型面对噪声干扰或错误输入时的容错能力,同时确保了目标检测的高精度输出,且该容错加固方式仅需对激活层的上界阈值向量进行优化求解,无需引入额外校验位或辅助子网络,降低了容错加固的计算与存储开销。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122840281A_ABST
    Figure CN122840281A_ABST
Patent Text Reader

Abstract

This invention discloses a neural network fault-tolerant reinforcement method and system for target detection. By adding upper bound thresholds to each activation layer of a baseline neural network to obtain a truncated neural network, and using the upper bound threshold vector formed by the upper bound thresholds of each activation layer in the truncated neural network as the optimization object, an optimization algorithm's objective function, constraints, and search space are constructed. Iterative optimization is performed within the search space under constraints to minimize the objective function value, obtaining the optimal upper bound threshold vector. This dynamically adds the optimal upper bound threshold to the ReLU activation function of the key layer, effectively truncating and suppressing the abnormal propagation of parameter errors during the target detection inference process. This enhances the model's fault tolerance capability when facing noise interference or erroneous inputs, while ensuring high-precision output for target detection. Furthermore, this fault-tolerant reinforcement method only requires optimizing the upper bound threshold vector of the activation layers, reducing computational and storage overhead.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of artificial intelligence and computer vision technology, and more specifically, relates to a neural network fault-tolerant reinforcement method and system for target detection. Background Technology

[0002] With the rapid development of artificial intelligence technology, target detection technology based on Convolutional Neural Networks (CNNs) has demonstrated outstanding performance in many fields. Introducing this technology into aerospace scenarios is of significant strategic importance for promoting the intelligent transformation of space target detection missions. However, space is filled with high-energy particles such as protons, heavy ions, and electrons, which can easily induce single-event upsets (SEUs), causing distortion of the weights or feature data stored on the spaceborne computing platform, resulting in decreased network inference accuracy or even network collapse. Therefore, improving the fault tolerance of target detection networks is crucial to ensuring the reliable execution of space missions.

[0003] Currently, fault tolerance technologies for neural networks are mainly divided into two categories: active fault tolerance (such as error detection and correction codes, network structure optimization) follows the principle of "detecting errors first, then correcting them." Although the fault location is accurate, it requires the introduction of additional check bits or auxiliary sub-networks, resulting in large computational and storage overhead, making it difficult to adapt to extreme conditions such as limited resources on spaceborne platforms and multiple concurrent faults. Passive fault tolerance (such as error blocking, weight mapping) suppresses error propagation through built-in mechanisms. Although the overhead is lower, it often requires compromise on model accuracy and generalization ability when pursuing lightweight design, and it is easy to lose effective information or limit model expression.

[0004] In summary, existing solutions struggle to balance high performance and low cost, either consuming excessive resources or sacrificing model accuracy. Currently, there is a lack of a fault-tolerant hardening method that can guarantee the normal operation accuracy of the target detection network while maintaining extremely low resource overhead. Therefore, there is an urgent need to develop a highly reliable, low-overhead fault-tolerant protection technology to meet the stringent requirements of space intelligence missions. Summary of the Invention

[0005] In view of the above-mentioned defects or improvement needs of the existing technology, the present invention provides a neural network fault-tolerant reinforcement method and system for target detection, thereby solving the technical problem that the existing solutions are unable to achieve both high performance and low resource cost.

[0006] To achieve the above objectives, according to a first aspect of the present invention, a neural network fault-tolerant reinforcement method for target detection is provided, comprising: A benchmark neural network for object detection is obtained, and forward inference is performed on the benchmark neural network using a validation image dataset to obtain the first activation matrix output by each activation layer for each sample image and the first object detection accuracy value of the benchmark neural network; the activation layers of the benchmark neural network adopt the ReLU activation function. An upper bound threshold to be optimized is added to each activation layer of the baseline neural network to obtain a truncated neural network. Using the upper bound threshold vector formed by the upper bound thresholds of each activation layer in the truncated neural network as the object to be optimized, the objective function, constraints, and search space of the optimization algorithm are constructed. Iterative optimization is performed in the search space under the constraints to minimize the objective function value and obtain the optimal upper bound threshold vector. In one round of optimization, the truncated neural network is configured based on the upper bound threshold vector of the current attempt. Forward inference is performed on the truncated neural network using the verification image dataset to obtain the second activation matrix output by each activation layer for each sample image and the second target detection accuracy value of the truncated neural network. Faults are injected into the convolutional layers in the truncated neural network. Forward inference is performed on the truncated neural network after the fault injection using the verification image dataset to obtain the third activation matrix output by each activation layer for each sample image and the third target detection accuracy value of the truncated neural network after the fault injection, so as to calculate the current objective function value. The objective function includes a fault-free accuracy loss term, a fault performance loss term, and an activation value stability loss term. The closer the first target detection accuracy value is to the second target detection accuracy value, the smaller the fault-free accuracy loss term. The closer the third target detection accuracy value is to the second target detection accuracy value, the smaller the fault performance loss term. The closer the statistics of the second activation matrix output by each activation layer for each sample image are to the statistics of the first activation matrix output by the corresponding activation layer for each sample image, the smaller the activation value stability loss term. The closer the statistics of the third activation matrix output by each activation layer for each sample image are to the statistics of the second activation matrix output by the corresponding activation layer for each sample image, the smaller the activation value stability loss term. Based on the optimal upper bound threshold vector, the truncated neural network is configured to obtain an optimized target detection model.

[0007] According to the above-mentioned neural network fault-tolerant reinforcement method for target detection, faults are injected into the convolutional layers of the truncated neural network. Forward inference is then performed on the fault-injected truncated neural network using the verification image dataset to obtain the third activation matrix output by each activation layer for each sample image and the third target detection accuracy value of the fault-injected truncated neural network. Specifically, this includes: Step S31: Sort all pairs of fault types and convolutional layers according to fault type, and then select the current pair from the sorting results. Step S32: Select the current failure rate from multiple preset failure rates, and inject the failure corresponding to the failure type in the current tuple into the convolutional layer in the current tuple based on the failure rate; Step S33: Use the verification image dataset to perform forward inference on the truncated neural network after the injection of fault, and obtain the activation matrix output by each activation layer for each sample image under the current fault rate and the target detection accuracy value of the truncated neural network after the injection of fault. Step S34: Iteratively execute steps S32 and S33 until all preset failure rates have been traversed, and obtain the activation matrix of each sample image corresponding to the current tuple of each activation layer and the target detection accuracy value corresponding to the current tuple; wherein, the activation matrix of any sample image of the current tuple of any activation layer is the average value of the activation matrix output by the activation layer for the sample image under each failure rate, and the target detection accuracy value corresponding to the current tuple is the average value of the target detection accuracy value of the truncated neural network after injecting faults under each failure rate. Step S35: If any pair is not selected, proceed to step S31. Step S36: Calculate the average value of the activation matrix of the same sample image corresponding to all pairs of activation layers, and use it as the third activation matrix output by the corresponding activation layer for the corresponding sample image. Calculate the average value of the target detection accuracy value corresponding to all pairs of activation layers, and use it as the third target detection accuracy value.

[0008] According to the above-described neural network fault-tolerant reinforcement method for object detection, the upper and lower bounds of the search space corresponding to each activation layer are determined based on the following formula: ; in, This is the upper bound for the search of any activation layer. This serves as the lower bound for the search of any of the activation layers; and These are the 99.9th percentile and standard deviation of all elements in the first activation matrix output by any activation layer in the benchmark neural network for each sample image.

[0009] According to the above-described neural network fault-tolerant reinforcement method for target detection, the fault-free accuracy loss term is: ; in, For the fault-free accuracy loss item, This is the detection accuracy value for the second target. This is the detection accuracy value for the first target.

[0010] According to the above-described neural network fault-tolerant reinforcement method for target detection, the fault performance loss term is: ; in, For the aforementioned performance loss due to the fault, This is the detection accuracy value for the third target. This is the detection accuracy value for the second target.

[0011] According to the above-described neural network fault-tolerant reinforcement method for target detection, the activation value stability loss term is: ; ; ; in, This is the activation value stability loss term. To avoid stability loss during fault-free activation, This represents the activation stability loss under fault conditions, where n is the number of activation layers. Let be the coefficient of variation of each second activation matrix output by the i-th activation layer in the truncated neural network. Let p be the coefficient of variation of each first activation matrix output by the i-th activation layer in the baseline neural network, p be the number of sample images, and m be the number of elements in the activation matrix. Let j be the j-th element in the third activation matrix output by the i-th activation layer for the k-th sample image. It is the j-th element in the second activation matrix output by the i-th activation layer for the k-th sample image.

[0012] According to the above-described neural network fault-tolerant reinforcement method for target detection, the constraint condition is: ; in, This is the detection accuracy value for the second target. This is the detection accuracy value for the first target. This is the lower bound constraint coefficient for accuracy.

[0013] Based on the aforementioned neural network fault-tolerant reinforcement method for target detection, and based on the optimal upper bound threshold vector, the truncated neural network is configured to obtain an optimized target detection model, specifically including: Based on the optimal upper bound threshold vector, the truncated neural network is configured to obtain an initial optimization model; Based on the training image dataset, the initial optimization model is subjected to quantization perception training to obtain the target detection optimization model.

[0014] According to a second aspect of the present invention, a neural network fault-tolerant reinforcement system for target detection is provided, comprising: The benchmark model data acquisition unit is used to acquire a benchmark neural network for target detection, and to perform forward inference on the benchmark neural network using a validation image dataset to obtain the first activation matrix output by each activation layer for each sample image and the first target detection accuracy value of the benchmark neural network; the activation layers of the benchmark neural network adopt the ReLU activation function; A truncation mechanism insertion unit is used to add an upper bound threshold to be optimized to each activation layer of the baseline neural network to obtain a truncated neural network. An optimization unit is used to construct the objective function, constraints, and search space of an optimization algorithm, using the upper bound threshold vector formed by the upper bound thresholds of each activation layer in the truncated neural network as the object to be optimized. The optimization unit then performs iterative optimization within the search space under the constraints to minimize the objective function value and obtain the optimal upper bound threshold vector. In one round of optimization, the truncated neural network is configured based on the upper bound threshold vector of the current attempt. Forward inference is performed on the truncated neural network using the verification image dataset to obtain the second activation matrix output by each activation layer for each sample image and the second target detection accuracy value of the truncated neural network. Faults are injected into the convolutional layers in the truncated neural network. Forward inference is performed on the truncated neural network after the fault injection using the verification image dataset to obtain the third activation matrix output by each activation layer for each sample image and the third target detection accuracy value of the truncated neural network after the fault injection, so as to calculate the current objective function value. The objective function includes a fault-free accuracy loss term, a fault performance loss term, and an activation value stability loss term. The closer the first target detection accuracy value is to the second target detection accuracy value, the smaller the fault-free accuracy loss term. The closer the third target detection accuracy value is to the second target detection accuracy value, the smaller the fault performance loss term. The closer the statistics of the second activation matrix output by each activation layer for each sample image are to the statistics of the first activation matrix output by the corresponding activation layer for each sample image, the smaller the activation value stability loss term. The closer the statistics of the third activation matrix output by each activation layer for each sample image are to the statistics of the second activation matrix output by the corresponding activation layer for each sample image, the smaller the activation value stability loss term. The target model acquisition unit is used to configure the truncated neural network based on the optimal upper bound threshold vector to obtain an optimized target detection model.

[0015] According to a third aspect of the present invention, an electronic device is provided, comprising: a computer-readable storage medium and a processor; The computer-readable storage medium is used to store executable instructions; The processor is configured to read executable instructions stored in the computer-readable storage medium and execute the method as described in the first aspect.

[0016] According to a fourth aspect of the invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to perform the method as described in the first aspect.

[0017] According to a fifth aspect of the invention, a computer program product is provided, comprising a computer program or instructions that, when executed by a processor, implement the method as described in the first aspect.

[0018] In summary, compared with the prior art, the above-described technical solutions conceived by this invention can achieve the following beneficial effects: By adding upper bound thresholds to each ReLU activation layer of the baseline neural network to be optimized, a truncated neural network is obtained. The upper bound threshold vector formed by the upper bound thresholds of each activation layer in the truncated neural network is used as the object to be optimized. The objective function, constraints, and search space of the optimization algorithm are constructed. Iterative optimization is performed in the search space under the constraints to minimize the objective function value and obtain the optimal upper bound threshold vector. This dynamically adds the optimal upper bound threshold to the ReLU activation function of the key layer, which can effectively truncate and suppress the abnormal propagation of parameter errors in the object detection inference process, significantly enhance the model's fault tolerance capability in the face of noise interference or erroneous input, and ensure high-precision output of object detection. Moreover, this fault tolerance reinforcement method only needs to optimize the upper bound threshold vector of the activation layer, without introducing additional check bits or auxiliary subnetworks, reducing the computational and storage overhead of fault tolerance reinforcement.

[0019] Furthermore, by employing the parameter quantization method of quantization-aware training, floating-point parameters are converted into low-bit-width integer representations. This not only effectively limits the risk of exponential distortion of parameters in low-bit space and significantly improves the model's fault tolerance and robustness to parameter numerical distortion, but also significantly reduces model size and computational overhead, reduces storage resource consumption, and enables the target detection optimization model to meet the hardware constraints of real-time inference with higher throughput. Attached Figure Description

[0020] Figure 1 This is a flowchart illustrating the neural network fault-tolerant reinforcement method for target detection provided in an embodiment of the present invention.

[0021] Figure 2 This is a flowchart illustrating the optimization process provided in an embodiment of the present invention.

[0022] Figure 3 This is a schematic diagram of the quantitative perception training process provided in an embodiment of the present invention. Detailed Implementation

[0023] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.

[0024] This invention provides a method for fault-tolerant reinforcement of neural networks for target detection, such as... Figure 1 As shown, it includes: Step S1: Obtain a benchmark neural network for target detection, and perform forward inference on the benchmark neural network using a validation image dataset to obtain the first activation matrix output by each activation layer for each sample image and the first target detection accuracy value of the benchmark neural network; the activation layers of the benchmark neural network use the ReLU activation function. Step S2: Add an upper bound threshold to be optimized for each activation layer of the baseline neural network to obtain a truncated neural network. Step S3: Using the upper bound threshold vector formed by the upper bound thresholds of each activation layer in the truncated neural network as the object to be optimized, construct the objective function, constraints and search space of the optimization algorithm, and perform iterative optimization in the search space under the constraints to minimize the objective function value of the objective function, thereby obtaining the optimal upper bound threshold vector. In one round of optimization, the truncated neural network is configured based on the upper bound threshold vector of the current attempt. Forward inference is performed on the truncated neural network using the verification image dataset to obtain the second activation matrix output by each activation layer for each sample image and the second target detection accuracy value of the truncated neural network. Faults are injected into the convolutional layers in the truncated neural network. Forward inference is performed on the truncated neural network after the fault injection using the verification image dataset to obtain the third activation matrix output by each activation layer for each sample image and the third target detection accuracy value of the truncated neural network after the fault injection, so as to calculate the current objective function value. The objective function includes a fault-free accuracy loss term, a fault performance loss term, and an activation value stability loss term. The closer the first target detection accuracy value is to the second target detection accuracy value, the smaller the fault-free accuracy loss term. The closer the third target detection accuracy value is to the second target detection accuracy value, the smaller the fault performance loss term. The closer the statistics of the second activation matrix output by each activation layer for each sample image are to the statistics of the first activation matrix output by the corresponding activation layer for each sample image, the smaller the activation value stability loss term. The closer the statistics of the third activation matrix output by each activation layer for each sample image are to the statistics of the second activation matrix output by the corresponding activation layer for each sample image, the smaller the activation value stability loss term. Step S4: Based on the optimal upper bound threshold vector, configure the truncated neural network to obtain the target detection optimization model.

[0025] Here, a target neural network for object detection is obtained, and the activation functions of each activation layer in this target neural network are uniformly configured as ReLU activation functions to obtain a baseline neural network. Subsequently, a preset proportion (e.g., 20%) of sample data can be randomly extracted from a pre-constructed image dataset as a validation image dataset. The sample data (including sample images and the target locations they contain) in this validation image dataset is input into the network model for forward inference to obtain the first activation matrix output by each activation layer for each sample image and the first object detection accuracy value of the baseline neural network. This first object detection accuracy value characterizes the accuracy of the baseline neural network in detecting targets in the sample images of the validation image dataset. In some embodiments, the object detection accuracy value of the model can be quantified and evaluated using the mAP50 metric, including the first object detection accuracy value here and the second and third object detection accuracy values ​​described later. After acquiring the first activation matrix output by each activation layer for each sample image, the statistics of the first activation matrix output by each activation layer can be calculated to characterize the characteristics of the output data of that activation layer. In some embodiments, the statistics of the first activation matrix include standard deviation, mean, and 99th percentile, etc. Specifically, for any activation layer, the first activation matrix output by the activation layer for each sample image can be fully expanded into a one-dimensional vector, and then the standard deviation, mean, and 99th percentile of each element in the one-dimensional vector can be calculated.

[0026] Then, a truncation mechanism is introduced to take advantage of the characteristics of the ReLU activation function. An upper bound threshold to be optimized is added to the ReLU activation function of each activation layer, which is denoted as the CReLU (Clamp ReLU) activation function, and its expression is as follows: ; Where c is the upper bound threshold of the current activated layer.

[0027] The upper bound thresholds of each activation layer form an upper bound threshold vector C, which serves as the target for subsequent optimization. The expression is as follows: .

[0028] in This represents the upper bound threshold of the i-th activation layer.

[0029] The neural network with the added upper threshold is called a truncated neural network.

[0030] The value of the upper bound threshold vector directly affects the accuracy and fault tolerance of the truncated neural network. Therefore, it is necessary to use optimization algorithms to find the optimal upper bound threshold vector.

[0031] Specifically, taking the upper bound threshold vector, formed by truncating the upper bound thresholds of each activation layer in the neural network, as the object to be optimized, the objective function, constraints, and search space of the optimization algorithm are constructed. Iterative optimization is then performed within the search space under these constraints to minimize the objective function value, thereby obtaining the optimal upper bound threshold vector. In some embodiments, the optimization algorithm may employ a Bayesian optimization algorithm.

[0032] In some embodiments, the upper and lower bounds of the search space corresponding to each activation layer are determined based on the following formula: ; in, This is the upper bound for the search of any activation layer. This serves as the lower bound for the search of the activation layer; and These are the 99.9th percentile and standard deviation of all elements in the first activation matrix output by the activation layer of the benchmark neural network for each sample image.

[0033] In other embodiments, the constraints are: ; in, This is the accuracy value for the second target detection. This is the accuracy value for the first target detection. This is the lower bound constraint coefficient for accuracy.

[0034] During one round of optimization, the truncated neural network can be configured based on the upper bound threshold vector of the current attempt. Then, forward inference is performed on the truncated neural network using the aforementioned validation image dataset to obtain the second activation matrix output by each activation layer for each sample image and the second object detection accuracy value of the truncated neural network. Based on the second activation matrix output by each activation layer for each sample image, statistics of these second activation matrices can also be collected. These statistics are consistent with the statistics of the first activation matrix mentioned above, and will not be repeated here.

[0035] Simultaneously, during each round of iterative optimization, fault injection simulations are performed on the network parameters of the truncated neural network to evaluate its performance under abnormal conditions, thereby guiding the optimization algorithm. Specifically, faults can be injected into the convolutional layers of the truncated neural network, and then forward inference is performed on the fault-injected truncated neural network using a validation image dataset. This yields the third activation matrix output by each activation layer for each sample image and the third object detection accuracy value of the fault-injected truncated neural network, which is then used to calculate the current objective function value.

[0036] In some embodiments, the fault injection process includes the following steps: Step S31: After sorting all pairs of fault types and convolutional layers according to fault type, select the current pair from the sorting results. Pairs of the same fault type can be grouped together, allowing for the sequential injection of the corresponding fault into each convolutional layer for the same fault type. Fault types include bit flipping, zeroing error, and 1ting error.

[0037] Step S32: Select the current failure rate from multiple preset failure rates, and inject the failure corresponding to the failure type in the current tuple into the convolutional layer of the current tuple based on the current failure rate. When injecting the failure, the parameters of the convolutional layer can be changed based on the failure type in the current tuple.

[0038] Step S33: Using the verification image dataset, perform forward inference on the truncated neural network after the injection of faults, and obtain the activation matrix output by each activation layer for each sample image at the current fault rate, as well as the target detection accuracy value of the truncated neural network after the injection of faults. In some embodiments, this step can be repeated a preset number of times (e.g., 20 times) and the average value of the results obtained from these preset number of times can be used as the activation matrix output by each activation layer for each sample image at the current fault rate, as well as the target detection accuracy value of the truncated neural network after the injection of faults.

[0039] Step S34: Iteratively execute steps S32 and S33 until all preset failure rates have been traversed, obtaining the activation matrix of each sample image corresponding to the current tuple for each activation layer and the target detection accuracy value corresponding to the current tuple. Specifically, the activation matrix of any sample image corresponding to the current tuple for any activation layer is the average value of the activation matrix output by that activation layer for that sample image at each failure rate, and the target detection accuracy value corresponding to the current tuple is the average target detection accuracy value of the truncated neural network after injecting faults at each failure rate.

[0040] In step S35, if any pair is not selected, proceed to step S31.

[0041] Step S36: Calculate the average value of the activation matrix of the same sample image corresponding to all pairs of activation layers, and use it as the third activation matrix output by the corresponding activation layer for the corresponding sample image. Calculate the average value of the target detection accuracy value corresponding to all pairs of activation layers as the third target detection accuracy value.

[0042] In each round of optimization, after obtaining the second activation matrix and second target detection accuracy value output by each activation layer for each sample image, as well as the third activation matrix output by each activation layer for each sample image and the third target detection accuracy value of the truncated neural network after the injection of faults, the current objective function value can be calculated. Subsequently, if the upper bound threshold vector of the current attempt satisfies the constraint condition, the current objective function value is recorded; otherwise, the current objective function value is updated to a preset penalty value (a large value).

[0043] In some embodiments, when the feature dimension of the activation layer At this point, the convergence efficiency and optimization accuracy of Bayesian optimization decrease significantly. Therefore, dimensionality reduction can be performed on the upper bound threshold vector C based on the search space and fault tolerance evaluation results. The specific dimensionality reduction strategy follows two main principles: Firstly, error-sensitive weak activation layer removal: Prioritize removing activation layers that maintain high detection accuracy under fault injection scenarios. This can be achieved by sequentially injecting faults into each convolutional layer of the truncated neural network (only one convolutional layer at a time), obtaining the target detection accuracy value corresponding to each convolutional layer under fault injection conditions (the target detection accuracy value of the entire network after any convolutional layer is injected with a fault is the target detection accuracy value corresponding to that convolutional layer). Then, the target detection accuracy value corresponding to each convolutional layer is used as the detection accuracy of the activation layer connecting the corresponding convolutional layers. ,like ( If the preset accuracy threshold is used, then the active layer is determined to be less sensitive to errors and can be marked as a discarded layer.

[0044] Secondly, optimize the elimination of redundant activation layers in the interval: if the difference between the upper and lower bounds of an activation layer is less than or equal to a preset value (e.g., 1), and the change in target detection accuracy caused by setting its upper bound threshold to the upper and lower bounds is less than a preset value (e.g., 0.5%), it indicates that the optimization space of the activation layer is limited, and it can be marked as a elimination layer.

[0045] By defining the upper bound threshold corresponding to all elimination layers as their corresponding search upper bound, the optimization algorithm can avoid optimizing the solution of the upper bound threshold of the elimination layers, thereby achieving dimensionality reduction of the upper bound threshold vector.

[0046] When the iteration termination condition is reached, the upper bound threshold vector corresponding to the minimum objective function value can be selected as the optimal upper bound threshold vector. In some embodiments, the optimal upper bound threshold vector can also be verified based on a test image set to confirm that the optimal upper bound threshold vector can achieve good model performance.

[0047] In some embodiments, the objective function includes a fault-free accuracy loss term, a fault performance loss term, and an activation value stability loss term. The fault-free accuracy loss term ensures the accuracy of the truncated neural network under fault-free conditions; the fault performance loss term ensures the model accuracy of the truncated neural network under fault conditions; and the activation value stability loss term ensures the stability of activation values ​​under both fault-free and fault-free conditions, thus preventing extreme solutions caused by the spread of local fault deviations to global activation value instability, thereby enhancing the robustness of the truncated neural network.

[0048] Specifically, the closer the accuracy value of the first target detection is to the accuracy value of the second target detection, the smaller the error-free accuracy loss term; the closer the accuracy value of the third target detection is to the accuracy value of the second target detection, the smaller the error-free performance loss term; the closer the statistics of the second activation matrix output by each activation layer for each sample image are to the statistics of the first activation matrix output by the corresponding activation layer for each sample image, the smaller the activation value stability loss term; the closer the statistics of the third activation matrix output by each activation layer for each sample image are to the statistics of the second activation matrix output by the corresponding activation layer for each sample image, the smaller the activation value stability loss term.

[0049] In some embodiments, the fault-free accuracy loss term is: ; in, This is the error-free accuracy loss item. This is the accuracy value for the second target detection. This is the accuracy value for the first target detection.

[0050] In other embodiments, the performance loss due to failure is: ; in, This is a performance loss item due to failure. This represents the accuracy value for the third target detection.

[0051] In other embodiments, the activation value stability loss term is: ; ; ; in, This is the activation value stability loss term. To avoid stability loss during fault-free activation, This represents the activation stability loss under fault conditions, where n is the number of activation layers. To truncate the coefficients of variation of each second activation matrix output from the i-th activation layer in a neural network, Let be the coefficient of variation of each first activation matrix output by the i-th activation layer in the baseline neural network, where p is the number of sample images and m is the number of elements in the activation matrix. Let j be the j-th element in the third activation matrix output by the i-th activation layer for the k-th sample image. Let be the j-th element of the second activation matrix output by the i-th activation layer for the k-th sample image. Here, the coefficient of variation of each first activation matrix output by the i-th activation layer is the ratio of the standard deviation to the mean of all elements in each first activation matrix, and the coefficient of variation of each second activation matrix output by the i-th activation layer is the ratio of the standard deviation to the mean of all elements in each second activation matrix.

[0052] The entire optimization process is as follows Figure 2 As shown.

[0053] After obtaining the optimal upper bound threshold vector, the truncated neural network can be configured based on this optimal upper bound threshold vector to obtain the target detection optimization model.

[0054] In some embodiments, the truncated neural network can be configured first based on the optimal upper bound threshold vector to obtain an initial optimized model. Then, based on the training image dataset, the initial optimized model can be subjected to quantization perception training to obtain an object detection optimized model.

[0055] Specifically, such as Figure 3 As shown in Table 1, after preparing the training image dataset, pseudo-quantization nodes are inserted into the initial optimization model, and the quantization configuration parameters are initialized. Then, the model with inserted pseudo-quantization nodes is fine-tuned using the training image dataset until the preset training epochs are reached, resulting in the object detection optimization model.

[0056] Table 1

[0057] In summary, the method provided by this invention adds an upper bound threshold to each ReLU activation layer of the baseline neural network to obtain a truncated neural network. Using the upper bound threshold vector formed by the upper bound thresholds of each activation layer in the truncated neural network as the optimization object, the objective function, constraints, and search space of the optimization algorithm are constructed. Iterative optimization is performed within the search space under constraints to minimize the objective function value, resulting in the optimal upper bound threshold vector. This dynamically adds the optimal upper bound threshold to the ReLU activation function of the key layer, effectively truncating and suppressing the abnormal propagation of parameter errors during object detection inference. This significantly enhances the model's fault tolerance capability when facing noise interference or erroneous input, while ensuring high-precision output of object detection. Furthermore, this fault-tolerant reinforcement method only requires optimizing the upper bound threshold vector of the activation layer, without introducing additional check bits or auxiliary subnetworks, reducing the computational and storage overhead of fault-tolerant reinforcement.

[0058] Furthermore, by employing the parameter quantization method of quantization-aware training, floating-point parameters are converted into low-bit-width integer representations. This not only effectively limits the risk of exponential distortion of parameters in low-bit space and significantly improves the model's fault tolerance and robustness to parameter numerical distortion, but also significantly reduces model size and computational overhead, reduces storage resource consumption, and enables the target detection optimization model to meet the hardware constraints of real-time inference with higher throughput.

[0059] The neural network fault-tolerant reinforcement system for target detection provided by the present invention will be described below. The neural network fault-tolerant reinforcement system for target detection described below can be referred to in correspondence with the neural network fault-tolerant reinforcement method for target detection described above.

[0060] This invention provides a neural network fault-tolerant reinforcement system for target detection, comprising: The benchmark model data acquisition unit is used to acquire a benchmark neural network for target detection, and to perform forward inference on the benchmark neural network using a validation image dataset to obtain the first activation matrix output by each activation layer for each sample image and the first target detection accuracy value of the benchmark neural network; the activation layers of the benchmark neural network adopt the ReLU activation function; A truncation mechanism insertion unit is used to add an upper bound threshold to be optimized to each activation layer of the baseline neural network to obtain a truncated neural network. An optimization unit is used to construct the objective function, constraints, and search space of an optimization algorithm, using the upper bound threshold vector formed by the upper bound thresholds of each activation layer in the truncated neural network as the object to be optimized. The optimization unit then performs iterative optimization within the search space under the constraints to minimize the objective function value and obtain the optimal upper bound threshold vector. In one round of optimization, the truncated neural network is configured based on the upper bound threshold vector of the current attempt. Forward inference is performed on the truncated neural network using the verification image dataset to obtain the second activation matrix output by each activation layer for each sample image and the second target detection accuracy value of the truncated neural network. Faults are injected into the convolutional layers in the truncated neural network. Forward inference is performed on the truncated neural network after the fault injection using the verification image dataset to obtain the third activation matrix output by each activation layer for each sample image and the third target detection accuracy value of the truncated neural network after the fault injection, so as to calculate the current objective function value. The objective function includes a fault-free accuracy loss term, a fault performance loss term, and an activation value stability loss term. The closer the first target detection accuracy value is to the second target detection accuracy value, the smaller the fault-free accuracy loss term. The closer the third target detection accuracy value is to the second target detection accuracy value, the smaller the fault performance loss term. The closer the statistics of the second activation matrix output by each activation layer for each sample image are to the statistics of the first activation matrix output by the corresponding activation layer for each sample image, the smaller the activation value stability loss term. The closer the statistics of the third activation matrix output by each activation layer for each sample image are to the statistics of the second activation matrix output by the corresponding activation layer for each sample image, the smaller the activation value stability loss term. The target model acquisition unit is used to configure the truncated neural network based on the optimal upper bound threshold vector to obtain an optimized target detection model.

[0061] This invention provides an electronic device, including: a computer-readable storage medium and a processor; The computer-readable storage medium is used to store executable instructions; The processor is configured to read executable instructions stored in the computer-readable storage medium and execute the method as described in any of the above embodiments.

[0062] This invention provides a computer-readable storage medium storing computer instructions that cause a processor to perform the method described in any of the above embodiments.

[0063] This invention provides a computer program product, including a computer program or instructions, which, when executed by a processor, implement the method described in any of the above embodiments.

[0064] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for fault-tolerant reinforcement of neural networks for target detection, characterized in that, include: A benchmark neural network for object detection is obtained, and forward inference is performed on the benchmark neural network using a validation image dataset to obtain the first activation matrix output by each activation layer for each sample image and the first object detection accuracy value of the benchmark neural network; the activation layers of the benchmark neural network adopt the ReLU activation function. An upper bound threshold to be optimized is added to each activation layer of the baseline neural network to obtain a truncated neural network. Using the upper bound threshold vector formed by the upper bound thresholds of each activation layer in the truncated neural network as the object to be optimized, the objective function, constraints, and search space of the optimization algorithm are constructed. Iterative optimization is performed in the search space under the constraints to minimize the objective function value and obtain the optimal upper bound threshold vector. In one round of optimization, the truncated neural network is configured based on the upper bound threshold vector of the current attempt. Forward inference is performed on the truncated neural network using the verification image dataset to obtain the second activation matrix output by each activation layer for each sample image and the second target detection accuracy value of the truncated neural network. Faults are injected into the convolutional layers in the truncated neural network. Forward inference is performed on the truncated neural network after the fault injection using the verification image dataset to obtain the third activation matrix output by each activation layer for each sample image and the third target detection accuracy value of the truncated neural network after the fault injection, so as to calculate the current objective function value. The objective function includes a fault-free accuracy loss term, a fault performance loss term, and an activation value stability loss term. The closer the first target detection accuracy value is to the second target detection accuracy value, the smaller the fault-free accuracy loss term. The closer the third target detection accuracy value is to the second target detection accuracy value, the smaller the fault performance loss term. The closer the statistics of the second activation matrix output by each activation layer for each sample image are to the statistics of the first activation matrix output by the corresponding activation layer for each sample image, the smaller the activation value stability loss term. The closer the statistics of the third activation matrix output by each activation layer for each sample image are to the statistics of the second activation matrix output by the corresponding activation layer for each sample image, the smaller the activation value stability loss term. Based on the optimal upper bound threshold vector, the truncated neural network is configured to obtain an optimized target detection model.

2. The neural network fault-tolerant reinforcement method for target detection as described in claim 1, characterized in that, Injecting faults into the convolutional layers of the truncated neural network, and using the verification image dataset to perform forward inference on the fault-injected truncated neural network, the third activation matrix output by each activation layer for each sample image and the third target detection accuracy value of the fault-injected truncated neural network are obtained, specifically including: Step S31: Sort all pairs of fault types and convolutional layers according to fault type, and then select the current pair from the sorting results. Step S32: Select the current failure rate from multiple preset failure rates, and inject the failure corresponding to the failure type in the current tuple into the convolutional layer in the current tuple based on the failure rate; Step S33: Use the verification image dataset to perform forward inference on the truncated neural network after the injection of fault, and obtain the activation matrix output by each activation layer for each sample image under the current fault rate and the target detection accuracy value of the truncated neural network after the injection of fault. Step S34: Iteratively execute steps S32 and S33 until all preset failure rates have been traversed, and obtain the activation matrix of each sample image corresponding to the current tuple of each activation layer and the target detection accuracy value corresponding to the current tuple; wherein, the activation matrix of any sample image of the current tuple of any activation layer is the average value of the activation matrix output by the activation layer for the sample image under each failure rate, and the target detection accuracy value corresponding to the current tuple is the average value of the target detection accuracy value of the truncated neural network after injecting faults under each failure rate. Step S35: If any pair is not selected, proceed to step S31. Step S36: Calculate the average value of the activation matrix of the same sample image corresponding to all pairs of activation layers, and use it as the third activation matrix output by the corresponding activation layer for the corresponding sample image. Calculate the average value of the target detection accuracy value corresponding to all pairs of activation layers, and use it as the third target detection accuracy value.

3. The neural network fault-tolerant reinforcement method for target detection as described in claim 1 or 2, characterized in that, The upper and lower bounds of the search space corresponding to each activation layer are determined based on the following formula: ; in, This is the upper bound for the search of any activation layer. This serves as the lower bound for the search of any of the activation layers; and These are the 99.9th percentile and standard deviation of all elements in the first activation matrix output by any activation layer in the benchmark neural network for each sample image.

4. The neural network fault-tolerant reinforcement method for target detection as described in claim 1, characterized in that, The fault-free accuracy loss term is: ; in, For the fault-free accuracy loss item, This is the detection accuracy value for the second target. This is the detection accuracy value for the first target.

5. The neural network fault-tolerant reinforcement method for target detection as described in claim 1, characterized in that, The failure performance loss item is: ; in, For the aforementioned performance loss due to the fault, This is the detection accuracy value for the third target. This is the detection accuracy value for the second target.

6. The neural network fault-tolerant reinforcement method for target detection as described in claim 1, characterized in that, The activation value stability loss term is: ; ; ; in, This is the activation value stability loss term. To avoid stability loss during fault-free activation, This represents the activation stability loss under fault conditions, where n is the number of activation layers. Let be the coefficient of variation of each second activation matrix output by the i-th activation layer in the truncated neural network. Let p be the coefficient of variation of each first activation matrix output by the i-th activation layer in the baseline neural network, p be the number of sample images, and m be the number of elements in the activation matrix. Let j be the j-th element in the third activation matrix output by the i-th activation layer for the k-th sample image. It is the j-th element in the second activation matrix output by the i-th activation layer for the k-th sample image.

7. The neural network fault-tolerant reinforcement method for target detection as described in claim 1, characterized in that, The constraints are as follows: ; in, This is the detection accuracy value for the second target. This is the detection accuracy value for the first target. This is the lower bound constraint coefficient for accuracy.

8. The neural network fault-tolerant reinforcement method for target detection as described in claim 1, characterized in that, Based on the optimal upper bound threshold vector, the truncated neural network is configured to obtain an optimized target detection model, specifically including: Based on the optimal upper bound threshold vector, the truncated neural network is configured to obtain an initial optimization model; Based on the training image dataset, the initial optimization model is subjected to quantization perception training to obtain the target detection optimization model.

9. A neural network fault-tolerant reinforcement system for target detection, characterized in that, include: The benchmark model data acquisition unit is used to acquire a benchmark neural network for target detection, and to perform forward inference on the benchmark neural network using a validation image dataset to obtain the first activation matrix output by each activation layer for each sample image and the first target detection accuracy value of the benchmark neural network; the activation layers of the benchmark neural network adopt the ReLU activation function; A truncation mechanism insertion unit is used to add an upper bound threshold to be optimized to each activation layer of the baseline neural network to obtain a truncated neural network. An optimization unit is used to construct the objective function, constraints, and search space of an optimization algorithm, using the upper bound threshold vector formed by the upper bound thresholds of each activation layer in the truncated neural network as the object to be optimized. The optimization unit then performs iterative optimization within the search space under the constraints to minimize the objective function value and obtain the optimal upper bound threshold vector. In one round of optimization, the truncated neural network is configured based on the upper bound threshold vector of the current attempt. Forward inference is performed on the truncated neural network using the verification image dataset to obtain the second activation matrix output by each activation layer for each sample image and the second target detection accuracy value of the truncated neural network. Faults are injected into the convolutional layers in the truncated neural network. Forward inference is performed on the truncated neural network after the fault injection using the verification image dataset to obtain the third activation matrix output by each activation layer for each sample image and the third target detection accuracy value of the truncated neural network after the fault injection, so as to calculate the current objective function value. The objective function includes a fault-free accuracy loss term, a fault performance loss term, and an activation value stability loss term. The closer the first target detection accuracy value is to the second target detection accuracy value, the smaller the fault-free accuracy loss term. The closer the third target detection accuracy value is to the second target detection accuracy value, the smaller the fault performance loss term. The closer the statistics of the second activation matrix output by each activation layer for each sample image are to the statistics of the first activation matrix output by the corresponding activation layer for each sample image, the smaller the activation value stability loss term. The closer the statistics of the third activation matrix output by each activation layer for each sample image are to the statistics of the second activation matrix output by the corresponding activation layer for each sample image, the smaller the activation value stability loss term. The target model acquisition unit is used to configure the truncated neural network based on the optimal upper bound threshold vector to obtain an optimized target detection model.

10. An electronic device, characterized in that, include: Computer-readable storage media and processors; The computer-readable storage medium is used to store executable instructions; The processor is configured to read executable instructions stored in the computer-readable storage medium and execute the neural network fault-tolerant reinforcement method for target detection as described in any one of claims 1-8.