Target detection robustness enhancement method based on texture feature similarity
By preprocessing and smoothing the target image, combining adaptive mechanism and attention mechanism, key feature pairs are selected, which solves the problem of insufficient robustness of traditional methods in complex environments and improves the accuracy and reliability of target detection.
Patent Information
- Application Number
- CN202510353295.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-11
AI Technical Summary
Traditional object detection methods lack adaptive processing capabilities for texture feature changes in complex environments, resulting in insufficient detection robustness and missed detection and false detection.
By preprocessing and smoothing the image to be detected, multi-level features are extracted using the Backbone network, adaptive mechanisms are used to filter out the feature pairs with the greatest differences, and abnormal detection scores are calculated through the attention mechanism, and abnormal samples are discarded to improve detection accuracy.
It enhances the robustness of the target detection system in complex environments, reduces false detection and missed detection, and ensures high-performance detection results.
Smart Images

Figure CN120298720A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to computer vision and image processing technologies, and particularly to a method for enhancing the robustness of object detection based on texture feature similarity. Background Art
[0002] With the continuous development of artificial intelligence and deep learning technologies, object detection systems have been widely applied in fields such as autonomous driving, security monitoring, and industrial inspection. In the field of object detection, robustness has always been a key and challenging issue. With the wide application of computer vision technologies in scenarios such as autonomous driving, security monitoring, and industrial inspection, the requirements for the stable and accurate operation of object detection algorithms in complex and changing environments are increasing. In practical applications, the images of objects to be detected are often affected by various interference factors, such as illumination changes, noise pollution, occlusion, and the pose changes of the objects themselves. These interferences will cause changes in the texture features of the images, thereby reducing the performance of the object detection model and resulting in situations such as missed detections and false detections. When dealing with these complex situations, traditional object detection methods mostly rely on fixed feature extraction methods and single model architectures, lacking the ability to adaptively process changes in image texture features and being difficult to maintain the robustness of detection under different interference conditions. Therefore, developing a method that can effectively enhance the robustness of object detection and make full use of image texture features has important practical significance and application value. Summary of the Invention
[0003] To solve the problems in the background art, the present invention provides a method for enhancing the robustness of object detection based on texture feature similarity, including:
[0004] S1: Obtain the image of the object to be detected and perform preprocessing;
[0005] S2: Smooth the image of the object to be detected to obtain a smoothed image;
[0006] S3: Input the image of the object to be detected and the smoothed image into two Backbone networks respectively for feature extraction to obtain the multi-level features of the image of the object to be detected and the multi-level features of the smoothed image correspondingly;
[0007] S4: Use an adaptive mechanism to screen out the top L pairs of multi-level features with the largest difference between the image of the object to be detected and the smoothed image;
[0008] S5: Calculate the anomaly detection score of the image of the object to be detected through an attention mechanism according to the difference between the top L pairs of multi-level features with the largest difference between the image of the object to be detected and the smoothed image;
[0009] S6: Determine whether the anomaly detection score of the target image to be detected is greater than the set threshold. If so, discard the sample; otherwise, input the sample into the target detection model for inference to obtain the detection result.
[0010] The present invention has at least the following beneficial effects
[0011] The present invention provides a method for enhancing the robustness of target detection based on texture feature similarity. By preprocessing and smoothing the obtained target image to be detected, it is possible to reduce the influence of interference factors such as noise and uneven illumination in the image, making the subsequent feature extraction process more stable. An adaptive mechanism is used to screen out the top L hierarchical feature pairs with the largest difference between the target image to be detected and the smoothed image. This process can highlight the most representative and discriminative features in the image, avoid the influence of redundant features or features with large interference on the detection result, and improve the accuracy and effectiveness of feature selection. According to the difference between the selected feature pairs, the attention mechanism is used to calculate the anomaly detection score of the target image to be detected, which can quantitatively evaluate the anomaly degree of the target in the image. By judging the relationship between the anomaly detection score and the set threshold, abnormal samples are discarded to avoid the interference of abnormal samples on the inference result of the target detection model, thereby improving the accuracy and reliability of the inference of the target detection model, making the finally output detection result more in line with the actual situation, reducing the occurrence of false detections and missed detections, and enhancing the robustness and practicality of the target detection system in complex environments. Ensure that the target detection system can still maintain high performance in various environments, and has strong practical value and application prospects. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 It is a schematic flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0013] The following specific examples illustrate the embodiments of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0014] Please refer to Figure 1 , the present invention provides a method for enhancing the robustness of target detection based on texture feature similarity, including:
[0015] S1: Obtain the target image to be detected and perform preprocessing;
[0016] Preferably, the preprocessing of the target image to be detected includes: normalizing the target image to be detected, and the formula is as follows:
[0017]
[0018] wherein, represents the pixel value of the i-th pixel after the normalization process of the target image to be detected, and z i represents the pixel value of the i-th pixel in the target image to be detected, and max(z) and min(z) respectively represent the maximum and minimum pixel values in the target image to be detected; adjust the size of the target image to be detected after the normalization process to be the same as the input of the target detection model.
[0019] In this embodiment, the input image is first preprocessed to ensure that it meets the requirements of the target detection system. The preprocessing steps include size adjustment and normalization. The size adjustment adjusts the image to a unified size to meet the input requirements of the target detection model. The target image to be detected includes: face images, vehicle images, flow field images, or plant and animal images, etc., which are images for target detection;
[0020] The methods for obtaining images in this embodiment include:
[0021] When the entity executing the method of the present invention is a computer, the present invention can obtain image data from the outside through interfaces such as RS485, RS232, and USB. At the same time, it can also be downloaded from the network through communication protocols such as 4G and 5G, for example, from browsers such as Baidu and Google. When the entity executing the method of the present invention is a mobile phone, since the mobile phone has a built-in camera function, in addition to the acquisition methods when the entity in S101 is a computer, the present invention can also obtain image data by taking pictures on its own. When the entity executing the method of the present invention is an aerial photography drone, since the aerial photography drone is equipped with a camera and can perform real-time image transmission and shooting, which is a flying vehicle with a photography function, in addition to the acquisition method when the entity in S102 is a mobile phone, the target image data to be detected obtained includes aerial photography pictures taken by the aerial photography drone at high altitude. When the entity executing the method of the present invention is an access control system, since the access control system has the ability to collect face image data, in addition to the acquisition method when the entity in S103 is an aerial photography drone, the target image data to be detected obtained includes the face images of residents collected by the access control system. When the entity executing the method of the present invention is a traffic monitoring system, since the traffic monitoring system obtains a scene video sequence through cameras deployed on the road and has the function of image acquisition, in addition to the acquisition method when the entity in S104 is an access control system, the target image data to be detected obtained includes the traffic scene video sequence collected by the traffic monitoring system. When the target detection model in the present invention is applied to face recognition, the target detection model in the present invention is mainly based on the CNN convolutional neural network, and the target image data obtained mainly includes: face images of various age groups, and face images of people of different age groups in static and moving states. The face images in static and moving states include normal face images, distorted face images, and face images with different degrees of noise.
[0022] S2: Smooth the target image to be detected to obtain a smoothed image;
[0023] In this embodiment, through smoothing processing, such as using common image smoothing algorithms such as mean filtering and Gaussian filtering, the pixel values in the image can be made more continuous and stable. Mean filtering replaces the current pixel value by calculating the average value of neighboring pixels, thereby reducing the influence of noise; Gaussian filtering performs weighted averaging on neighboring pixels according to the Gaussian distribution, and can better retain important features such as the edges of the image while smoothing the image. The smoothed image obtained through smoothing processing not only enables the subsequent Backbone network to be more stable in feature extraction and reduces feature extraction errors caused by noise, but also forms a contrast with the original target image to be detected, providing an image data basis from different perspectives for subsequent screening of key feature pairs, helping to analyze image features more comprehensively and accurately, and thus improving the robustness of target detection.
[0024] S3: Input the target image to be detected and the smoothed image into two Backbone networks respectively for feature extraction, and correspondingly obtain the multi-level features of the target image to be detected and the multi-level features of the smoothed image;
[0025] S4: Use an adaptive mechanism to screen out the top L pairs of level features with the largest difference between the target image to be detected and the smoothed image;
[0026] Preferably, the Backbone network includes: an input layer, M cascaded feature extraction modules, extract the output features of the M feature extraction modules to obtain M-level features of the target image to be detected or the smoothed image; calculate the response intensity between each pair of level features of the target image to be detected and the smoothed image, and select the top L pairs of level features with the largest response intensity.
[0027] In this embodiment, the Backbone network mainly consists of an input layer and M cascaded feature extraction modules. Here, M = 4 is taken as an example for illustration, and in practical applications, the value of M can be adjusted according to requirements. The input layer is used to receive the target image to be detected or the smoothed image, and its number of input channels depends on the type of the image. For example, for a common color image, the number of input channels is 3 (corresponding to the three RGB channels). The input layer contains a convolutional layer that uses a convolutional kernel of size 3×3, a stride of 1, and a padding of 1. The purpose is to perform preliminary feature extraction on the input image, convert the number of channels of the input image to 64 (this value can be adjusted according to the actual situation), and without changing the size of the image. Immediately after the convolutional operation is a batch normalization layer (BatchNorm2d) for normalizing the features output by the convolutional layer and accelerating the training convergence of the network. Finally, a ReLU activation function is passed to introduce non-linearity so that the network can learn more complex features. Feature Extraction Modules The first feature extraction module (layer1): This module is composed of 2 cascaded residual blocks (Residual Block). Each residual block contains two 3×3 convolutional layers inside, and both are followed by a batch normalization layer and a ReLU activation function. The role of the residual block is to avoid the problems of gradient disappearance and gradient explosion through skip connections while deepening the number of layers of the network, so as to extract features more effectively. In the first residual block, the number of input channels is 64 (the same as the output channel number of the input layer), the number of output channels is also 64, and the stride is 1. When the input feature map passes through the first residual block, it will first perform convolution, batch normalization, and activation operations, then add the original input feature map passing through the shortcut connection, and then pass through another activation to get the output. The structure of the second residual block is similar to the first one, with both the input and output channel numbers being 64 and the stride being 1. After passing through layer1, the number of channels of the feature map remains 64, and the size also remains basically unchanged (because the stride is 1 and the padding is appropriate). The second feature extraction module (layer2): It consists of 2 residual blocks. The number of input channels of the first residual block is 64 (the output channel number of layer1), the number of output channels is 128, and the stride is 2. The setting of the stride of 2 makes the size of the feature map halve (both the width and height become half of the original) after passing through this residual block, and at the same time the number of channels increases to 128. The number of input and output channels of the second residual block are both 128, and the stride is 1, further extracting and processing the features. After passing through layer2, the number of channels of the feature map becomes 128, and the size becomes one-fourth of the original (because of one operation with a stride of 2).
[0028] The third feature extraction module (layer3): It also contains 2 residual blocks. The input channel number of the first residual block is 128 (the output channel number of layer2), the output channel number is 256, and the stride is 2, which halves the size of the feature map again and increases the channel number to 256. The input and output channel numbers of the second residual block are both 256, and the stride is 1. After passing through layer3, the channel number of the feature map becomes 256, and the size becomes one-eighth of the original.
[0029] The fourth feature extraction module (layer4): It consists of 2 residual blocks. The input channel number of the first residual block is 256 (the output channel number of layer3), the output channel number is 512, and the stride is 2, which continues to halve the size of the feature map and the channel number becomes 512. The input and output channel numbers of the second residual block are both 512, and the stride is 1. After passing through layer4, the channel number of the feature map becomes 512, and the size becomes one-sixteenth of the original. During the forward propagation of the network, the target image to be detected or the smoothed image first undergoes preliminary processing through the input layer, and then passes through these 4 feature extraction modules in sequence. Each module will output the feature map of the corresponding level. Finally, the output features of these 4 feature extraction modules are extracted, that is, the 4-level features of the target image to be detected or the smoothed image are obtained, which are used to calculate the response intensity between each pair of level features in the subsequent calculation, and the top L pairs of level features with the largest response intensity are selected.
[0030] Preferably, the response intensity between each pair of level features of the target image to be detected and the smoothed image includes:
[0031]
[0032] where Layer m represents the feature difference degree of the m-th level feature pair of the target image to be detected and the smoothed image; C, H, and W respectively represent the number of channels, height, and width; represents the pixel value of the m-th level feature of the target image to be detected at the c-th channel, h-th height, and w-th width; represents the pixel value of the m-th level feature of the smoothed image at the c-th channel, h-th height, and w-th width.
[0033] S5: Calculate the anomaly detection score of the target image to be detected through the attention mechanism according to the difference between the top L level feature pairs with the largest difference between the target image to be detected and the smoothed image;
[0034] Preferably, the calculation of the anomaly detection score of the target image to be detected includes:
[0035] S51: Calculate the covariance matrix of the features at each level of the filtered target image to be detected and the smoothed image, and calculate the overall difference between the features at each level of the filtered target image to be detected and the smoothed image through the Frobenius norm;
[0036] S52: Calculate the cosine similarity and KL divergence between the feature pairs at each level of the filtered target image to be detected and the smoothed image, and obtain the comprehensive similarity between the feature pairs at each level based on the calculated cosine similarity and KL divergence;
[0037] S53: Use the SE attention mechanism module to calculate the fluctuation of the features at each level of the filtered target image to be detected, and calculate the anomaly detection score of the target image to be detected based on the overall difference and comprehensive similarity between the feature pairs at each level of the filtered target image to be detected and the smoothed image, and the fluctuation of the features at each level of the filtered target image to be detected.
[0038] In this embodiment, by calculating the covariance matrix of the features at each level of the filtered target image to be detected and the smoothed image, and using the Frobenius norm to measure the overall difference between them (S51), the change relationship between features can be accurately captured from the statistical property level. The covariance matrix reflects the correlation between different feature dimensions, and the Frobenius norm provides a way to measure the overall "size" of the matrix, which makes the evaluation of the difference between feature pairs more comprehensive and accurate. This quantization method can effectively highlight the deviation degree of the abnormal image from the normal image in the feature distribution, providing important basic data for subsequent anomaly scoring.
[0039] In the process of calculating the cosine similarity and KL divergence between the feature pairs at each level and obtaining the comprehensive similarity based on this (S52), the advantages of the two measurement methods are fully utilized. The cosine similarity can measure the similarity in the direction of the feature vectors, reflecting the consistency of the features in the spatial distribution; the KL divergence focuses on measuring the difference between two probability distributions and can reflect the deviation of the feature distribution in the feature representation. Combining the two can comprehensively evaluate the similarity degree between feature pairs from multiple perspectives, further refining the judgment of abnormal features and improving the accuracy and reliability of the anomaly detection score.
[0040] The SE (Squeeze-and-Excitation) attention mechanism module is used to calculate the fluctuations of each level of features of the screened target image to be detected (S53), which can automatically learn the importance of different levels of features. The SE attention mechanism can highlight key features and suppress noise features by performing compression and excitation operations on the feature channels, thus paying more attention to the feature fluctuations that have important indication effects on anomaly detection. The anomaly detection score is calculated based on the overall difference, comprehensive similarity, and feature fluctuations, so that the score result comprehensively considers various attributes of the features. It can not only identify the numerical differences of the features, but also capture the anomalies in the importance distribution and fluctuations of the features, greatly enhancing the sensitivity and accuracy of the anomaly target detection, and then significantly improving the robustness of the target detection, effectively reducing the false detection and missed detection caused by image interference in complex environments.
[0041] Preferably, the calculation of the overall difference between each level of feature pairs of the screened target image to be detected and the smoothed image includes:
[0042]
[0043] Among them, represents the overall difference between the l-th level feature pair of the screened target image to be detected and the smoothed image; and respectively represent the channel covariance matrices of the l-th level features of the screened target image to be detected and the smoothed image in the i-th and j-th channels; C l represents the covariance matrix of the l-th level feature of the screened target image to be detected; F l represents the l-th level feature of the screened target image to be detected; μ l represents the mean of the l-th level feature of the screened target image to be detected; T represents transpose, ‖.‖ F represents the Frobenius norm, which is used to calculate the overall difference of the matrix.
[0044] Preferably, the comprehensive similarity between each level of feature pairs of the target image to be detected and the smoothed image includes:
[0045]
[0046]
[0047] Among them, represents the comprehensive similarity between the l-th level feature pair of the screened target image to be detected and the smoothed image; α and β represent weight parameters; represents the cosine similarity between the l-th level feature pair of the screened target image to be detected and the smoothed image; represents the KL divergence between the screened target image to be detected and the l - th level feature pair of the smoothed image; represents the pixel value of the l - th level feature of the screened target image to be detected at the c - th channel, h - th height, and w - th width; represents the pixel value of the l - th level feature of the screened smoothed image at the c - th channel, h - th height, and w - th width; log represents the natural logarithm, is for to perform normalization processing, is the normalization processing for .
[0048] Preferably, the anomaly detection score of the target image to be detected includes:
[0049]
[0050] where S final represents the anomaly detection score of the target image to be detected, α l represents the weight parameter; γ, λ, and η represent weight parameters; represents the overall difference between the screened target image to be detected and the l - th level feature pair of the smoothed image; represents the fluctuation of the l - th level feature of the target image to be detected; σ represents the sigmoid function; W1 and W2 are the weights of the fully - connected layer; δ represents the ReLU activation function; C, H, and W respectively represent the number of channels, height, and width; μ is the mean of represents the pixel value of the l - th level feature of the screened target image to be detected at the c - th channel, h - th height, and w - th width.
[0051] S6: Determine whether the anomaly detection score of the target image to be detected is greater than the set threshold. If so, discard the sample; otherwise, input the sample into the target detection model for inference to obtain the detection result.
[0052] Preferably, use the binary cross - entropy (BCE) loss to distinguish the normal sample label y = 0 and the abnormal sample label y = 1, and update the weights W1 and W2 of the fully - connected layer of the SE module and the weight parameters γ, λ, and η through the Adam optimizer. The loss function Loss is set as follows:
[0053] Loss=-[ylog(S final )+(1 - y)log(1 - S final )]
[0054] where, log represents the natural logarithm; y represents the sample label; S final represents the anomaly detection score of the sample.
[0055] Preferably, the set threshold includes: calculating the mean of the anomaly detection scores of all normal samples in the training set corresponding to the target detection model as the set threshold. Samples determined to be normal based on the anomaly scores will continue to be input into the target detection system for inference. This system can adapt to any type of target detection model (such as YOLO, Faster R-CNN, SSD, etc.) and can adjust the model prediction strategy in combination with the anomaly detection results to improve the robustness of target detection.
[0056] In summary, the present invention provides a method for enhancing the robustness of target detection based on texture feature similarity. By preprocessing and smoothing the acquired target image to be detected, it is possible to reduce the influence of interference factors such as noise and uneven illumination in the image, making the subsequent feature extraction process more stable. An adaptive mechanism is used to select the top L hierarchical feature pairs with the largest difference between the target image to be detected and the smoothed image. This process can highlight the most representative and discriminative features in the image, avoiding the influence of redundant features or features with large interference on the detection results, and improving the accuracy and effectiveness of feature selection. According to the difference between the selected feature pairs, the attention mechanism is used to calculate the anomaly detection score of the target image to be detected, which can quantitatively evaluate the anomaly degree of the target in the image. By judging the relationship between the anomaly detection score and the set threshold, abnormal samples are discarded to avoid the interference of abnormal samples on the inference results of the target detection model, thereby improving the accuracy and reliability of the inference of the target detection model, making the finally output detection results more in line with the actual situation, reducing the occurrence of false detections and missed detections, and enhancing the robustness and practicality of the target detection system in complex environments. Ensure that the target detection system can still maintain high performance in various environments, with strong practical value and application prospects.
[0057] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the purpose and scope of the present technical solution, and they should all be covered by the scope of the claims of the present invention.
Claims
1. A method for enhancing the robustness of object detection based on texture feature similarity, characterized in that Including: S1: Obtain the target image to be detected and perform preprocessing; S2: Smooth the target image to be detected to obtain a smoothed image; S3: Input the target image to be detected and the smoothed image into two Backbone networks respectively for feature extraction to obtain the multi-level features of the target image to be detected and the multi-level features of the smoothed image correspondingly; S4: Use an adaptive mechanism to screen out the top L pairs of hierarchical features with the largest difference between the target image to be detected and the smoothed image; S5: Calculate the anomaly detection score of the target image to be detected through an attention mechanism based on the difference between the top L pairs of hierarchical features with the largest difference between the target image to be detected and the smoothed image; S6: Determine whether the anomaly detection score of the target image to be detected is greater than the set threshold. If so, discard the sample; otherwise, input the sample into the target detection model for inference to obtain the detection result.
2. The robust enhancement method for object detection based on texture feature similarity according to claim 1, wherein The preprocessing of the target image to be detected includes: normalizing the target image to be detected, and the formula is as follows: Among them, represents the pixel value of the i-th pixel after the normalization process of the target image to be detected, z i represents the pixel value of the i-th pixel in the target image to be detected, max(z) and min(z) respectively represent the maximum and minimum pixel values in the target image to be detected; the size of the target image to be detected after the normalization process is adjusted to be the same as the input of the target detection model.
3. A method for enhancing the robustness of object detection based on texture feature similarity according to claim 1, wherein The Backbone network includes: an input layer, M cascaded feature extraction modules, extract the output features of the M feature extraction modules to obtain M-level features of the target image to be detected or the smoothed image; calculate the response intensity between each pair of hierarchical features of the target image to be detected and the smoothed image, and select the top L pairs of hierarchical features with the largest response intensity.
4. A method for enhancing the robustness of object detection based on texture feature similarity according to claim 3, characterized in that The response intensity between each pair of hierarchical features of the target image to be detected and the smoothed image includes: Among them, Layer m represents the feature difference degree of the feature pair of the target image to be detected and the smoothed image at the m-th layer level; C, H, and W respectively represent the number of channels, height, and width; represents the pixel value of the m-th layer level feature of the target image to be detected at the c-th channel, h-th height, and w-th width; represents the pixel value of the m-th layer level feature of the smoothed image at the c-th channel, h-th height, and w-th width.
5. A method for enhancing the robustness of object detection based on texture feature similarity according to claim 1, characterized in that, The calculation of the anomaly detection score of the target image to be detected includes: S51: Calculate the covariance matrix of each hierarchical feature of the selected target image to be detected and the smoothed image, and calculate the overall difference between each pair of hierarchical features of the selected target image to be detected and the smoothed image through the Frobenius norm; S52: Calculate the cosine similarity and KL divergence between each pair of hierarchical features of the selected target image to be detected and the smoothed image, and obtain the comprehensive similarity between each pair of hierarchical features based on the calculated cosine similarity and KL divergence; S53: Use the SE attention mechanism module to calculate the fluctuation of each hierarchical feature of the selected target image to be detected, and calculate the anomaly detection score of the target image to be detected based on the overall difference and comprehensive similarity between each pair of hierarchical features of the selected target image to be detected and the smoothed image, and the fluctuation of each hierarchical feature of the selected target image to be detected.
6. The robust enhancement method for object detection based on texture feature similarity according to claim 5, wherein, The calculation of the overall difference between each pair of hierarchical features of the selected target image to be detected and the smoothed image includes: Among them, represents the overall difference between the l-th level feature pairs of the screened target image to be detected and the smoothed image; and respectively represent the channel covariance matrices of the l-th level features of the screened target image to be detected and the smoothed image in the i-th and j-th channels; C l represents the covariance matrix of the l-th level feature of the screened target image to be detected; F l represents the l-th level feature of the screened target image to be detected; μ l represents the mean value of the l-th level feature of the screened target image to be detected; T represents transpose, ‖.‖ F represents the Frobenius norm, which is used to calculate the overall difference of the matrix.
7. A method for enhancing the robustness of object detection based on texture feature similarity according to claim 5, characterized in that The comprehensive similarity between each pair of hierarchical features of the target image to be detected and the smoothed image includes: Among them, represents the comprehensive similarity between the filtered target image to be detected and the l-th level feature pair of the smoothed image; α and β represent weight parameters; represents the cosine similarity between the filtered target image to be detected and the l-th level feature pair of the smoothed image; represents the KL divergence between the filtered target image to be detected and the l-th level feature pair of the smoothed image; represents the pixel value of the l-th level feature of the filtered target image to be detected at the c-th channel, h-th height, and w-th width; represents the pixel value of the l-th level feature of the filtered smoothed image at the c-th channel, h-th height, and w-th width; log represents the natural logarithm, is subject to normalization processing, is the normalization processing of .
8. A method for enhancing the robustness of object detection based on texture feature similarity according to claim 5, characterized in that The anomaly detection score of the target image to be detected includes: Among them, S final represents the anomaly detection score of the target image to be detected, and α l represents a weight parameter; γ, λ, and η represent weight parameters; represents the overall difference between the l-th level features of the selected target image to be detected and the smoothed image; represents the fluctuation of the l-th level feature of the target image to be detected; σ represents the sigmoid function; W1 and W2 are the weights of the fully connected layer; δ represents the ReLU activation function; C, H, and W represent the number of channels, height, and width respectively; μ is the mean value of; represents the pixel value of the l-th level feature of the selected target image to be detected at the c-th channel, h-th height, and w-th width.
9. A method for enhancing the robustness of object detection based on texture feature similarity according to claim 1, characterized in that, Use binary cross-entropy BCE loss to distinguish the normal sample label y = 0 and the abnormal sample label y = 1, update the weights W1 and W2 of the fully connected layer of the SE module and the weight parameters γ, λ, and η through the Adam optimizer, and set its loss function Loss as follows: Loss=-[ylog(S final )+(1-y)log(1-S final )] Among them, log represents the natural logarithm; y represents the sample label; S final represents the anomaly detection score of the sample.
10. A method for enhancing the robustness of object detection based on texture feature similarity according to claim 1, characterized in that, The set threshold includes: calculating the mean value of the anomaly detection scores of all normal samples in the training set corresponding to the target detection model as the set threshold.