Visual target detection method and device used in strong noise interference environment
By replacing modules in the YOLOv8 backbone network and introducing soft thresholding processing, designing attention mechanism modules, and optimizing the YOLOv8 network, the accuracy and speed imbalance of object detection in strong noise environments are solved, high-precision rapid detection and model lightweighting are achieved, and real-time object detection is suitable for complex scenarios.
Patent Information
- Application Number
- CN202510516350.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-07-18
AI Technical Summary
The prior art object detection method has problems with imbalance in detection accuracy and speed in a strong noise interference environment, and is difficult to be applied to different types and levels of noise interference environments. The model is complex and difficult to transplant into embedded devices for real-time detection.
Replace the C2f module in YOLOv8's backbone network as the deep residual shrinkage module DRSB, introduce soft thresholding processing, design attention mechanism module WSW, optimize feature extraction and information interaction, and establish a high-precision and rapid detection model.
It realizes high-precision and rapid detection in strong noise environments, reduces calculation complexity, is suitable for different types and levels of noise interference environments, and is suitable for transplanting to embedded devices for real-time object detection.
Smart Images

Figure CN120339644A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of target detection, and particularly relates to a visual target detection method and device for a strong noise interference environment. Background Art
[0002] Vision-based target detection is a research topic in interdisciplinary fields such as image processing, computer vision, and pattern recognition. It is widely used in fields such as video surveillance, intelligent transportation, human-computer interaction, and satellite navigation, and has extremely important theoretical significance and practical value. Target detection aims to separate the background and targets from image or video data, accurately and efficiently locate and discriminate object instances, and further obtain basic information such as the position, pose, and trajectory of the targets. Therefore, target detection is the premise for perceiving and controlling target objects, and also the basis for high-level understanding and application. Its performance directly affects subsequent high-level tasks such as target tracking, action recognition, and behavior understanding.
[0003] With the rapid development of deep learning, convolutional neural networks have achieved remarkable success in image recognition and are widely used in the field of target detection. Compared with traditional computer vision-based methods, using a deep learning framework to perform self-learning on samples and automatically extract features has greatly improved the accuracy and efficiency of target detection. Currently, there are still key technical problems in deep learning-based target detection, such as optimizing the performance of mainstream target detection algorithms, making full use of the context information of the environmental background, improving the detection accuracy of small targets, and lightweighting the detection model.
[0004] In practical applications, the spatial environment has factors such as image noise, motion blur, illumination change, object scale, and complex background, and image or sensor data is often interfered by different levels of noise. Currently, the main methods for target detection in a noise interference environment are as follows: suppressing static noise through polynomial fitting denoising methods; removing the background and part of the interference through high-pass filtering, and then using image enhancement techniques to further suppress noise interference; modeling the complex background through the structure of an encoder and a decoder, and achieving background suppression through the background residual cancellation method; extracting feature maps from a denoising diffusion probability model to improve the target detection ability of the YOLO model (a deep learning-based target detection algorithm, whose core idea is to transform the target detection problem into a regression problem, and use a neural network to directly predict the bounding box coordinates and target categories from image pixels to achieve real-time target detection). However, the application scenarios of the above methods are relatively single, and the versatility of the detection model is poor, resulting in limited promotion to practical applications. At the same time, it is difficult to achieve accurate and efficient detection of moving targets in a strong noise interference environment, which also greatly restricts the stability and robustness of the detection effect.
[0005] The prior art has improved single-stage and two-stage object detection algorithms, further enhancing detection accuracy. However, it has poor anti-interference ability in noisy environments, and it is difficult to ensure real-time detection after transplanting the model complexity to embedded devices, resulting in a great consumption of detection efficiency. The prior art learns the characteristics of specific noises based on filters or deep learning frameworks and eliminates redundant information. However, the formation of noises is diverse, making it difficult to form a general denoising method. The prior art captures global feature information, breaks through the limitations of long-distance modeling, and achieves efficient parallel computing. However, due to the characteristics of global computing, it will increase the computational complexity, with longer training and inference times, restricting the application of long-sequence tasks and there is also a risk of model overfitting.
[0006] In summary, in an environment with strong noise interference, the object detection method based on a deep learning framework has problems in balancing detection accuracy and detection speed. At the same time, it is necessary to solve the problem of how to make the object detection method applicable to different types and levels of noise interference environments and transplant it into embedded devices to achieve accurate and real-time object detection in multiple extreme scenarios. Summary of the Invention
[0007] To solve the above technical problems, the present invention adopts the following technical solutions:
[0008] A visual object detection method for an environment with strong noise interference, comprising:
[0009] Step 1: Add a Deep Residual Shrinkage Block (DRSB); in the backbone network of YOLOv8, replace the C2f module with the Deep Residual Shrinkage Block (DRSB) to improve the original backbone network of YOLOv8.
[0010] Step 2: Introduce soft thresholding processing; effectively eliminate noise features and improve the feature extraction process of the original backbone network of YOLOv8.
[0011] Step 3: Design an attention mechanism module WSW; before the SPPF module in the backbone network of YOLOv8, design an attention mechanism module WSW, and use the operation of sliding windows to calculate multi-head self-attention within and between the windows of the feature map in turn to improve the interactivity of feature information and improve the original backbone network of YOLOv8.
[0012] Step 4: Establish a high-precision and fast detection model for objects in a spatial strong noise interference environment; connect the improved backbone network of YOLOv8 with the feature fusion network and detection head of the YOLOv8 network to establish a complete model, and perform training and testing to verify the effectiveness and generality of the visual object detection method for an environment with strong noise interference.
[0013] A visual target detection device for a strong noise interference environment, comprising:
[0014] A DRSB addition module of a deep residual shrinkage module. In the backbone network of YOLOv8, the C2f module is replaced with the deep residual shrinkage module DRSB to improve the original backbone network of YOLOv8.
[0015] A soft thresholding processing introduction module introduces soft thresholding processing to effectively eliminate noise features and improve the feature extraction process of the original backbone network of YOLOv8.
[0016] An attention mechanism module WSW design module designs the attention mechanism module WSW. Before the SPPF module in the backbone network of YOLOv8, the attention mechanism module WSW is designed. By using the operation of a sliding window, the multi-head self-attention is calculated within and between the windows of the feature map in turn to improve the interaction of feature information and improve the original backbone network of YOLOv8.
[0017] A detection model establishment module establishes a high-precision and fast detection model for targets in a strong noise interference environment in space. The improved backbone network of YOLOv8 is connected to the feature fusion network and the detection head of the YOLOv8 network to establish a complete model, and training and testing are carried out to verify the effectiveness and generality of the visual target detection method for a strong noise interference environment.
[0018] An electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the visual target detection method for a strong noise interference environment described above are implemented.
[0019] A non-transitory computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the visual target detection method for a strong noise interference environment described above are implemented.
[0020] The present invention has the following beneficial effects:
[0021] The method of the present invention relies on the backbone network, feature fusion network, and detection head of YOLOv8. It improves and optimizes its backbone network, replaces the C2f module with the deep residual shrinkage module DRSB, incorporates the denoising idea of soft thresholding, adaptively learns the thresholds of image features, adjusts the weights of each feature, excludes the interference of noise, and enhances the feature extraction ability. In addition, an attention mechanism module WSW is designed during the extraction of features by the backbone network to reduce the computational complexity. By dividing the feature map into several small windows and calculating the local self-attention, and using the sliding window to promote the information exchange and fusion between windows, the expression ability and global perception ability of the model are improved, so as to ensure that subsequent operations such as fusion, classification, and regression are based on effective image features.
[0022] The present invention aims at object detection in a strong noise environment, achieving a balance between detection accuracy and speed, while improving the versatility of the detection method, and can be widely applied to different types and levels of noise interference environments. The present invention can effectively reduce the computational complexity of the model, greatly saving the training and inference time; has high detection accuracy under strong noise interference, and also has the characteristics of fast detection; can effectively denoise by adaptively learning the characteristics of different noises; the present invention is suitable for transplantation into embedded devices to achieve real-time object detection in complex scenarios, and can be extended to application scenarios such as autonomous navigation, in-air refueling, and industrial sites in extreme environments, having important research significance and application value. Brief Description of the Drawings
[0023] Figure 1 is a flowchart of the visual object detection method for a strong noise interference environment of the present invention;
[0024] Figure 2 is a schematic diagram of the residual structure in the deep residual shrinkage module DRSB;
[0025] Figure 3 is a schematic diagram of the deep residual shrinkage module DRSB;
[0026] Figure 4 is a schematic diagram of the window operation of the attention mechanism module WSW. Among them, (a) is a schematic diagram of the window division operation, (b) is a schematic diagram of the sliding window operation, and (c) is a schematic diagram of the window movement and merging operation;
[0027] Figure 5Dataset diagram of the present invention, where (a) is a picture with mild noise simulating scenarios such as rain and snow, sand and dust, and dust; (b) is a picture with mild noise simulating scenarios such as motion blur; (c) is a picture with mild noise simulating scenarios such as light changes; (d) is a picture with moderate noise simulating scenarios such as rain and snow, sand and dust, and dust; (e) is a picture with moderate noise simulating scenarios such as motion blur; (f) is a picture with moderate noise simulating scenarios such as light changes; (g) is a picture with severe noise simulating scenarios such as rain and snow, sand and dust, and dust; (h) is a picture with severe noise simulating scenarios such as motion blur; (i) is a picture with severe noise simulating scenarios such as light changes. Detailed implementation manners
[0028] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0029] The present invention aims to establish a high-precision and fast detection model for targets in a strong spatial noise interference environment, realize a visual target detection method for a strong noise interference environment, and conduct experiments and tests on this model to verify the effectiveness and generality of the method of the present invention.
[0030] The basic idea of the present invention is as follows: relying on the basic framework of the backbone network, feature fusion network, and detection head of the YOLOv8 network (a target detection algorithm based on a deep convolutional neural network and belonging to a model in the YOLO (You Only Look Once) series), the backbone network is improved and optimized. First, in the backbone network of YOLOv8, the C2f module (a residual structure that realizes cross-stage partial fusion with 2 convolutional layers) is replaced with the deep residual shrinkage module DRSB. In the DRSB, the residual structure with a cross-layer identity path can alleviate the problems of gradient disappearance and gradient explosion, thereby increasing the network depth and facilitating the learning of more abstract feature information. Second, soft thresholding processing is introduced in the DRSB. A small network adaptively learns the threshold of the input features, and the features are soft-thresholded using the threshold to re-shrink the weight magnitudes of the features, effectively removing redundant features such as noise and ensuring the efficient transmission of useful features of the image. Third, before the pooling operation in the SPPF module (fast spatial pyramid pooling layer) in the backbone network of YOLOv8, an attention mechanism module WSW is designed. The feature map is divided into windows of a fixed size, and a sliding window operation is adopted to calculate multi-head self-attention within and between the windows. Using window calculation instead of long sequence calculation greatly reduces the computational complexity of the model, enhances the interaction of feature information between windows, captures the context information in the image, breaks through the constraints of local information, and improves the efficiency of target detection. Then, the improved backbone network is connected to the feature fusion network and detection head of YOLOv8 to establish a high-precision and fast detection model for targets in a strong spatial noise interference environment. A real and effective strong noise dataset is prepared, the model is trained and tested, various indicators are evaluated, and the performance is further optimized. Finally, the model is transplanted to a suitable embedded device and applied to real complex scenarios to verify the feasibility, effectiveness, and generality of the method.
[0031] The present invention can be used for target detection in a noise-free interference environment and noise environments of different types and levels. Taking the visual target detection method for a strong noise interference environment as an example, the present invention is further described in detail below. The present invention mainly studies 3 types of strong noise, namely image noise, motion blur noise, and brightness noise, and the corresponding complex scenarios are target detection in harsh weather such as rain and snow, sand and dust, detecting fast-moving targets, and the brightness changes generated by all-weather supervised targets, respectively.
[0032] As Figure 1 shown, the visual target detection method for a strong noise interference environment of the present invention includes:
[0033] Step 1: Add the Deep Residual Shrinkage Block (DRSB). In the backbone network of YOLOv8, replace the C2f module with the Deep Residual Shrinkage Block (DRSB) to improve the original backbone network of YOLOv8. In many feature learning tasks, the samples will contain some irrelevant interference information such as noise, which will affect the effect of feature learning. The Deep Residual Shrinkage Block (DRSB) is a further improvement of the residual structure. By adding a small sub-network inside the residual structure, a set of thresholds can be learned for different features. After subsequent soft-thresholding operations, the weights of all input features can be effectively adjusted to eliminate noise features. In the residual structure as shown in Figure 2 The size of the input feature map is , is the number of channels of the input feature map, is the width of the input feature map, is the height of the input feature map. After the input feature map passes through operation ① convolution, normalization, and activation function, the size of the output feature map is . After passing through operation ② convolution, normalization, and activation function, the output feature map is added to the input feature map after operation ③ to obtain the final output feature map. The residual structure effectively alleviates the training difficulty of deep networks through the cross-layer identity path. Using this operation, a deeper network model can be designed. As shown in Figure 3 , the Deep Residual Shrinkage Block is based on the residual structure and replaces the activation function in operation ② with a feature shrinkage operation, that is, after convolution and normalization in operation ②, calculate the absolute value through operation ③, ④ global average pooling, ⑤ fully connected layer, ⑥ normalization, activation function, ⑦ fully connected layer, ⑧ Sigmoid function, ⑨ multiplication, ⑩ soft-thresholding to achieve feature shrinkage. The feature map is reduced to a one-dimensional vector by operation ④, and the average value of the absolute values of the features of each channel is obtained from . After operation ⑤, neurons are obtained. Then, operations ⑥, ⑦, and ⑧ are performed on each neuron to obtain coefficients . Each coefficient is independent and does not affect each other, and is located in the interval . Multiply and one by one to get , forming a set of adaptively learned thresholds . The calculation formula is as follows. The threshold of each channel feature is a positive number and less than or equal to .
[0034] (1)
[0035] Using a threshold Perform a soft thresholding operation on the input features to readjust the importance of each feature, further removing redundant features with low importance such as noise, thus achieving feature contraction. Finally, a complete deep residual shrinkage module DRSB is implemented through a cross-layer identity path. It can be seen that feature contraction plays the role of a non-linear transformation, and the core operation is soft thresholding.
[0036] Step 2: Introduce soft thresholding processing. Soft thresholding is the core of signal denoising. It Figure 3 removes the features in each channel of the feature map output by operation ③ in whose absolute values are less than the corresponding threshold and shrinks the features whose absolute values are greater than this threshold towards zero, readjusting the weight of each feature , thereby improving the feature extraction process of the backbone network of the original YOLOv8, effectively removing noise features. The function expression and derivative expression of soft thresholding are as follows:
[0037] (2)
[0038] (3)
[0039] Among them, is Figure 3 the feature in each channel of the feature map output by operation ② in is the corresponding adaptively learned threshold, is the adjusted feature weight, . It can be seen that the derivative of the soft thresholding function has only two values, 0 or 1. Therefore, the problem of gradient disappearance or gradient explosion in deep learning models can be avoided. The thresholds for different samples are generally different. Through Figure 3 operations ③, ④, ⑤, ⑥, ⑦, ⑧, ⑨ in the sub-network can adaptively learn the thresholds of all features , where the absolute values of corresponding to features with low importance such as noise are small. After the soft thresholding operation, their weights are set to 0, thereby achieving the removal of redundant information and only retaining the useful features of the detection target, effectively improving the accuracy of target detection under strong noise.
[0040] Step 3: Design the attention mechanism module WSW. Before the SPPF module in the backbone network of YOLOv8, design the attention mechanism module WSW to improve the original backbone network of YOLOv8 and further enhance the feature extraction ability of the backbone network. In WSW, first calculate the multi-head self-attention within the window, and then after the sliding window operation, calculate the multi-head self-attention between the windows. The multi-head self-attention is obtained by calculating the single-head self-attention and then splicing and linearly transforming the output of the single-head self-attention. Specifically, it includes:
[0041] Step 3.1, in the attention mechanism module WSW, calculate the multi-head self-attention within the window. Figure 4 is a schematic diagram of the window operation of the attention mechanism module WSW. Divide the input feature map into multiple windows of a fixed size. Taking an 8×8 window as an example, as Figure 4 shown in (a) of, divide the input feature map into 4 8×8 windows, and calculate the multi-head self-attention separately within each window. Specifically, it includes:
[0042] Step 3.11, calculate the query , key , and value matrices of the input feature map of the window. Linearly transform the feature matrix in each window to obtain , , and , representing the matrices of the query, key, and value respectively. , , and The elements in the matrix come from different positions of the input feature map, that is, , , and are linearly transformed based on the same input feature matrix , and the calculation formula is as follows:
[0043] (4)
[0044] where , , and are trainable parameter matrices, is used to calculate the weight of attention, is used to interact with the query matrix to jointly determine the degree of relevance, decides the information representation after attention weighting. , , and project the input feature matrix into different subspaces respectively, and obtain , , and 。 , and are obtained through model training and optimized by backpropagation and gradient descent during the training process, enabling the self-attention mechanism to capture the interactions between different features.
[0045] Step 3.12, calculate the self-attention score matrix . By calculating the dot product of the query and key matrices , obtain the similarity matrix of the input features. The elements in the similarity matrix represent the similarity magnitudes between different positions of the input feature map. Divide each element in the similarity matrix by for scaling to obtain the self-attention score matrix , and the calculation formula is as follows:
[0046] (5)
[0047] where is the dimension of the key matrix, that is, the dimension of multiple parallel and independent single-head self-attention calculation units in the multi-head self-attention. The scaling operation ensures the stability of gradient update when training the model.
[0048] Step 3.13, calculate the weight matrix of the self-attention scores . Add the self-attention score matrix obtained in Step 3.12 to the relative position bias matrix . The matrix is constructed according to the relative position indices of the input features in the window and is used to introduce relative position information to solve the problem of absolute position encoding when processing sequence data. Perform , sum of (normalized exponential function) normalization operation so that each element after normalization is within the interval to obtain the weight matrix of the self-attention scores. The larger the numerical value of the elements in the weight matrix of the self-attention scores, the higher the correlation degree between the corresponding two positions in the input feature map of the window. The calculation formula of the weight matrix of the self-attention scores is as follows:
[0049] (6)
[0050] Step 3.14, calculate the output Output of the single-head self-attention. Use the weight matrix of the self-attention scores to perform weighted summation on the value matrix , that is, multiply the weight matrix of the self-attention scores by Multiply to obtain the output Output of the final single-head self-attention. The calculation formula is as follows:
[0051] (7)
[0052] Step 3.15, calculate the output of the multi-head self-attention. For the input feature matrix Define multiple groups of different parameter matrices. At the same time, calculate multiple groups of query, key, and value matrices, self-attention score matrices, and their weight matrices in parallel. Concatenate the outputs of multiple single-head self-attention within each window, that is, the Output obtained in Step 3.14, to obtain the output of the multi-head self-attention within each window, and perform a linear transformation to keep the dimension of the output of the multi-head self-attention the same as the dimension of the input feature map of each window.
[0053] Step 3.2, in the attention mechanism module WSW, calculate the multi-head self-attention between windows. Specifically, it includes:
[0054] Step 3.21, slide the windows divided from the input feature map in Step 3.1. As shown in (b) of Figure 4 , slide the windows 4 pixels to the right and down from the upper left corner respectively. After the sliding window operation, the number of windows changes from 4 to 9, and the 8×8 window in the middle contains the feature information of the 4 windows in (a) of Figure 4 , realizing the information interaction between windows and promoting the transmission of local information.
[0055] Step 3.22, move and merge the newly generated 9 windows to form 4 windows again, as shown in (c) of Figure 4 . Through the operations of sliding windows, window moving, and window merging, on the one hand, the information of the windows divided in Step 3.1 is made to interact, and on the other hand, the computational amount of 4 windows is still maintained without increasing the computational complexity.
[0056] Step 3.23, based on the window merging operation in Step 3.22, calculate the multi-head self-attention separately in each window. Repeat Step 3.11 to calculate the query, key, and value matrices of the input feature map of the window, and Step 3.12 to calculate the self-attention score matrix.
[0057] Step 3.24, due to the window merging operation in Step 3.22, there is information crossover between different windows. After obtaining the self-attention score matrix, it cannot be directly operated according to Step 3.13, otherwise information chaos will occur. Therefore, an additional mask matrix needs to be set to isolate information in different regions. Since the window area distributions are different, the mask formats are also different. The setting of the mask matrix needs to meet the following conditions: after adding the self-attention score matrix, the relative position bias matrix, and the mask matrix, the product of elements from different windows can be set to 0 after normalization. Using this condition, after calculating the self-attention score matrix according to Step 3.12 , the mask matrix can be solved inversely.
[0058] Step 3.25, based on Step 3.24, perform normalization on the sum of the self-attention score matrix, the relative position bias matrix, and the mask matrix to obtain the weight matrix of the self-attention scores.
[0059] Step 3.26, repeat Step 3.14 to calculate the output of single-head self-attention and Step 3.15 to calculate the output of multi-head self-attention to obtain the multi-head self-attention between windows. After the calculation is completed, the windows need to be restored to their original positions.
[0060] Step 4: Establish a high-precision and fast detection model for targets in a strong spatial noise interference environment. After improving the backbone network of YOLOv8 according to Steps 1 to 3, connect it to the feature fusion network and detection head of the YOLOv8 network to establish a complete high-precision and fast detection model for targets in a strong spatial noise interference environment, and achieve target detection in a strong noise interference environment. Input appropriate datasets into this model for experiments and tests. The experimental parameters of the present invention are set as shown in Table 1, and the datasets of the present invention are as Figure 5 shown. Among them, Figure 5 the (a) of Figure 5 the (d) of Figure 5 are pictures simulating scenes such as rain, snow, sand, and dust; Figure 5 the (b) of Figure 5 the (e) of Figure 5 are pictures simulating scenes such as motion blur; Figure 5 the (c) of Figure 5 the (f) of Figure 5 are pictures simulating scenes such as light changes. Figure 5 the (a) of Figure 5 the (b) of Figure 5 the (c) are pictures with mild noise, Figure 5 the (d) of Figure 5 the (e) of Figure 5The picture of (f) is a picture with moderate noise, Figure 5 the picture of (g), Figure 5 the picture of (h), Figure 5 and the picture of (i) is a picture with severe noise. The present invention mainly studies three types of strong noises, namely image noise, motion blur noise, and brightness noise. The settings of their noise ratios, convolution kernel sizes, and brightness factors are shown in Table 2.
[0061] Table 1
[0062] Table 2
[0063] As shown in Table 3, the method of the present invention is compared with the object detection method based on YOLOv8. Under the condition of strong noise, it is trained for 150 rounds. The precision rate of the method of the present invention is 0.022 higher than that of the object detection method based on YOLOv8, and the recall rate and mAP50 are 0.007 and 0.005 lower than those of the object detection method based on YOLOv8. The training time consumption of the present invention is 30.729% of that of the object detection method based on YOLOv8. It can be seen that the method of the present invention not only has high detection accuracy for object detection in a strong noise environment, but also greatly saves training time. By inferring a same 1224-frame video for both methods, it can be known that the inference speed of the method of the present invention is faster than that of the object detection method based on YOLOv8, and it achieves an ideal lightweight, which is suitable for running in embedded devices and resource-constrained environments.
[0064] Table 3
[0065] In summary, the present invention effectively balances the accuracy and speed of object detection under strong noise, can effectively remove various noises, is no longer limited to a single noise, and realizes the lightweight of the object detection method, and can be promoted to complex application scenarios such as industrial sites, autonomous navigation, and aerospace by using embedded technology.
[0066] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes. The solutions in the embodiments of the present invention can be implemented in various computer languages. For example, object-oriented programming languages such as Java and interpreted scripting languages such as JavaScript.
[0067] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing device generate means for implementing the specified functions in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0068] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the specified functions in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0069] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the specified functions in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0070] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments and all changes and modifications that fall within the scope of the present invention.
[0071] Obviously, those skilled in the art can make various changes and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. A visual target detection method for a strong noise interference environment, characterized in that Including: Step 1: Add the Deep Residual Shrinkage Block (DRSB); In the backbone network of YOLOv8, replace the C2f module with the Deep Residual Shrinkage Block (DRSB) to improve the original backbone network of YOLOv8. Step 2: Introduce soft thresholding processing; Effectively eliminate noise features and improve the feature extraction process of the original backbone network of YOLOv8. Step 3: Design the Window Self-Attention Mechanism (WSW) module; Before the SPPF module in the backbone network of YOLOv8, design the Window Self-Attention Mechanism (WSW) module. By using the sliding window operation, calculate the multi-head self-attention within and between the windows of the feature map in turn to improve the interaction of feature information and improve the original backbone network of YOLOv8. Step 4: Establish a high-precision and fast detection model for targets in a spatial strong noise interference environment; Connect the improved backbone network of YOLOv8 with the feature fusion network and detection head of the YOLOv8 network to establish a complete model, and perform training and testing to verify the effectiveness and generality of the visual target detection method for strong noise interference environments.
2. The visual target detection method for a strong noise interference environment according to claim 1, characterized in that, In step 1, the deep residual shrinkage module DRSB is a further improvement of the residual structure, replacing the activation function in operation ② with a feature shrinkage operation: the size of the input feature map is , is the number of channels of the input feature map, is the width of the input feature map, is the height of the input feature map; after the input feature map undergoes operations ① convolution, normalization, and activation function, the size of the output feature map is , and then after convolution and normalization in operation ②, the absolute value is calculated in operation ③, ④ global average pooling, ⑤ fully connected layer, ⑥ normalization and activation function, ⑦ fully connected layer, ⑧ sigmoid function, ⑨ multiplication, ⑩ soft thresholding to achieve feature shrinkage; the feature map is reduced to a one-dimensional vector by operation ④ ,from Get the average value of the absolute value of each channel feature , and then through operation ⑤ to obtain neurons, and then perform operations ⑥, ⑦, and ⑧ on each neuron to obtain Coefficient , each coefficient is independent of each other, does not affect each other, and is located in the interval within; will and Multiply one by one to get , forming a set of adaptive learning thresholds , the threshold of each channel feature All are positive and less than or equal to The threshold The calculation formula is as follows: (1) Using a threshold Perform a soft thresholding operation on the input features to readjust the importance of each feature, thus achieving feature shrinkage. Finally, a complete deep residual shrinkage block DRSB is realized through a cross-layer identity path.
3. The visual target detection method for a strong noise interference environment according to claim 2, characterized in that: In Step 2, soft thresholding is the key to feature shrinkage implementation: for the features of each channel in the feature map output by Operation ③ whose absolute values are less than the corresponding threshold are removed, and the features with absolute values greater than the threshold are shrunk towards zero to readjust the weight of each feature , thereby improving the feature extraction process of the backbone network of the original YOLOv8, effectively removing noise features. The function expression and its derivative expression of soft thresholding are as follows: (2) (3) Among them, is the feature of each channel in the feature map output by operation ②, is the threshold corresponding to adaptive learning, is the adjusted feature weight, .
4. The visual target detection method for a strong noise interference environment according to claim 1, characterized in that In Step 3, before the SPPF module in the backbone network of YOLOv8, design the Window Self-Attention Mechanism (WSW) module, specifically including: Step 3.1, Calculate the multi-head self-attention within the window in the Window Self-Attention Mechanism (WSW) module; Step 3.2, Calculate the multi-head self-attention between the windows in the Window Self-Attention Mechanism (WSW) module.
5. The visual target detection method for a strong noise interference environment according to claim 4, characterized in that Step 3.1 includes: Step 3.11, calculate the query , key , value matrices of the input feature map of the window; Step 3.12, calculate the self-attention score matrix ; By calculating the dot product of the query and key matrices , obtain the similarity matrix of the input features, divide each element in the similarity matrix by for scaling to obtain the self-attention score matrix , and the calculation formula is as follows: (5) Among them, is the key dimension of the matrix; Step 3.13, calculate the weight matrix of the self-attention scores ; add the self-attention score matrix obtained in Step 3.12 to the relative position bias matrix ; the matrix is constructed according to the relative position indices of the input features in the window; perform , sum of normalization operation so that each normalized element is within the interval to obtain the weight matrix of the self-attention scores, and the calculation formula is as follows: (6) Step 3.14, calculate the output Output of the single-head self-attention; use the weight matrix of the self-attention scores for the value matrix weighted sum to obtain the output Output of the final single-head self-attention. The calculation formula is as follows: (7) Step 3.15, calculate the output of the multi-head self-attention; for the input feature matrix Define multiple groups of different parameter matrices. At the same time, calculate multiple groups of query, key, and value matrices, the self-attention score matrix, and its weight matrix in parallel. Concatenate the Output obtained in Step 3.14 to get the output of the multi-head self-attention within each window, and perform a linear transformation to keep the output dimension of the multi-head self-attention the same as the input feature map dimension of each window.
6. The visual target detection method for a strong noise interference environment according to claim 5, wherein Step 3.2 includes: Step 3.21, Slide the windows divided from the input feature map in Step 3.1 to achieve information transfer between the windows; Step 3.22, Move and merge the newly generated windows to maintain the computational amount before the sliding window operation; Step 3.23, Based on the window merging operation in Step 3.22, calculate the multi-head self-attention separately in each window; Repeat Step 3.11 to calculate the query, key, and value matrices of the input feature map of the window, and Step 3.12 to calculate the self-attention score matrix; Step 3.24, after obtaining the self-attention score matrix, an additional mask matrix is set to isolate the information of different regions; the setting of the mask matrix satisfies the following condition: after the self-attention score matrix, the relative position bias matrix, and the mask matrix are added together, the product of the elements from different windows can be set to 0 after the normalization operation; using this condition, after calculating the self-attention score matrix according to Step 3.12, the mask matrix is solved backward; Step 3.25, perform a normalization operation on the sum of the self-attention score matrix, the relative position bias matrix, and the mask matrix to obtain the weight matrix of the self-attention scores; Step 3.26, Repeat Step 3.14 to calculate the output of the single-head self-attention and Step 3.15 to calculate the output of the multi-head self-attention to obtain the multi-head self-attention between the windows; After the calculation is completed, restore the windows to their original positions.
7. The visual target detection method for a strong noise interference environment according to claim 1, wherein Step 4 includes: After improving the backbone network of YOLOv8 according to Steps 1 to 3, connect it with the feature fusion network and detection head of the YOLOv8 network to establish a complete high-precision and fast detection model for targets in a spatial strong noise interference environment.
8. A visual target detection device for a strong noise interference environment, characterized in that, Including: Deep Residual Shrinkage Block (DRSB) addition module, In the backbone network of YOLOv8, replace the C2f module with the Deep Residual Shrinkage Block (DRSB) to improve the original backbone network of YOLOv8. Soft thresholding processing introduction module, Introduce soft thresholding processing to effectively eliminate noise features and improve the feature extraction process of the original backbone network of YOLOv8. Attention mechanism module WSW design module, which designs the attention mechanism module WSW; before the SPPF module in the backbone network of YOLOv8, the attention mechanism module WSW is designed. By using the sliding window operation, the multi-head self-attention is calculated successively within and between the windows of the feature map to improve the interaction of feature information, and the backbone network of the original YOLOv8 is improved; Detection model establishment module, which establishes a high-precision and fast detection model for targets in a spatially strong noise interference environment; connects the improved backbone network of YOLOv8 with the feature fusion network and detection head of the YOLOv8 network to establish a complete model, and conducts training and testing to verify the effectiveness and generality of the visual target detection method for strong noise interference environments.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the visual target detection method for strong noise interference environments according to any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the visual target detection method for strong noise interference environments according to any one of claims 1 to 7.
Citation Information
Cited By
Intelligent denoising and enhancing method and system for ground penetrating radar signals
CN121559512A