Target detection method, system and equipment for construction site with high-noise background based on MNF-YOLO and medium

By combining the ALSFB and NF Upsample structures on the YOLOv8 model, the problem of insufficient target detection accuracy and generalization capabilities in high noise environments on the construction site is solved, and higher detection accuracy and adaptability are achieved.

CN120182789APending Publication Date: 2025-06-20XIJING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510595973.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The existing technology lacks the accuracy and generalization ability of target detection in high-noise environments on construction sites, resulting in mis-checking and missed inspection problems.

Method used

Based on the YOLOv8 model, an MNF-YOLO detection method is constructed dynamically by combining the ALSFB structure based on the attention mechanism and the feature shift mechanism and an NF Upsample structure containing adaptive noise reduction convolution blocks.

Benefits of technology

It significantly improves the accuracy and generalization ability of construction site target detection, effectively suppresses background noise interference, and enhances the adaptability and reliability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182789A_ABST
    Figure CN120182789A_ABST
Patent Text Reader

Abstract

The invention discloses an MNF-YOLO-based high-noise background construction site target detection method, system and device, and a medium. The method comprises the following steps: obtaining SODA and WOTR data sets; dividing into a training set, a verification set and a test set; constructing an ALSFB structure based on an attention mechanism and a feature shift mechanism; an NF Upsample structure containing the adaptive noise reduction convolution block is constructed; the method is based on a YOLOv8 model and comprises a backbone network, a neck network and a detection head, a C2f module of the backbone network is replaced by an ALSFB structure based on an attention mechanism and a feature shift mechanism, a UPsample structure in a path aggregation network PVN in the neck network is replaced by an NF Upsample structure containing an adaptive noise reduction convolution block, and the NF Upsample structure is used for reducing noise. An ALSFB structure based on an attention mechanism and a feature shift mechanism and an NF Upsample structure containing an adaptive noise reduction convolution block are combined to construct a comprehensive detection method MNF-YOLO structure; respectively training, verifying and testing the MNF-YOLO structure of the comprehensive detection method; evaluating a detection result; the system, the equipment and the medium are used for implementing the method. The method has the advantages of high detection precision and good generalization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision and deep learning technology, and in particular to a method, system, device and medium for detecting a target at a construction site with a high noise background based on MNF-YOLO. Background Art

[0002] The construction site is a diversified and dynamically changing complex system, with various types of personnel, equipment and construction materials distributed in an interlaced manner, and the environment is diverse and highly uncertain. In current construction projects, in order to ensure the efficiency and economy of the project, multiple construction units are usually required to carry out construction tasks at the same time. To achieve the efficiency and safety of multi-party collaboration, real-time monitoring of the location, quantity and status of various objects on the construction site is a vital guarantee. In addition, when carrying out construction arrangements and safety management, regular or irregular safety inspections of the construction site are also essential. However, the current management of personnel, equipment and materials on the construction site mostly relies on manual records and reporting, and safety inspections are also mainly completed manually. This traditional manual management model faces many challenges. For example: manual records and reporting are highly subjective; manual management is inefficient; and traditional manual management costs are high.

[0003] In recent years, the target detection method based on improved YOLO (You Only Look Once) has been widely used in various fields due to its fast, robust and simple characteristics. The algorithm based on YOLO has also been widely used in industrial scenarios such as surface defects, showing its strong practicality and adaptability. However, the research on the target detection task at the construction site is still relatively limited. The high noise environment at the construction site is one of the main challenges in the target detection task. In the feature extraction stage, high-intensity noise will cause the model to mistakenly judge the noise as a target, resulting in false detection. At the same time, the interference of noise will cover the key information of the target, resulting in incomplete feature extraction, further causing the problem of missed detection. In the feature fusion stage, the traditional upsampling structure will greatly expand the residual noise information of the previous layer, thereby further increasing the detection accuracy and generalization ability of the model in the target detection task at the construction site.

[0004] At present, most existing research on construction site detection tasks tends to directly apply the general YOLO algorithm to the task. Even if some models have been improved in pursuit of higher performance, these improvements are mainly focused on optimizing the model itself, without improving the specific characteristics of the construction site detection task. This limitation of ignoring the application requirements in the high-noise environment unique to the construction site has resulted in the existing methods performing unsatisfactorily in construction site target detection.

[0005] The patent application document with the publication number CN 116342862 A discloses a multi-target real-time detection algorithm in the construction scenario. By proposing an efficient behavior recognition algorithm for local-global spatio-temporal aggregation, it can accurately identify the illegal water addition behavior in concrete construction. However, due to the extremely complex construction site environment, there is electromagnetic interference generated by the operation of various large-scale mechanical equipment. At the same time, there are huge differences in lighting conditions in different construction areas, and there are a large number of dynamically changing occlusions, resulting in the problem that the algorithm is prone to false detection and missed detection in some complex areas, and further causing limitations in the application scope of this patent and the inability to further improve the prediction accuracy.

[0006] The patent application document with the publication number CN 112686285 A discloses an engineering quality detection method and system based on computer vision. By using the DIoU-NMS algorithm, it can improve the prediction speed of the model and the recognition effect of overlapping targets. However, since this method depends on the quality and quantity of historical engineering pictures, it will affect the training effect of the model, resulting in limited detection accuracy, thus limiting the further improvement of the prediction accuracy of this patent. Summary of the Invention

[0007] In order to overcome the above deficiencies of the prior art, the purpose of the present invention is to provide a target detection method, system, device and medium for construction sites with high-noise backgrounds based on MNF-YOLO. By taking YOLOv8 as the benchmark model and dynamically combining the ALSFB structure and the NF Upsample structure, it can effectively meet the performance requirements of the target detection task in construction sites, and has the advantages of high detection accuracy and good generalization.

[0008] In order to achieve the above purpose, the technical solution adopted by the present invention is:

[0009] A target detection method for construction sites with high-noise backgrounds based on MNF-YOLO includes the following steps:

[0010] Step 1: Obtain public datasets, and select the SODA for object detection in construction sites and the WOTR dataset for target detection of blind sidewalk obstacles and road surface markings.

[0011] Step 2: Divide both the SODA and WOTR datasets into training sets, validation sets and test sets according to a certain proportion.

[0012] Step 3: Construct an ALSFB structure based on the attention mechanism and the feature shift mechanism.

[0013] Step 4: Construct an NF Upsample structure containing an adaptive noise reduction convolutional block.

[0014] Step Five: Based on the YOLOv8 model, which includes a backbone network, a neck network, and a detection head, replace the C2f module in the backbone network with the ALSFB structure based on the attention mechanism and feature shift mechanism constructed in Step Three, and replace the Upsample structure in the Path Aggregation Network (PVN) in the neck network with the NF Upsample structure containing an adaptive noise reduction convolutional block constructed in Step Four. Combine the ALSFB structure based on the attention mechanism and feature shift mechanism and the NF Upsample structure containing an adaptive noise reduction convolutional block to construct the comprehensive detection method MNF - YOLO structure;

[0015] Step Six: Input the SODA and WOTR training sets divided in Step Two into the comprehensive detection method MNF - YOLO structure constructed in Step Five for training. By optimizing the model parameters, obtain the trained comprehensive detection method MNF - YOLO structure; Input the SODA and WOTR validation sets divided in Step Two into the trained comprehensive detection method MNF - YOLO structure for validation to obtain the comprehensive detection method MNF - YOLO structure with optimal model parameters;

[0016] Step Seven: Input the SODA and WOTR test sets divided in Step Two into the comprehensive detection method MNF - YOLO structure with optimal model parameters obtained in Step Six for testing to obtain the detection results and evaluate the detection results.

[0017] In the above Step Three, the ALSFB structure based on the attention mechanism and feature shift mechanism includes: BN for normalization, the SSAM structure for calculating the attention weights at each spatial position in the feature map, the SRF structure for helping to capture the long - range dependent feature relationships in the image, the CFAM structure for assigning different weights to different feature channels, and Dropout, 1×1 two - dimensional convolution, and the identity mapping of the input feature map Fx for processing the channel attention feature map Fcfam;

[0018] The SSAM structure consists of three parallel convolutional layers for convolving the input, an operation for calculating the attention weights, and a Dropout layer for preventing overfitting;

[0019] The SRF structure includes two 5×5 large convolutional kernels, two channel feature shift operations, three BN for normalization, and a SiLU activation function;

[0020] The CFAM structure includes a global average pooling layer, two linear layers, a ReLU activation function, and a sigmoid activation function;

[0021] The specific operations of the ALSFB structure based on the attention mechanism and feature shift mechanism are as follows:

[0022] Step 3.1: Apply BN (Batch Normalization) to the input feature map Fx for normalization; then, compress the feature dimension through a 1×1 two-dimensional convolution operation to obtain a feature map;

[0023] Step 3.2: Input the feature map obtained in Step 3.1 into the SSAM structure. First, pass the feature map through three parallel convolutions to obtain three feature maps; then calculate the attention weights for two of the feature maps to obtain an attention weight matrix; then perform a Dropout operation on the attention weight matrix; finally, element-wise add the Dropout-processed attention weight matrix to the third feature map to obtain the feature map Fssam;

[0024] Step 3.3: Input the feature map Fssam output by the SSAM structure in Step 3.2 into the SRF structure. Use the first 5×5 large convolution kernel to perform preliminary processing on the feature map Fssam, and perform two channel feature shift operations on the processed feature map Fssam respectively to obtain the feature maps Fsrf1 and Fsrf2; perform a normalization process of one BN on the feature maps Fsrf1 and Fsrf2 respectively, and element-wise add them to obtain the feature map Fsrf; combine the feature map obtained by processing the feature map Fsrf through the second 5×5 large convolution kernel with the feature map Fsrf, and after passing through the normalization process of BN and the SiLU activation function, obtain the output feature map Fsrf-out of the SRF structure;

[0025] Step 3.4: Input the output feature map Fsrf-ou obtained by the SRF structure in Step 3.3 into the CFAM structure. Perform global average pooling on the output feature map Fsrf-ou to obtain the pooled feature Fcawg on the channel. Input the pooled feature Fcawg into a linear layer to reduce the dimension to 1 / 16 of the input channel number, and perform activation processing through the ReLU activation function. Then restore the dimension to the original channel number through another linear layer, and finally obtain the weight feature vector of channel attention using the sigmoid function; element-wise add the weight feature vector of channel attention to the output feature map Fsrf-out of the SRF structure to obtain the final channel attention feature map Fcfam;

[0026] Step 3.5: Process the final channel attention feature map Fcfam obtained in Step 3.4 using Dropout, a 1×1 two-dimensional convolution, and the identity mapping of the input feature map Fx. Among them, Dropout is used to prevent the model from overfitting, the 1×1 two-dimensional convolution is used to restore the dimension of the feature map, and the identity mapping of the input feature map Fx is used to preserve the integrity of the original feature information, to obtain the output of the ALSFB structure.

[0027] In step 4, the NF Upsample structure containing the adaptive noise reduction convolutional block introduces an adaptive noise reduction convolutional block before the conventional upsampling operation, generating a noise reduction feature map through additional convolutional operations to offset the noise information in the input features;

[0028] The NF Upsample structure containing the adaptive noise reduction convolutional block specifically includes: an adaptive noise reduction convolutional layer and an upsampling operation;

[0029] The specific operations of the NF Upsample structure containing the adaptive noise reduction convolutional block are as follows:

[0030] First, take the output of the ALSFB structure based on the attention mechanism and feature shift mechanism in step 3.5 as the input, pass it through the adaptive noise reduction convolutional layer to obtain a new feature map; then add the new feature map and the output of the ALSFB structure based on the attention mechanism and feature shift mechanism element by element to obtain a fused feature map; finally, perform an upsampling operation on the fused feature map to obtain a noise reduction feature map.

[0031] In step 5, the specific method for constructing the comprehensive detection method MNF-YOLO structure is as follows:

[0032] The comprehensive detection method MNF-YOLO structure is based on the YOLOv8 model and includes a backbone network, a neck network, and a detection head. Replace the C2f module in the backbone network with the ALSFB structure based on the attention mechanism and feature shift mechanism constructed in step 3, and replace the Upsample structure in the path aggregation network PVN in the neck network with the NF Upsample structure containing the adaptive noise reduction convolutional block constructed in step 4; among them, the ALSFB structure based on the attention mechanism and feature shift mechanism consists of a BN for normalization processing, an SSAM structure for calculating the attention weights at each spatial position in the feature map, an SRF structure for helping to capture the long-range dependent feature relationships in the image, a CFAM structure for assigning differential weights to different feature channels, and a series connection of Dropout, a 1×1 two-dimensional convolution, and the identity mapping of the input feature map Fx. The ALSFB structure based on the attention mechanism and feature shift mechanism first inputs the feature map Fx and finally obtains the output of the ALSFB structure based on the attention mechanism and feature shift mechanism;

[0033] The NF Upsample structure containing the adaptive noise reduction convolutional block consists of an adaptive noise reduction convolutional layer and an upsampling operation; the NF Upsample structure containing the adaptive noise reduction convolutional block takes the output of the ALSFB structure based on the attention mechanism and feature shift mechanism as the input and finally obtains a noise reduction feature map.

[0034] In step six, the MNF-YOLO structure of the comprehensive detection method is trained, and the specific operations are as follows:

[0035] The SODA and WOTR training sets divided in step two are used to train the MNF-YOLO structure of the comprehensive detection method. By optimizing the parameters of the MNF-YOLO structure of the comprehensive detection method, the trained MNF-YOLO structure of the comprehensive detection method is obtained. Among them, the SODA training set for object detection at the construction site is used as the main training set to verify the core method, and the WOTR training set for object detection of blind sidewalk obstacles and road surface markings is used to evaluate the generalization performance of the MNF-YOLO structure of the comprehensive detection method in different application scenarios.

[0036] The specific method of step seven includes:

[0037] First, the SODA and WOTR test sets divided in step two are input into the MNF-YOLO structure of the comprehensive detection method with the optimal model parameters obtained in step six for testing to obtain the detection results. Then, the mean average precision (mAP) between the detection results and the true position and category information of the targets in the true result annotation data under multiple intersection over union (IoU) is selected as the main evaluation index. The threshold values of these intersection over union (IoU) range from 0.5 to 0.95.

[0038] Compared with the prior art, the beneficial effects of the present invention are:

[0039] 1. Based on the YOLOv8 model, the present invention dynamically combines the ALSFB structure based on the attention mechanism and the feature shift mechanism and the NF Upsample structure containing the adaptive noise reduction convolution block, and proposes an object detection method for the construction site with a high-noise background based on MNF-YOLO. The attention and feature shift mechanism of the ALSFB structure is used to suppress background noise, and the adaptive noise reduction convolution block of the NF Upsample structure is used to suppress noise expansion and residual noise. The two cooperate to improve the detection accuracy of the MNF-YOLO structure of the comprehensive detection method, enhance the adaptability and reliability, and broaden the application scope. The present invention can effectively meet the performance requirements of the object detection task at the construction site, provide an efficient and reliable solution for the object detection task in a high-noise environment, and provide a theoretical basis and practical reference for future related research.

[0040] 2. The datasets used in the present invention are two high-quality publicly available datasets, namely SODA for object detection at construction sites and WOTR for object detection of blind path obstacles and road surface markings. Among them, the SODA training set is used as the main training set to verify the core method, which can fully reflect the complex environment of construction sites; the WOTR training set is used to evaluate the generalization performance of the comprehensive detection method MNF-YOLO structure in different application scenarios, and comprehensively test its adaptability and reliability in complex scenarios.

[0041] 3. The SSAM structure used in the ALSFB structure based on the attention mechanism and feature shift mechanism designed in the present invention, which is used to calculate the attention weights of each spatial position in the feature map, enhances the attention to the target area, thereby achieving the purpose of suppressing background noise.

[0042] 4. The SRF structure used in the ALSFB structure based on the attention mechanism and feature shift mechanism designed in the present invention, which helps to capture the long-range dependent feature relationships in the image, significantly expands the receptive field of the model, enables the comprehensive detection method MNF-YOLO structure to capture the long-range dependent feature relationships in the image, thereby deeply understanding the relationship between foreground objects and the background, and further achieving the effect of suppressing background noise.

[0043] 5. The CFAM structure used in the ALSFB structure based on the attention mechanism and feature shift mechanism designed in the present invention, which is used to assign differential weights to different feature channels, enhances the attention to high-level semantic features, thereby enabling the comprehensive detection method MNF-YOLO structure to achieve the effect of further suppressing background noise.

[0044] 6. The NF Upsample structure containing adaptive noise reduction convolutional blocks designed in the present invention can effectively suppress the residual internal noise of the target and a small amount of background noise in the previous layer, thereby significantly improving the detection accuracy and generalization ability of the model for targets in complex construction site scenarios.

[0045] 7. A target detection method for construction sites with high-noise backgrounds designed in the present invention, which is based on the YOLOv8 model and dynamically combines the ALSFB structure based on the attention mechanism and feature shift mechanism and the NF Upsample structure containing adaptive noise reduction convolutional blocks, effectively meets the performance requirements of construction site target detection tasks, provides an efficient and reliable solution for target detection tasks in high-noise environments, and provides a theoretical basis and practical reference for future related research.

[0046] In summary, based on the YOLOv8 model, the present invention dynamically combines the ALSFB structure based on the attention mechanism and the feature shift mechanism and the NF Upsample structure containing the adaptive noise reduction convolutional block, and proposes a target detection method for construction sites with high-noise backgrounds. It significantly suppresses the interference of background noise on target feature extraction, solves the problem of insufficient performance of traditional feature extraction modules in high-noise environments, provides an efficient and reliable solution for target detection tasks in high-noise environments, and provides a theoretical basis and practical reference for future related research. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] Figure 1 It is a flowchart of the detection method of the present invention.

[0048] Figure 2 It is a structural diagram of ALSFB based on the attention mechanism and the feature shift mechanism.

[0049] Figure 3 It is a structural diagram of NF Upsample containing the adaptive noise reduction convolutional block.

[0050] Figure 4 It is a diagram showing the detection results of each module in the ablation experiment.

[0051] Figure 5 It is a visualized heat map of the ablation experiment.

[0052] Figure 6 It is a diagram showing the detection results of MNF-YOLO and SOTA models.

[0053] Figure 7 It is a diagram showing the detection results of MNF-YOLO and SOTA models on the WOTR dataset. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0054] The following is an explanation of the present invention with reference to the accompanying drawings.

[0055] Aiming at the problems of target detection accuracy and generalization ability in construction sites. Based on YOLOv8, the present invention proposes a target detection method for construction sites with high-noise backgrounds, which dynamically combines the ALSFB (Attention-based Low-dimensional Shifted Receptive Field Block) structure that can enhance the model's attention to foreground targets and high-level semantic information and the NF Upsample (Noise-Filtered Upsample) structure that can solve the problem of noise expansion and significantly suppress the internal noise of targets, and has achieved excellent performance in the construction site target detection task.

[0056] As Figure 1 shown, a target detection method for a construction site with a high-noise background based on MNF-YOLO includes: obtaining a public dataset, selecting two high-quality datasets, SODA and WOTR, to verify the effectiveness and generalization ability of the method; dividing both the SODA and WOTR datasets into a training set, a validation set, and a test set at a ratio of 7:2:1; designing an ALSFB structure based on a powerful attention mechanism and a feature shift mechanism; introducing an NF Upsample structure with an adaptive noise reduction convolutional block; designing a comprehensive detection method, the MNF-YOLO structure, based on the YOLOv8 model and combining the ALSFB structure and the NF Upsample structure; training the comprehensive detection method, the MNF-YOLO structure, and retaining the optimal weights; using the trained optimal weights for testing and evaluating the detection results.

[0057] The specific steps include:

[0058] Step 1: Obtain a public dataset, select two high-quality datasets, SODA for object detection at construction sites and WOTR for target detection of blind sidewalk obstacles and road surface markings, to improve the effectiveness and generalization ability of the verification method;

[0059] In Step 1, two high-quality public datasets, SODA and WOTR, are selected. Among them, SODA (SiteObject Detection Dataset) is a target detection dataset for object detection at construction sites jointly released by Tsinghua University and South China University of Technology. WOTR is a target detection dataset for blind sidewalk obstacles and road surface markings released by Guangxi Normal University.

[0060] The SODA (Site Object Detection Dataset) dataset contains 19,846 images and 286,201 annotated targets, covering a variety of shooting perspectives, different lighting conditions, and construction stages, and covering 15 common construction targets, including person, helmet, vest, board, wood, rebar, brick, scaffold, handcart, cutter, ebox, hopper, hook, fence, and slogan.

[0061] The WOTR dataset contains 13,928 images and 189,998 labeled objects, covering 15 common obstacles including tree, ashcan, roadblock, fire hydrant, reflective cone, warning column, pole, person, dog, bicycle, truck, motorcycle, bus, car, tricycle, red light, crosswalk, sign, green light, and blind road, as well as 5 road judgment objects.

[0062] Step 2: Divide both the SODA and WOTR datasets into a training set, a validation set, and a test set in a ratio of 7:2:1.

[0063] Step 3: Construct an ALSFB structure based on the attention mechanism and the feature shift mechanism to achieve the purpose of suppressing background noise.

[0064] As Figure 2 shown, Step 3: The ALSFB structure constructed in the present invention based on the attention mechanism and the feature shift mechanism includes: BN for normalization processing, an SSAM (Spatial Self - Attention Mechanism) structure for calculating the attention weights at each spatial position in the feature map, an SRF (Shifted Receptive Field) structure helpful for capturing the long - range dependent feature relationships in the image, a CFAM (Channel Feature Attention Mechanism) structure for assigning different weights to different feature channels, and Dropout, a 1×1 two - dimensional convolution, and an identity mapping of the input feature map Fx for processing the channel attention feature map Fcfam.

[0065] The SSAM structure consists of three parallel convolutional layers for convolutional operations on the input, an operation for calculating the attention weights, and a Dropout layer for preventing overfitting.

[0066] The SSAM structure enhances the attention to the target region by calculating the attention weights at each spatial position in the feature map, thereby achieving the purpose of suppressing background noise.

[0067] The SRF structure includes two 5×5 large convolutional kernels, two channel feature shift operations, three BN for normalization processing, and a SiLU activation function.

[0068] The SRF structure significantly expands the receptive field of the model through feature map shifting and large convolutional kernel operations, prompting the model to capture the feature relationships of long-range dependencies in the image, thereby deeply understanding the relationship between foreground objects and the background, and further achieving the effect of suppressing background noise.

[0069] The CFAM structure includes a global average pooling layer, two linear layers, a ReLU activation function, and a sigmoid activation function.

[0070] The CFAM structure enhances the attention to high-level semantic features by suppressing the model's dependence on low-level features such as texture and edges, thereby prompting the model to further suppress background noise.

[0071] The specific operations are as follows:

[0072] Step 3.1: Apply BN (Batch Normalization) to the input feature map Fx for normalization to improve the stability of training. Then, compress the feature dimension through a 1×1 two-dimensional convolution operation to expand the receptive field and effectively reduce the subsequent computational burden, obtaining a feature map;

[0073] Step 3.2: To enhance the model's ability to distinguish between targets and noisy backgrounds, input the feature map obtained in Step 3.1 into the SSAM structure. First, pass the feature map through three parallel convolutions to obtain three feature maps; then calculate the attention weight for two of the feature maps to obtain an attention weight matrix; then perform a Dropout operation on the attention weight matrix to reduce the co-adaptation space between neurons, thereby avoiding the problem of the model over-relying on certain specific weights; finally, element-wise add the attention weight matrix after Dropout processing to the third feature map to obtain the feature map Fssam;

[0074] Step 3.3: Input the feature map Fssam output by the SSAM structure in Step 3.2 into the SRF structure. Use the first 5×5 large convolutional kernel to preliminarily process the feature map Fssam, and perform two channel feature shifting operations on the processed feature map Fssam respectively to obtain the feature maps Fsrf1 and Fsrf2; perform a BN normalization process on the feature maps Fsrf1 and Fsrf2 respectively, and element-wise add them to obtain the feature map Fsrf; combine the feature map obtained by processing the feature map Fsrf through the second 5×5 large convolutional kernel with the feature map Fsrf, and after BN normalization processing and SiLU activation function processing, obtain the output feature map Fsrf-out of the SRF structure;

[0075] Step 3.4: Input the output feature map Fsrf-ou obtained from the SRF structure in Step 3.3 into the CFAM structure. Perform global average pooling on the output feature map Fsrf-ou to obtain the pooled feature Fcawg on the channels. Input the pooled feature Fcawg into a linear layer to reduce the dimension to 1 / 16 of the number of input channels, and perform activation processing through the ReLU activation function. Then, restore the dimension to the original number of channels through another linear layer. Finally, use the sigmoid function to obtain the weight feature vector of channel attention; Add the weight feature vector of channel attention and the output feature map Fsrf-out of the SRF structure element-wise to obtain the final channel attention feature map Fcfam;

[0076] Step 3.5: Process the final channel attention feature map Fcfam obtained in Step 3.4 using Dropout, 2D convolution with a kernel size of 1×1, and the identity mapping of the input feature map Fx. Among them, Dropout is used to prevent the model from overfitting, 2D convolution with a kernel size of 1×1 is used to restore the dimension of the feature map, and the identity mapping of the input feature map Fx is used to preserve the integrity of the original feature information, obtaining the output of the ALSFB structure based on the attention mechanism and the feature shift mechanism.

[0077] As Figure 3 shown, Step 4: Construct an NF Upsample structure containing an adaptive noise reduction convolution block to solve the noise expansion problem of traditional Upsample and the noise problem inside the target;

[0078] In the above Step 4, when constructing an NF Upsample structure containing an adaptive noise reduction convolution block, an adaptive noise reduction convolution block is introduced before the conventional upsampling operation. Through additional convolution operations, a noise reduction feature map is generated to offset the noise information in the input features;

[0079] The NF Upsample structure containing an adaptive noise reduction convolution block specifically includes: an adaptive noise reduction convolution layer and an upsampling operation;

[0080] The NF Upsample structure containing an adaptive noise reduction convolution block can effectively suppress the internal noise of the target remaining from the previous layer and a small amount of background noise, significantly improving the detection accuracy and generalization ability of the model for targets in complex construction site scenarios.

[0081] The specific operations of the NF Upsample structure containing an adaptive noise reduction convolution block are as follows:

[0082] First, take the output of the ALSFB structure based on the attention mechanism and the feature shift mechanism in step 3.5 as the input, pass it through the convolutional layer with adaptive noise reduction to obtain a new feature map; then add the new feature map and the output of the ALSFB structure based on the attention mechanism and the feature shift mechanism element by element to obtain a fused feature map; finally, perform an upsampling operation on the fused feature map to obtain a noise-reduced feature map.

[0083] The main function of the ALSFB structure is to suppress the influence of background noise, while the NF Upsample structure mainly addresses the noise expansion of traditional Upsample and the noise problem inside the target.

[0084] Step Five: Based on the YOLOv8 model, which includes a backbone network, a neck network, and a detection head, replace the C2f module in the backbone network with the ALSFB structure based on the attention mechanism and the feature shift mechanism constructed in step three, and replace the Upsample structure in the Path Aggregation Network (PVN) in the neck network with the NF Upsample structure containing the adaptive noise reduction convolutional block constructed in step four. Combine the ALSFB structure based on the attention mechanism and the feature shift mechanism and the NF Upsample structure containing the adaptive noise reduction convolutional block to construct the comprehensive detection method MNF-YOLO structure; dynamically combine the ALSFB structure that can enhance the model's attention to foreground targets and high-level semantic information and the NF Upsample structure that can solve the problem of noise expansion and significantly suppress the noise inside the target. The two cooperate to improve the detection accuracy of the comprehensive detection method MNF-YOLO structure, enhance the adaptability and reliability, broaden the application scope, and effectively meet the performance requirements of the target detection task at the construction site.

[0085] In the above step five, the ALSFB structure and the NF Upsample structure proposed by the MNF-YOLO structure respectively improve the C2f module in the backbone and the Upsample structure in PAN. The C2f module is a component of the YOLOv8 backbone network. PAN (Path Aggregation Network) is an important structure in the YOLOv8 neck network, and the Upsample structure is a key component in PAN. The ALSFB structure improves the C2f module in the backbone to suppress the influence of background noise; the NF Upsample structure improves the Upsample structure in PAN to address the noise expansion of traditional Upsample and the noise problem inside the target.

[0086] In the above step five, the specific method for constructing the comprehensive detection method MNF-YOLO structure is as follows:

[0087] The comprehensive detection method MNF-YOLO structure is based on the YOLOv8 model, including a backbone network, a neck network, and a detection head. The C2f module of the backbone network is replaced with the ALSFB structure based on the attention mechanism and feature shift mechanism constructed in step three, and the Upsample structure in the path aggregation network PVN in the neck network is replaced with the NF Upsample structure containing an adaptive noise reduction convolutional block constructed in step four. Among them, the ALSFB structure based on the attention mechanism and feature shift mechanism consists of a BN for normalization processing, an SSAM structure for calculating the attention weights of each spatial position in the feature map, an SRF structure for helping to capture the long-range dependence feature relationships in the image, a CFAM structure for assigning differential weights to different feature channels, and a series connection of Dropout, a 1×1 two-dimensional convolution, and the identity mapping of the input feature map Fx for processing the channel attention feature map Fcfam. The ALSFB structure based on the attention mechanism and feature shift mechanism first inputs the feature map Fx and finally obtains the output of the ALSFB structure based on the attention mechanism and feature shift mechanism.

[0088] The NF Upsample structure containing an adaptive noise reduction convolutional block consists of an adaptive noise reduction convolutional layer and an upsampling operation. The NF Upsample structure containing an adaptive noise reduction convolutional block takes the output of the ALSFB structure based on the attention mechanism and feature shift mechanism as the input and finally obtains a noise reduction feature map.

[0089] Step six: Input the SODA and WOTR training sets divided in step two into the comprehensive detection method MNF-YOLO structure constructed in step five for training. By optimizing the model parameters, obtain the trained comprehensive detection method MNF-YOLO structure. Input the SODA and WOTR validation sets divided in step two into the trained comprehensive detection method MNF-YOLO structure for validation to obtain the comprehensive detection method MNF-YOLO structure with the optimal model parameters.

[0090] In the said step six, the model training is carried out on a server configured with an i9-14900KF 3.2GHz CPU, 128G RAM, and an NVIDIA GeForce RTX 3090 24GB GPU. The experimental process uses the Python programming language and is completed with the help of the PyTorch deep learning framework. Among them, the Python version is 3.8.16, the PyTorch version is 1.13.1, and the CUDA version is 11.7.

[0091] In the said step six, the specific operation of training the comprehensive detection method MNF-YOLO structure is as follows:

[0092] Use the SODA and WOTR training sets divided in Step 2 to train the comprehensive detection method MNF-YOLO structure. By optimizing the parameters of the comprehensive detection method MNF-YOLO structure, the trained comprehensive detection method MNF-YOLO structure is obtained. Among them, the SODA training set for object detection at the construction site is used as the main training set to verify the core method, and the WOTR training set for object detection of blind sidewalk obstacles and road surface markings is used to evaluate the generalization performance of the comprehensive detection method MNF-YOLO structure in different application scenarios.

[0093] The specific method for optimizing the parameters of the comprehensive detection method MNF-YOLO structure is as follows:

[0094] Use the Stochastic Gradient Descent (SGD) optimizer to iteratively optimize the parameters of the comprehensive detection method MNF-YOLO structure, and dynamically adjust the learning rate. To ensure the fairness of training, the batch_size of all comprehensive detection method MNF-YOLO structures is set to 60% of the video memory capacity. To ensure sufficient training, the maximum number of training epochs is set to 1000 times, and an early stopping strategy to prevent overfitting is adopted. The determination iteration number for early stopping is 100 times. By testing the performance of the MNF-YOLO structure under different parameter combinations on the validation set, the parameter settings of the best-performing MNF-YOLO structure are found.

[0095] Step 7: Input the SODA and WOTR test sets divided in Step 2 into the comprehensive detection method MNF-YOLO structure with the optimal model parameters obtained in Step 6 for testing, obtain the detection results, and evaluate the detection results.

[0096] The specific method of Step 7 includes:

[0097] First, input the SODA and WOTR test sets divided in Step 2 into the comprehensive detection method MNF-YOLO structure with the optimal model parameters obtained in Step 6 for testing, and obtain the detection results. Then, select the mean Average Precision (mAP) at multiple Intersection over Union (IoU) values between the detection results and the true position and category information of the targets in the true result annotation data as the main evaluation index. The IoU threshold values range from 0.5 to 0.95 with a step size of 0.05. By calculating this index, the detection performance of the comprehensive detection method MNF-YOLO structure in complex scenarios can be better reflected.

[0098] In Step 7, select mAP50:95 at high IoU as the main evaluation index to better reflect the detection performance of the model in complex scenarios. At the same time, to evaluate the deployment performance of the model, use Parameter to evaluate the volume size of the model. Here, Parameter represents the number of variables inside the model, and these variables are learned through training data.

[0099] The detection results are compared with other models through ablation experiments.

[0100] In the comparison experiment with SOTA models, the MNF-YOLO structure of the comprehensive detection method constructed by the present invention is compared with these SOTA models such as YOLOv5, YOLOv6, YOLOv7, and YOLOv8 on the SODA dataset; in the generalization verification experiment, it is compared with these SOTA models on the WOTR dataset.

[0101] Finally, to evaluate the deployment performance of the MNF-YOLO structure of the comprehensive detection method, the volume size of the model finally output by the MNF-YOLO structure of the comprehensive detection method is evaluated using Parameter, which represents the number of variables inside the MNF-YOLO structure of the comprehensive detection method.

[0102] As Figure 4 shown, it can be seen that the number of missed detections and misdetections of the MNF-YOLO structure of the comprehensive detection method designed by the present invention is significantly reduced compared with YOLOv8, AL-YOLO, and NU-YOLO.

[0103] Taking YOLOv8 as the baseline, the model after introducing the ALSFB structure is named AL-YOLO, and the model after introducing the NF Upsample structure is named NU-YOLO. The model after introducing the ALSFB structure and the NF Upsample structure is named MNF-YOLO.

[0104] As Figure 5 shown, the MNF-YOLO structure of the comprehensive detection method designed by the present invention has stronger adaptability when dealing with targets in complex construction site scenarios compared with YOLOv8, AL-YOLO, and NU-YOLO, can fully suppress various noises, and at the same time maintain accurate capture of target features.

[0105] Among them, the red area represents the confidence level, that is, the focus of the model. The higher the attention, the redder the color, and the lower the attention, the bluer the color.

[0106] Affected by the construction site noise, the focus of the baseline model YOLOv8 has obvious deviation and it is difficult to accurately focus on the target area. The attention points of the MNF-YOLO structure obtained by combining the ALSFB structure and the NF Upsample structure can not only accurately locate the position of the target, but also accurately cover the complete target, reflecting the comprehensive extraction ability of target information.

[0107] As Figure 6As shown, it presents partial detection results of the comprehensive detection method MNF-YOLO structure and various SOTA models. The MNF-YOLO structure of the comprehensive detection method can successfully detect each target in a complex background environment, while some misdetections and missed detections occur in other SOTA models.

[0108] The SOTA models include YOLOv5, YOLOv6, YOLOv7, and YOLOv8.

[0109] The MNF-YOLO structure of the comprehensive detection method constructed in the present invention demonstrates higher detection accuracy and robustness, fully verifying its strong competitiveness and practical application potential in the object detection task in a complex construction site scenario.

[0110] As Figure 7 shown, it presents the detection results of the MNF-YOLO structure of the comprehensive detection method constructed in the present invention and various SOTA models on the WOTR dataset. The MNF-YOLO structure of the comprehensive detection method can successfully detect each target in the WOTR dataset, while different degrees of missed detections exist in other SOTA models. It proves that the MNF-YOLO structure of the comprehensive detection method has excellent detection accuracy and generalization performance and can meet the object detection requirements of the construction site.

[0111] In addition, the detection results of the MNF-YOLO structure of the comprehensive detection method constructed in the present invention and various SOTA models on the WOTR dataset also prove that the MNF-YOLO structure of the comprehensive detection method constructed in the present invention performs excellently in the blind sidewalk obstacle detection task, fully meeting the expectations of the present invention for the generalization and reliability of the model in practical applications.

[0112] To achieve accurate detection of construction site targets, the present invention takes YOLOv8 as a benchmark, dynamically combines the ALSFB structure and the NF Upsample structure, and designs the MNF-YOLO structure of the comprehensive detection method. Among them, the main function of the ALSFB structure is to suppress the influence of background noise, and the main function of the NF Upsample structure is to address the noise expansion of the traditional Upsample and the noise problem inside the target. The MNF-YOLO structure of the comprehensive detection method constructed in the present invention can effectively suppress the object detection model with high noise interference and meet the performance requirements of the construction site object detection task.

[0113] The constructed comprehensive detection method MNF-YOLO structure of the present invention includes an ALSFB structure and an NF Upsample structure. The ALSFB structure suppresses background noise by means of an attention and feature shift mechanism, and the NF Upsample structure solves the problem of noise expansion with an adaptive noise reduction convolutional block. In the experiments on the SODA and WOTR datasets, the MNF-YOLO structure of the comprehensive detection method has excellent detection accuracy and generalization performance, and can meet the target detection requirements of the construction site.

[0114] The present invention also provides a target detection system for a construction site with a high-noise background based on MNF-YOLO, including:

[0115] A dataset acquisition module for implementing the acquisition of a public dataset in Step 1, and selecting the SODA dataset for object detection at the construction site and the WOTR dataset for target detection of blind sidewalk obstacles and road surface markings;

[0116] A dataset processing module for implementing the division of the SODA and WOTR datasets into a training set, a validation set, and a test set according to a ratio in Step 2;

[0117] An ALSFB structure construction module for implementing the construction of an ALSFB structure based on an attention mechanism and a feature shift mechanism in Step 3;

[0118] An NF Upsample structure construction module for implementing the construction of an NF Upsample structure containing an adaptive noise reduction convolutional block in Step 4;

[0119] A comprehensive detection method MNF-YOLO structure construction module for implementing, based on the YOLOv8 model in Step 5, including a backbone network, a neck network, and a detection head, replacing the C2f module of the backbone network with the ALSFB structure based on an attention mechanism and a feature shift mechanism constructed in Step 3, replacing the Upsample structure in the path aggregation network PVN in the neck network with the NF Upsample structure containing an adaptive noise reduction convolutional block constructed in Step 4, and combining the ALSFB structure based on an attention mechanism and a feature shift mechanism and the NF Upsample structure containing an adaptive noise reduction convolutional block to construct the comprehensive detection method MNF-YOLO structure;

[0120] The comprehensive detection method MNF-YOLO structure training and verification module is used to implement the steps in Step 6, where the SODA and WOTR training sets divided in Step 2 are input into the comprehensive detection method MNF-YOLO structure constructed in Step 5 for training. By optimizing the model parameters, the trained comprehensive detection method MNF-YOLO structure is obtained. The SODA and WOTR validation sets divided in Step 2 are input into the trained comprehensive detection method MNF-YOLO structure for verification, and the comprehensive detection method MNF-YOLO structure with the optimal model parameters is obtained.

[0121] The comprehensive detection method MNF-YOLO structure testing and evaluation module is used to implement the steps in Step 7, where the SODA and WOTR test sets divided in Step 2 are input into the comprehensive detection method MNF-YOLO structure with the optimal model parameters obtained in Step 6 for testing, and the detection results are obtained and evaluated.

[0122] The present invention also provides a target detection device for a construction site with a high-noise background based on MNF-YOLO, including:

[0123] A memory: storing a computer program of the above-mentioned target detection method for a construction site with a high-noise background based on MNF-YOLO, which is a computer-readable device;

[0124] A processor: used to implement the above-mentioned target detection method for a construction site with a high-noise background based on MNF-YOLO when executing the computer program.

[0125] The present invention also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it can implement the above-mentioned target detection method for a construction site with a high-noise background based on MNF-YOLO.

[0126] The above are only the preferred embodiments of the present invention, which do not limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A target detection method for construction sites with high noise background based on MNF-YOLO, characterized in that: The following steps are involved: Step 1: Obtain public datasets, such as SODA for construction site object detection and WOTR for target detection of blind path obstacles and road surface markings; Step 2: Divide the SODA and WOTR datasets into training sets, validation sets, and test sets in proportion; Step 3: Construct an ALSFB structure based on the attention mechanism and feature shift mechanism; Step 4: Construct the NF Upsample structure containing the adaptive denoising convolution block; Step 5: Based on the YOLOv8 model, the YOLOv8 model includes a backbone network, a neck network and a detection head. The C2f module of the backbone network is replaced with the ALSFB structure based on the attention mechanism and the feature shift mechanism constructed in step 3. The UPsample structure in the path aggregation network PVN in the neck network is replaced with the NF Upsample structure containing the adaptive denoising convolution block constructed in step 4. Combine the ALSFB structure based on the attention mechanism and the feature shift mechanism and the NF Upsample structure containing the adaptive denoising convolution block to construct a comprehensive detection method MNF-YOLO structure. Step 6: Input the SODA and WOTR training sets divided in step 2 into the comprehensive detection method MNF-YOLO structure constructed in step 5 for training, and obtain the trained comprehensive detection method MNF-YOLO structure by optimizing the model parameters; input the SODA and WOTR verification sets divided in step 2 into the trained comprehensive detection method MNF-YOLO structure for verification, and obtain the comprehensive detection method MNF-YOLO structure with the optimal model parameters; Step 7: Input the SODA and WOTR test sets divided in step 2 into the MNF-YOLO structure of the comprehensive detection method with the optimal model parameters obtained in step 6 for testing, obtain the detection results, and evaluate the detection results.

2. According to claim 1, a method for detecting targets at construction sites with high noise background based on MNF-YOLO, characterized in that: In the step 3, the ALSFB structure based on the attention mechanism and feature shift mechanism includes: BN for normalization, SSAM structure for calculating the attention weight of each spatial position in the feature map, SRF structure for capturing long-range dependent feature relationships in the image, CFAM structure for assigning differentiated weights to different feature channels, Dropout for processing the channel attention feature map Fcfam, 1×1 two-dimensional convolution and identity mapping of the input feature map Fx; The SSAM structure consists of three parallel convolutional layers that perform convolution operations on the input, an operation to calculate attention weights, and a Dropout layer to prevent overfitting; The SRF structure includes two 5×5 large convolution kernels, two channel feature shift operations, three BNs for normalization, and a SiLU activation function; The CFAM structure includes a global average pooling layer, two linear layers, a ReLU activation function, and a sigmoid activation function; The specific operations of the ALSFB structure based on the attention mechanism and feature shift mechanism are as follows: Step 3.1: Apply BN (Batch Normalization) to the input feature map Fx for normalization; then, compress the feature dimension through a 1×1 two-dimensional convolution operation to obtain a feature map; Step 3.2: Input the feature map obtained in step 3.1 into the SSAM structure. First, the feature map is subjected to three parallel convolutions to obtain three feature maps. Then, the attention weights are calculated for two of the feature maps to obtain the attention weight matrix. Then, the attention weight matrix is ​​subjected to the Dropout operation. Finally, the attention weight matrix after the Dropout operation is added element by element to the third feature map to obtain the feature map Fssam. Step 3.3: Input the feature map Fssam output by the SSAM structure in step 3.2 into the SRF structure, use the first 5×5 large convolution kernel to perform preliminary processing on the feature map Fssam, perform two channel feature shift operations on the processed feature map Fssam, and obtain feature maps Fsrf1 and Fsrf2; perform a BN normalization process on the feature maps Fsrf1 and Fsrf2, and add them element by element to obtain the feature map Fsrf; combine the feature map Fsrf obtained by processing the second 5×5 large convolution kernel with the feature map Fsrf, and after BN normalization and SiLU activation function processing, obtain the output feature map Fsrf-out of the SRF structure; Step 3.4: Input the output feature map Fsrf-ou obtained by the SRF structure in step 3.3 into the CFAM structure, perform global average pooling on the output feature map Fsrf-ou to obtain the pooling feature Fcawg on the channel, input the pooling feature Fcawg into the linear layer, reduce the dimension to 1 / 16 of the number of input channels, and activate it through the ReLU activation function, and then restore the dimension to the original number of channels through another linear layer, and finally use the sigmoid function to obtain the weight feature vector of the channel attention; add the weight feature vector of the channel attention to the output feature map Fsrf-out of the SRF structure element by element to obtain the final channel attention feature map Fcfam; Step 3.5: Use Dropout, 1×1 two-dimensional convolution and identity mapping of the input feature map Fx to process the final channel attention feature map Fcfam obtained in step 3.4, where Dropout is used to prevent the model from overfitting, 1×1 two-dimensional convolution is used to restore the dimension of the feature map, and the identity mapping of the input feature map Fx is used to retain the integrity of the original feature information to obtain the output of the ALSFB structure.

3. According to claim 1, a method for detecting targets at construction sites with high noise background based on MNF-YOLO, characterized in that: In the step 4, the NF Upsample structure including the adaptive denoising convolution block introduces an adaptive denoising convolution block before the conventional upsampling operation, and generates a denoising feature map through an additional convolution operation to offset the noise information in the input feature; The NF Upsample structure containing the adaptive denoising convolution block specifically includes: an adaptive denoising convolution layer and an upsampling operation; The specific operations of the NF Upsample structure containing the adaptive denoising convolution block are as follows: First, the output of the ALSFB structure based on the attention mechanism and feature shift mechanism in step 3.5 is taken as input, and a new feature map is obtained through a convolutional layer with adaptive denoising. Then, the new feature map is added element by element with the output of the ALSFB structure based on the attention mechanism and feature shift mechanism to obtain a fused feature map. Then, the fused feature map is upsampled to obtain a denoised feature map.

4. The method for detecting a target at a construction site with a high noise background based on MNF-YOLO according to claim 1, characterized in that: In step 5, the specific method for constructing the comprehensive detection method MNF-YOLO structure is: The MNF-YOLO structure of the comprehensive detection method is based on the YOLOv8 model, including a backbone network, a neck network and a detection head. The C2f module of the backbone network is replaced by the ALSFB structure based on the attention mechanism and feature shift mechanism constructed in step 3, and the UPsample structure in the path aggregation network PVN in the neck network is replaced by the NF Upsample structure containing an adaptive denoising convolution block constructed in step 4; wherein, the ALSFB structure based on the attention mechanism and feature shift mechanism is composed of a BN for normalization, an SSAM structure for calculating the attention weight of each spatial position in the feature map, an SRF structure that helps capture the feature relationship of long-range dependencies in the image, a CFAM structure for assigning differentiated weights to different feature channels, and a Dropout for processing the channel attention feature map Fcfam, a 1×1 two-dimensional convolution and an identity mapping of the input feature map Fx in series. The ALSFB structure based on the attention mechanism and feature shift mechanism first inputs the feature map Fx, and finally obtains the output of the ALSFB structure based on the attention mechanism and feature shift mechanism; The NF Upsample structure including the adaptive denoising convolution block consists of an adaptive denoising convolution layer and an upsampling operation; the NF Upsample structure including the adaptive denoising convolution block takes the output of the ALSFB structure based on the attention mechanism and feature shift mechanism as input, and finally obtains the denoising feature map.

5. The method for detecting a target at a construction site with a high noise background based on MNF-YOLO according to claim 1, characterized in that: In step 6, the comprehensive detection method MNF-YOLO structure is trained, and the specific operations are as follows: Use the SODA and WOTR training sets divided in step 2 to train the MNF-YOLO structure of the comprehensive detection method, and obtain the trained MNF-YOLO structure of the comprehensive detection method by optimizing the MNF-YOLO structure parameters of the comprehensive detection method; The SODA training set for object detection on construction sites is used as the main training set to verify the core method, and the WOTR training set for target detection of blind obstacles and road surface signs is used to evaluate the generalization performance of the comprehensive detection method MNF-YOLO structure in different application scenarios.

6. The method for detecting a target at a construction site with a high noise background based on MNF-YOLO according to claim 1, characterized in that: The specific method of step seven includes: First, the SODA and WOTR test sets divided in step 2 are input into the comprehensive detection method MNF-YOLO structure with the optimal model parameters obtained in step 6 for testing to obtain the detection results; then, the mean average precision (mAP) between the detection results and the true position and category information of the target in the true result annotation data under multiple intersection-over-union (IoU) ratios (mAP) is selected as the main evaluation indicator; these intersection-over-union (IoU) thresholds range from 0.5 to 0.

95.

7. A target detection system for construction sites with high noise background based on MNF-YOLO based on the method according to any one of claims 1 to 6, characterized in that: include: The dataset acquisition module is used to obtain public datasets, including the SODA dataset for construction site object detection and the WOTR dataset for target detection of blind path obstacles and road surface markings. The dataset processing module is used to divide the SODA and WOTR datasets into training sets, validation sets, and test sets in proportion; ALSFB structure building module, used to build the ALSFB structure based on attention mechanism and feature shift mechanism; NF Upsample structure building module, used to build the NF Upsample structure containing the adaptive denoising convolution block; The MNF-YOLO structure building module of the comprehensive detection method is used to implement the YOLOv8 model, including the backbone network, the neck network and the detection head. The C2f module of the backbone network is replaced with the ALSFB structure based on the attention mechanism and the feature shift mechanism, and the UPsample structure in the path aggregation network PVN in the neck network is replaced with the NFUpsample structure containing the adaptive denoising convolution block. The ALSFB structure based on the attention mechanism and the feature shift mechanism and the NF Upsample structure containing the adaptive denoising convolution block are combined to construct the MNF-YOLO structure of the comprehensive detection method. The comprehensive detection method MNF-YOLO structure training and verification module is used to input the SODA and WOTR training sets into the comprehensive detection method MNF-YOLO structure for training, and obtain the trained comprehensive detection method MNF-YOLO structure by optimizing the model parameters; input the SODA and WOTR verification sets into the trained comprehensive detection method MNF-YOLO structure for verification, and obtain the comprehensive detection method MNF-YOLO structure with the optimal model parameters; The comprehensive detection method MNF-YOLO structure testing and evaluation module is used to input the SODA and WOTR test sets into the comprehensive detection method MNF-YOLO structure with optimal model parameters for testing, obtain the detection results, and evaluate the detection results.

8. A target detection device for construction sites with high noise background based on MNF-YOLO, characterized in that: include: Memory: a computer program storing a method for detecting a target at a construction site with a high noise background based on MNF-YOLO as described in any one of claims 1 to 6, which is a computer-readable device; Processor: used to implement the target detection method for a construction site with a high noise background based on MNF-YOLO as described in any one of claims 1 to 6 when executing the computer program.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it can implement the target detection method for a construction site with a high noise background based on MNF-YOLO as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Engineering quality detection method and system based on computer vision

    CN112686285A

  • Multi-target real-time detection algorithm in building construction scene

    CN116342862A