Target detection method and system based on LCSD-YOLO network

By building the LCSD-YOLO network, using lightweight backbone network and feature interaction modules, the feature extraction and fusion of the YOLO network is optimized, and the problem of insufficient accuracy in small object detection is solved, and efficient and real-time safety helmet identification for high-altitude workers is achieved.

CN120339695AActive Publication Date: 2025-07-18UNIV OF ELECTRONICS SCI & TECH OF CHINA ZHONGSHAN INST
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510413222.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-18
Estimated Expiration
2045-04-03

AI Technical Summary

Technical Problem

The existing lightweight YOLO network has problems such as impaired feature expression capabilities and insufficient detection accuracy in small-object detection. Especially in the identification of safety helmets for high altitude workers, the existing technology is difficult to meet the needs of high precision, real-time and low latency.

Method used

By introducing the LCSD-YOLO network, a lightweight backbone network is built by introducing the PP-LCNet network architecture, SDDS module and attention-based internal scale feature interaction module, combining deep separation convolution and dynamic upsampling, optimizing feature extraction and fusion, using a non-monotonic focusing mechanism to improve the loss function, and improving the detection accuracy and efficiency of the model.

Benefits of technology

It significantly improves the accuracy and efficiency of small object detection, reduces calculation costs, meets the real-time detection needs of edge equipment, and is suitable for intelligent identification of safety helmets for high-altitude workers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339695A_ABST
    Figure CN120339695A_ABST
Patent Text Reader

Abstract

The invention discloses a target detection method and system based on an LCSD-YOLO network, and the method comprises the steps: obtaining target image data, carrying out the data preprocessing, and obtaining a preprocessed image data set; the method comprises the following steps: based on a YOLO11 network model, introducing a PP-LCNet network architecture, an SDDS module and an attention-based internal scale feature interaction module, and constructing an LCSD-YOLO network model; and performing target identification detection on the preprocessed image data set based on an LCSD-YOLO network model to obtain a target identification detection result. According to the invention, the small target detection capability of the model and the extraction precision and efficiency of target features can be improved, and the precision and efficiency of small target detection are improved. The target detection method and system based on the LCSD-YOLO network can be widely applied to the technical field of target recognition and detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of target recognition and detection, and in particular to a target detection method and system based on the LCSD-YOLO network. Background Art

[0002] At present, methods based on deep learning have significantly optimized the detection performance through the improvement of hardware computing power. However, their inherent defects such as a large number of model parameters and high computational complexity lead to a bottleneck in real-time performance when deployed on edge devices. Although existing lightweight technologies can compress the model scale through structural optimization means such as parameter pruning, low-rank decomposition, compact convolution design, and knowledge distillation, they have the defect that the feature expression ability is damaged, resulting in a decrease in the recognition accuracy of small targets, and the compression rate and the improvement of operation efficiency are not linearly related. Traditional compression cannot meet the requirements of high precision, real-time performance, and low latency in practical applications. For the detection of small targets, such as the intelligent recognition of whether a safety helmet is worn by a high-altitude operator, although existing YOLO series algorithms (such as YOLO11n) perform excellently in the speed-accuracy balance, the images of high-altitude operators collected by ground cameras looking up have a low target pixel ratio (usually less than 0.1% of the image area). Therefore, in recent years, some enhanced networks and feature fusion technologies have been gradually introduced to further improve the robustness of detection, and various attention mechanisms have been added to improve the accuracy of small target detection, but this will bring a large number of parameters to the model, resulting in a decrease in the model detection speed and affecting its real-time performance in practical applications. Summary of the Invention

[0003] In order to solve the above technical problems, the purpose of the present invention is to provide a target detection method and system based on the LCSD-YOLO network, which can improve the small target detection ability of the model, the extraction accuracy and efficiency of target features, and the accuracy and efficiency of small target detection.

[0004] The first technical solution adopted by the present invention is: a target detection method based on the LCSD-YOLO network, comprising the following steps:

[0005] Obtain target image data and perform data preprocessing to obtain a preprocessed image data set;

[0006] Based on the YOLO11 network model, introduce the PP-LCNet network architecture, the SDDS module, and the internal scale feature interaction module based on attention to construct the LCSD-YOLO network model;

[0007] Based on the LCSD-YOLO network model, perform target recognition and detection on the preprocessed image data set to obtain target recognition and detection results.

[0008] Further, the step of obtaining target image data and performing data preprocessing to obtain a preprocessed image data set specifically includes:

[0009] Collect images of the target object through a camera to obtain a target image data set;

[0010] Name the target image data set according to the format of the Pascal VOC data set, and create a target image data file with three folders: Annotations, ImageSets, and JPEGImages;

[0011] Perform data marking processing on the target image data file to obtain a preprocessed image data set.

[0012] Further, the LCSD-YOLO network model specifically includes a backbone network module, a neck network, and a head network. The backbone network module, the neck network, and the head network are connected in sequence, where:

[0013] The backbone network module adopts the PP-LCNet network architecture. The backbone network module includes a first convolution module, an SDDS module, a depthwise separable convolution module, and an attention-based internal scale feature interaction module. The SDDS module includes a first SDDS module, a second SDDS module, and a third SDDS module. The depthwise separable convolution module includes a first depthwise separable convolution module, a second depthwise separable convolution module, and a third depthwise separable convolution module;

[0014] The neck network includes a DySample module, a Concat module, a C3K2 module, an SE module, and a convolution module. The DySample module includes a first DySample module and a second DySample module. The Concat module includes a first Concat module, a second Concat module, a third Concat module, and a fourth Concat module. The first C3K2 module, the second C3K2 module, the third C3K2 module, and the fourth C3K2 module. The convolution module includes a second convolution module and a third convolution module;

[0015] The head network includes a first detection head, a second detection head, and a third detection head.

[0016] Further, the LCSD-YOLO network model constructs a target recognition loss function by introducing a non-monotonic focusing mechanism. The expression of the target recognition loss function is specifically as follows:

[0017] L WIoU =r×L WIoUv1

[0018]

[0019] In the above formula, L WIoU represents the target recognition loss function, r represents the non-monotonic focusing coefficient, and L WIoUv1 represents the metric function of the overlap degree between the predicted box and the ground truth box, β represents WIoUv3 with dynamic non-monotonic FM applied to WIoUv1, and α and δ represent variable hyperparameters. represents the monotonic focusing coefficient, and L IoU represents the intersection ratio loss function between the predicted box and the ground truth box, and IoU represents the intersection ratio between the predicted box and the ground truth box.

[0020] Further, the step of performing target recognition and detection on the preprocessed image dataset based on the LCSD-YOLO network model to obtain the target recognition and detection result specifically includes:

[0021] Input the preprocessed image dataset into the LCSD-YOLO network model;

[0022] Based on the backbone network module of the LCSD-YOLO network model, perform feature recognition processing on the preprocessed image dataset to obtain target image features;

[0023] Based on the neck network of the LCSD-YOLO network model, perform multi-scale feature fusion processing on the target image features to obtain the fused target image features;

[0024] Based on the head network of the LCSD-YOLO network model, perform decoding prediction on the fused target image features to obtain the target recognition and detection result.

[0025] Further, the step of performing feature recognition processing on the preprocessed image dataset based on the backbone network module of the LCSD-YOLO network model to obtain target image features specifically includes:

[0026] Input the preprocessed image dataset into the backbone network module of the LCSD-YOLO network model;

[0027] Based on the first convolution module of the backbone network module, perform channel adjustment and feature extraction on the preprocessed image dataset to obtain a preliminary target feature image;

[0028] Based on the SDDS module of the backbone network module, perform convolution and spatial channel dimension conversion processing on the preliminary target feature image to obtain a first intermediate target feature image;

[0029] Based on the depthwise separable convolution module of the backbone network module, perform depthwise separable convolution processing on the intermediate target feature image to obtain a second intermediate target feature image;

[0030] The attention-based internal scale feature interaction module based on the backbone network module performs multi-scale feature interaction redundancy processing on the second intermediate target feature image to obtain the target image feature.

[0031] Further, the step of the SDDS module based on the backbone network module performing convolution and spatial-channel dimension conversion processing on the preliminary target feature image to obtain the first intermediate target feature image specifically includes:

[0032] Input the preliminary target feature image into the SDDS module of the backbone network module, and the SDDS module includes a depthwise convolution module, a pointwise convolution module, and a spatial-depth conversion module;

[0033] Based on the depthwise convolution module of the SDDS module, perform depthwise convolution processing on the preliminary target feature image to obtain the first convolution target feature image;

[0034] Based on the pointwise convolution module of the SDDS module, perform channel aggregation processing on the first convolution target feature image to obtain the second convolution target feature image;

[0035] Based on the spatial-depth conversion module of the SDDS module, connect the second convolution target feature image along the channel dimension to obtain the first intermediate target feature image.

[0036] Further, the step of the neck network based on the LCSD-YOLO network model performing multi-scale feature fusion processing on the target image feature to obtain the fused target image feature specifically includes:

[0037] Input the target image feature into the neck network of the LCSD-YOLO network model;

[0038] Based on the DySample module of the neck network, perform feature reconstruction processing on the target image feature to obtain the reconstructed target image feature;

[0039] Based on the Concat module of the neck network, perform feature splicing processing on the reconstructed target image feature to obtain the spliced target image feature;

[0040] Based on the C3K2 module of the neck network, perform multi-level spatial feature extraction processing on the spliced target image feature to obtain the deeper target image feature;

[0041] Based on the SE module of the neck network, perform channel feature weight adjustment processing on the deeper target image feature to obtain the adjusted target image feature;

[0042] The convolution module based on the neck network performs convolution processing on the adjusted target image features and outputs the fused target image features.

[0043] The second technical solution adopted by the present invention is: a target detection system based on the LCSD-YOLO network, including:

[0044] The first module is used to obtain target image data and perform data preprocessing to obtain a preprocessed image data set;

[0045] The second module is used to construct the LCSD-YOLO network model based on the YOLO11 network model by introducing the PP-LCNet network architecture, the SDDS module, and the internal scale feature interaction module based on attention;

[0046] The third module is used to perform target recognition and detection on the preprocessed image data set based on the LCSD-YOLO network model to obtain the target recognition and detection results.

[0047] The beneficial effects of the method and system of the present invention are: by obtaining target image data and performing data preprocessing, and then based on the YOLO11 network model, introducing the PP-LCNet network architecture, the SDDS module, and the internal scale feature interaction module based on attention to construct the LCSD-YOLO network model, using the lightweight backbone network PP-LCNet to replace the original yolo11n backbone network can reduce the model complexity and reduce the computational resource requirements. Furthermore, by integrating the depthwise separable downsampling convolution SDDS with spatial depth conversion, the loss of fine-grained information is reduced and the learning efficiency of feature representation is improved, significantly improving the problem of insufficient small target detection ability caused by the impaired feature expression ability in directly compressing the model. Further, through the internal scale feature interaction module based on attention for in-scale interaction, while improving the model's feature extraction ability, the computational cost is further reduced, and effectively solves the problem of insufficient small target detection accuracy caused by the damaged feature pyramid in the lightweight model. Description of the Drawings

[0048] Figure 1 is the flowchart of the steps of a target detection method based on the LCSD-YOLO network of the present invention;

[0049] Figure 2 is the structural block diagram of a target detection system based on the LCSD-YOLO network of the present invention;

[0050] Figure 3 is the structural schematic diagram of the LCSD-YOLO network model provided by a specific embodiment of the present invention;

[0051] Figure 4It is a schematic diagram of the parameters of the PP-LCNet network architecture provided by a specific embodiment of the present invention;

[0052] Figure 5 It is an example schematic diagram of depthwise convolution provided by a specific embodiment of the present invention;

[0053] Figure 6 It is an example schematic diagram of pointwise convolution provided by a specific embodiment of the present invention;

[0054] Figure 7 It is an example schematic diagram of spatial-depth conversion provided by a specific embodiment of the present invention;

[0055] Figure 8 It is an example schematic diagram of the DySample module provided by a specific embodiment of the present invention;

[0056] Figure 9 It is a schematic diagram of the target recognition and detection result provided by a specific embodiment of the present invention. Detailed implementation manners

[0057] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. For the step numbers in the following embodiments, they are only set for the convenience of description and explanation, and no limitation is made on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.

[0058] First of all, it should be noted that the embodiments of the present invention are directed to target detection under high altitude. In the following detailed implementation manners, the embodiments of the present invention will be described by taking the intelligent recognition of whether a safety helmet is worn by high-altitude workers as an example.

[0059] Referring to Figure 1 , the present invention provides a target detection method based on the LCSD-YOLO network, and the method includes the following steps:

[0060] S100. Obtain target image data and perform data preprocessing to obtain a preprocessed image data set;

[0061] Specifically, image acquisition of the target object is performed through a camera to obtain a target image data set; the target image data set is named according to the format of the Pascal VOC data set, and a target image data file with three folders, namely Annotations, ImageSets, and JPEGImages, is created; data marking processing is performed on the target image data file to obtain a preprocessed image data set.

[0062] In this embodiment, a self-made dataset is adopted. The experimenter uses an industrial camera to collect images of the target object, names the images in the format of the Pascal VOC dataset, and creates three folders named Annotations, ImageSets, and JPEGImages at the same time. The target in the collected images is marked using the image annotation tool LabelImg, and the class labels are human and safety helmet respectively.

[0063] Since this dataset cannot meet the requirement of 2000 images needed to recognize two-category targets, the Augmentor image data augmentation library is used to augment the images. The experimenter selects the saving path of the images and the path of the marked information XML file, specifies the output paths of the augmented images and XML files, selects the required image augmenters such as brightness, cropping, Gaussian noise, etc., selects the number of augmentations and the augmentation method (sequential, combined, random, etc.) to augment the images to meet the recognition requirements.

[0064] Since the YOLO series neural network is used in the embodiment of the present invention, the augmented images and marked files need to be converted into the YOLO format and then the dataset is divided into a training set and a validation set, which are 75% and 25% respectively.

[0065] Furthermore, an experimental environment is constructed. After preprocessing the dataset, a dataset configuration file corresponding to the dataset is constructed next. In the dataset configuration file, the paths of the training set and the validation set and the class information in the dataset need to be written. Subsequently, the default.yaml file of YOLO11 is modified according to the experimental requirements, and the input image size is adjusted to 640×640. The training rounds of this model are 400, the batchsize is 64, the number of detected object categories is 2, the probability of using the Mosaic data augmentation method is 1, and the probability of using the copy-paste data augmentation method is 0.5. The environment information is as follows: cuda11.8, deep learning framework pytorch2.3.0, Intel core i9-13900ks CPU, 128G memory, GPU is NVIDIA GeForce RTX 4090, and the video memory is 24G.

[0066] S200. Based on the YOLO11 network model, the PP-LCNet network architecture, the SDDS module and the internal scale feature interaction module based on attention are introduced to construct the LCSD-YOLO network model;

[0067] Specifically, as Figure 3 shown, the LCSD-YOLO network model specifically includes a backbone network module, a neck network and a head network, and the backbone network module, the neck network and the head network are connected in sequence, where:

[0068] The backbone network module adopts the PP-LCNet network architecture. The backbone network module includes a first convolution module, an SDDS module, a depthwise separable convolution module, and an attention-based internal scale feature interaction module. The SDDS module includes a first SDDS module, a second SDDS module, and a third SDDS module. The depthwise separable convolution module includes a first depthwise separable convolution module, a second depthwise separable convolution module, and a third depthwise separable convolution module;

[0069] The neck network includes a DySample module, a Concat module, a C3K2 module, an SE module, and a convolution module. The DySample module includes a first DySample module and a second DySample module. The Concat module includes a first Concat module, a second Concat module, a third Concat module, and a fourth Concat module. The first C3K2 module, the second C3K2 module, the third C3K2 module, and the fourth C3K2 module. The convolution module includes a second convolution module and a third convolution module;

[0070] The head network includes a first detection head, a second detection head, and a third detection head.

[0071] In this embodiment, the LCSD-YOLO network model is based on the YOLO11 network architecture. The overall structure consists of three parts: a backbone network, a neck network, and a head network. First, by replacing the backbone network with the lightweight network PP-LCNet, the complexity and the number of parameters of the model are significantly reduced with a small loss of accuracy. To further optimize the performance, an SDDS module is introduced in PP-LCNet to replace the original downsampling convolution, and the last layer of DepthSepConv is replaced with an AIFI module, thus enhancing the feature extraction ability while improving the model inference speed, especially significantly improving the detection accuracy of small targets. Second, in the neck network, a lightweight dynamic upsampler DySample is used to replace the traditional nn.UPSample, significantly improving the performance of the model with almost no increase in model complexity. An SE attention mechanism is added to the small feature layer of the original YOLO11 neck network to further improve the detection accuracy of small targets. Finally, the head network is responsible for the target classification and bounding box regression tasks, and outputs the class prediction and position prediction results of the target object by further processing the feature map transmitted by the neck network.

[0072] In summary, the LCSD-YOLO network model is based on the YOLO11 network. The YOLO11 network includes a backbone network, a neck network, and a head network. The backbone network is replaced with PP-LCNet (a lightweight CPU network based on the MKLDNN acceleration strategy), and its parameters are as Figure 4 shown. The downsampling convolution in the replaced backbone network is replaced with the SDDS module (a depthwise separable downsampling convolution that integrates spatial depth conversion). The last layer in the replaced backbone network is replaced with the attention-based internal scale feature interaction module AIFI. The upsampler nn.UPSample in the original YOLO11 neck network is replaced with a more lightweight dynamic upsampler DySample, and a layer of SE attention mechanism is added to the small feature layer of the original YOLO11 neck network.

[0073] Finally, in this embodiment, the original loss function is replaced with a non-monotonic focusing mechanism that measures the WIoUv3 of the target by evaluating the outlier degree of the anchor box, so as to avoid the accuracy impact caused by too high or too low dataset quality.

[0074] WIoUv1 introduces distance as an attention metric. When the target box and the prediction box overlap within a certain range, reducing the penalty of the geometric metric enables the model to obtain better generalization ability. The formula for calculating WIoUv1 is as follows:

[0075] L WIoUv1 = R WIoU × L IoU

[0076]

[0077] L IoU = 1 - IoU

[0078] In the above formula, represents the x-axis position of the center of the ground truth box, represents the y-axis position of the center of the ground truth box, b cx represents the x-axis position of the center of the prediction box, b cy represents the y-axis position of the center of the prediction box, c w and c h represent the height and width of the minimum bounding box formed by the Prediction Box and the Real Box.

[0079] Furthermore, a non-monotonic focusing coefficient r is constructed using β and applied to WIoUv1 to obtain WIoUv3 with dynamic non-monotonic FM. By using a reasonable gradient gain allocation strategy, the weights of high-quality and low-quality bounding boxes are reduced, enabling the model to focus on average-quality samples, thereby improving performance. The expression for calculating WIoUv1 is specifically as follows:

[0080] L WIoU = r × L WIoUv1

[0081]

[0082] In the above formula, L WIoU represents the target recognition loss function, r represents the non-monotonic focusing coefficient, and L WIoUv1 represents the metric function for the overlap degree between the predicted bounding box and the ground truth bounding box, β represents WIoUv3 with dynamic non-monotonic FM applied to WIoUv1, and α and δ represent variable hyperparameters. represents the monotonic focusing coefficient, and L IoU represents the intersection ratio loss function between the predicted bounding box and the ground truth bounding box, and IoU represents the intersection ratio between the predicted bounding box and the ground truth bounding box.

[0083] S300. Perform target recognition detection on the preprocessed image dataset based on the LCSD-YOLO network model to obtain the target recognition detection result.

[0084] Specifically, input the preprocessed image dataset into the LCSD-YOLO network model; based on the backbone network module of the LCSD-YOLO network model, perform feature recognition processing on the preprocessed image dataset to obtain the target image features; based on the neck network of the LCSD-YOLO network model, perform multi-scale feature fusion processing on the target image features to obtain the fused target image features; based on the head network of the LCSD-YOLO network model, perform decoding prediction on the fused target image features to obtain the target recognition detection result.

[0085] First, use the self-made building worker safety helmet dataset with category labels of human and safety helmet images as the training set of the model, and divide it into a training set and a test set. There are 3000 training set images and 1000 validation set images. Import the improved LCSD-YOLO network model in the pre-configured deep learning environment, modify the parameter configuration file in the network model according to the actual environment information, and then use the processed dataset to train the model. The number of training rounds is 400, and 64 images are input in each round. Observe the training log in real time through Wandb during the training process, and save the weight file after the training is completed.

[0086] After the model training is completed, a corresponding weight file will be generated. Select the best.pt weight file, the image to be detected, and the label, and import them into the test program. After the program is executed, the detection effect diagram will be output. Evaluate whether the experiment meets the expected requirements by analyzing these images.

[0087] Further, for the backbone network module, the preprocessed image dataset is input into the backbone network module of the LCSD-YOLO network model; based on the first convolutional module of the backbone network module, channel adjustment and feature extraction are performed on the preprocessed image dataset to obtain a preliminary target feature image; based on the SDDS module of the backbone network module, convolution and spatial-channel dimension conversion processing are performed on the preliminary target feature image to obtain a first intermediate target feature image; based on the depthwise separable convolutional module of the backbone network module, depthwise separable convolution processing is performed on the intermediate target feature image to obtain a second intermediate target feature image; based on the attention-based internal scale feature interaction module of the backbone network module, multi-scale feature interaction redundancy processing is performed on the second intermediate target feature image to obtain target image features.

[0088] Furthermore, for the SDDS module in the backbone network module, the preliminary target feature image is input into the SDDS module of the backbone network module. The SDDS module includes a depthwise convolution module, a pointwise convolution module, and a spatial-depth conversion module; based on the depthwise convolution module of the SDDS module, depthwise convolution processing is performed on the preliminary target feature image to obtain a first convolution target feature image; based on the pointwise convolution module of the SDDS module, channel aggregation processing is performed on the first convolution target feature image to obtain a second convolution target feature image; based on the spatial-depth conversion module of the SDDS module, the second convolution target feature image is connected along the channel dimension to obtain a first intermediate target feature image.

[0089] For the neck network, the target image features are input into the neck network of the LCSD-YOLO network model; based on the DySample module of the neck network, feature reconstruction processing is performed on the target image features to obtain reconstructed target image features; based on the Concat module of the neck network, feature splicing processing is performed on the reconstructed target image features to obtain spliced target image features; based on the C3K2 module of the neck network, multi-level spatial feature extraction processing is performed on the spliced target image features to obtain deeper target image features; based on the SE module of the neck network, channel feature weight adjustment processing is performed on the deeper target image features to obtain adjusted target image features; based on the convolutional module of the neck network, convolution processing is performed on the adjusted target image features to output fused target image features.

[0090] In this embodiment, the proportion of target pixels in the images of aerial working personnel collected by the ground camera looking up is low (usually less than 0.1% of the image area). After improvement, the input image first undergoes a 3×3 convolution for channel adjustment and feature extraction. Subsequently, the feature map is downsampled through the SDDS layer that fuses the depthwise separable downsampling convolution with spatial depth conversion. The depthwise separable convolution consists of a depthwise convolution and a pointwise convolution. The depthwise convolution is used to extract spatial features, and the pointwise convolution is used to extract channel features. The depthwise separable convolution performs grouped convolution in the feature dimension, performs independent depthwise convolution on each channel, and aggregates all channels using a 1x1 convolution before output, as Figure 5 and Figure 6 shown. Since grouped convolution is required, the number of all input channels should always be a multiple of the number of output channels. To accelerate the model's feature extraction ability, the present invention incorporates spatial depth conversion. By connecting the feature map X obtained from the depthwise separable convolution, as Figure 7 shown, the default scale of this module is 2 to obtain 4 submaps f 0,0 、f 1,0 、f 0,1 and f 1,1 each of which has the shape (S / 2, S / 2, C1) and downsamples X by a factor of 2. Next, these sub-feature maps are concatenated along the channel dimension to obtain a feature map X′, whose spatial dimension is reduced by a scale factor and the channel dimension is increased by a scale factor of 2. In other words, SPD converts the feature map X(S, S, C1) into an intermediate feature map X′(S / sacle, S / sacle, scale 2C1) and can further improve the quality of the feature map through batch normalization, making subsequent operations easier to capture effective patterns. At the same time, combining with the Hardswish activation function can ensure that the transformed feature map still maintains good statistical characteristics, promoting the maximization of the overall architecture efficiency. The attention-based internal scale feature interaction module AIFI reduces the computational cost of feature extraction by reducing the redundancy of multi-level feature extraction through feature extraction only on the third depthwise separable convolution module, avoiding the situation that the high-level features extracted from low-level features contain too much rich semantic information about the object and the feature interaction redundancy of cascaded multi-scale features. By using a single-scale Transformer encoder for intra-scale interaction only on the third depthwise separable convolution module, the computational cost is further reduced. Subsequently, the feature map passes through the dynamic upsampling layer DySample. Based on the essential point sampling of upsampling, content-aware sampling points are generated, where the point offsets are generated by linear projection and the point values are resampled using the grid_sample function of PyTorch. The upsampling of the feature map is gradually optimized by controlling the initial sampling position, adjusting the offset range, and grouping the upsampling process. That is, the upsampling of the feature map is gradually optimized by controlling the initial sampling position, adjusting the offset range, and grouping the upsampling process. In addition, free convolution kernels are only enabled in the fourth C3K2 layer, which improves the extraction ability of the feature map without adding too much invalid computational amount. The optimization effect of the free convolution kernel on feature extraction is not obvious in the low layer.

[0091] As Figure 8 shown, the DySample dynamic upsampling module assumes the input feature, upsampled feature, generated offset, original grid, and sigmoid activation function. They are represented by X, X′, 0, G, and σ respectively. Figure 8 (a) in shows that the sampling set is generated by the sampling point generator, and the input feature is resampled through the grid sample function. Figure 8 (b) in shows that in the generator, the sampling set is the sum of the generated offset and the original grid position. The upper box shows the version with a "static range factor", where the offset is generated by a linear layer. The lower part describes the version with a "dynamic range factor", where the range factor is first generated and then used to modulate the offset.

[0092] In summary, the present invention replaces the original YOLO11n backbone network (a lightweight network using depthwise separable convolutions) with the lightweight backbone network PP-LCNet, reducing the model complexity and the computational resource requirements. Moreover, the original downsampling convolution in the PP-LCNet backbone network is replaced with a depthwise separable downsampling convolution SDDS that integrates spatial depth transformation, reducing the loss of fine-grained information and improving the learning efficiency of feature representation, significantly improving the problem of insufficient small object detection ability caused by impaired feature expression ability in directly compressing the model. At the same time, the Hardswish activation function and batch normalization are used, greatly improving the inference speed of the model and meeting the requirements of real-time detection at the edge. The SE mechanism is enabled to enhance the model's ability to extract target features by dynamically adjusting the channel weights, while the AIFI module only performs scale-internal interaction on the third depthwise separable convolution module, further reducing the computational cost while improving the model's feature extraction ability, and effectively solving the problem of insufficient small object detection accuracy caused by damaged feature pyramids in lightweight models. The dynamic upsampling layer DySample is adopted to further improve the model performance with almost no change to the original model complexity. This technical solution has the advantages of high computational efficiency, strong feature extraction ability, flexibility and scalability, and is very suitable for application in the safety guarantee of construction site personnel, meeting the requirements of high precision, real-time performance and low latency in practical applications. The recognition results are as Figure 9 shown.

[0093] Refer to Figure 2 , a target detection system based on the LCSD-YOLO network, comprising:

[0094] The first module 201 is used to obtain target image data and perform data preprocessing to obtain a preprocessed image data set;

[0095] The second module 202 is used to construct the LCSD-YOLO network model based on the YOLO11 network model by introducing the PP-LCNet network architecture, the SDDS module and the internal scale feature interaction module based on attention;

[0096] The third module 203 is used to perform target recognition and detection on the preprocessed image data set based on the LCSD-YOLO network model to obtain target recognition and detection results.

[0097] The content in the above method embodiments is applicable to the present system embodiment. The functions specifically implemented by the present system embodiment are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.

[0098] The above is a detailed description of the preferred embodiment of the present invention. However, the present invention is not limited to the described embodiment. Those skilled in the art can make various equivalent deformations or substitutions without departing from the spirit of the present invention, and these equivalent deformations or substitutions are all included within the scope defined by the claims of this application.

Claims

1. A target detection method based on the LCSD-YOLO network, characterized in that, It includes the following steps: Obtain target image data and perform data preprocessing to obtain a preprocessed image data set; Based on the YOLO11 network model, introduce the PP-LCNet network architecture, the SDDS module, and the attention-based internal scale feature interaction module to construct the LCSD-YOLO network model; Based on the LCSD-YOLO network model, perform target recognition detection on the preprocessed image data set to obtain target recognition detection results.

2. The object detection method based on the LCSD-YOLO network according to claim 1, wherein, The step of obtaining target image data and performing data preprocessing to obtain a preprocessed image data set specifically includes: Collect images of the target object through a camera to obtain a target image data set; Name the target image data set according to the format of the Pascal VOC data set, and create a target image data file with three folders: Annotations, ImageSets, and JPEGImages; Perform data marking processing on the target image data file to obtain a preprocessed image data set.

3. The object detection method based on the LCSD-YOLO network according to claim 2, wherein, The LCSD-YOLO network model specifically includes a backbone network module, a neck network, and a head network. The backbone network module, the neck network, and the head network are connected in sequence, where: The backbone network module adopts the PP-LCNet network architecture. The backbone network module includes a first convolution module, an SDDS module, a depthwise separable convolution module, and an attention-based internal scale feature interaction module. The SDDS module includes a first SDDS module, a second SDDS module, and a third SDDS module. The depthwise separable convolution module includes a first depthwise separable convolution module, a second depthwise separable convolution module, and a third depthwise separable convolution module; The neck network includes a DySample module, a Concat module, a C3K2 module, an SE module, and a convolution module. The DySample module includes a first DySample module and a second DySample module. The Concat module includes a first Concat module, a second Concat module, a third Concat module, and a fourth Concat module. The first C3K2 module, the second C3K2 module, the third C3K2 module, and the fourth C3K2 module. The convolution module includes a second convolution module and a third convolution module; The head network includes a first detection head, a second detection head, and a third detection head.

4. The object detection method based on the LCSD-YOLO network according to claim 3, characterized in that, The LCSD-YOLO network model constructs a target recognition loss function by introducing a non-monotonic focusing mechanism. The expression of the target recognition loss function is specifically as follows: L WIoU = r × L WIoUv1 In the above formula, L WIoU represents the target recognition loss function, r represents the non-monotonic focusing coefficient, L WIoUv1 represents the metric function of the overlap degree between the predicted box and the ground truth box, β represents WIoUv3 with dynamic non-monotonic FM applied to WIoUv1, and α and δ represent variable hyperparameters, represents the monotonic focusing coefficient, L IoU represents the intersection ratio loss function of the predicted box and the ground truth box, and IoU represents the intersection ratio of the predicted box and the ground truth box.

5. The object detection method based on the LCSD-YOLO network according to claim 4, wherein The step of performing target recognition detection on the preprocessed image data set based on the LCSD-YOLO network model to obtain target recognition detection results specifically includes: Input the preprocessed image data set into the LCSD-YOLO network model; Based on the backbone network module of the LCSD-YOLO network model, perform feature recognition processing on the preprocessed image data set to obtain target image features; The neck network based on the LCSD-YOLO network model performs multi-scale feature fusion processing on the target image features to obtain the fused target image features; The head network based on the LCSD-YOLO network model decodes and predicts the fused target image features to obtain the target recognition and detection results.

6. The object detection method based on the LCSD-YOLO network according to claim 5, wherein, The step of the backbone network module based on the LCSD-YOLO network model performing feature recognition processing on the preprocessed image dataset to obtain the target image features specifically includes: Inputting the preprocessed image dataset into the backbone network module of the LCSD-YOLO network model; Based on the first convolution module of the backbone network module, performing channel adjustment and feature extraction on the preprocessed image dataset to obtain a preliminary target feature image; Based on the SDDS module of the backbone network module, performing convolution and spatial-channel dimension conversion processing on the preliminary target feature image to obtain a first intermediate target feature image; Based on the depthwise separable convolution module of the backbone network module, performing depthwise separable convolution processing on the intermediate target feature image to obtain a second intermediate target feature image; Based on the attention-based internal scale feature interaction module of the backbone network module, performing multi-scale feature interaction redundancy processing on the second intermediate target feature image to obtain the target image features.

7. The object detection method based on the LCSD-YOLO network according to claim 6, wherein The step of the SDDS module based on the backbone network module performing convolution and spatial-channel dimension conversion processing on the preliminary target feature image to obtain a first intermediate target feature image specifically includes: Inputting the preliminary target feature image into the SDDS module of the backbone network module, and the SDDS module includes a depthwise convolution module, a pointwise convolution module, and a spatial-depth conversion module; Based on the depthwise convolution module of the SDDS module, performing depthwise convolution processing on the preliminary target feature image to obtain a first convolution target feature image; Based on the pointwise convolution module of the SDDS module, performing channel aggregation processing on the first convolution target feature image to obtain a second convolution target feature image; Based on the spatial-depth conversion module of the SDDS module, connecting the second convolution target feature image along the channel dimension to obtain a first intermediate target feature image.

8. The object detection method based on the LCSD-YOLO network according to claim 7, characterized in that, The step of the neck network based on the LCSD-YOLO network model performing multi-scale feature fusion processing on the target image features to obtain the fused target image features specifically includes: Inputting the target image features into the neck network of the LCSD-YOLO network model; Based on the DySample module of the neck network, performing feature reconstruction processing on the target image features to obtain the reconstructed target image features; Based on the Concat module of the neck network, performing feature splicing processing on the reconstructed target image features to obtain the spliced target image features; Based on the C3K2 module of the neck network, performing multi-level spatial feature extraction processing on the spliced target image features to obtain deeper target image features; Based on the SE module of the neck network, the channel feature weights of the deeper target image features are adjusted to obtain the adjusted target image features; Based on the convolution module of the neck network, the adjusted target image features are convolved to output the fused target image features.

9. A target detection system based on the LCSD-YOLO network, characterized in that, It includes the following modules: The first module is used to obtain the target image data and perform data preprocessing to obtain the preprocessed image data set; The second module is used to build the LCSD-YOLO network model based on the YOLO11 network model, introducing the PP-LCNet network architecture, the SDDS module and the internal scale feature interaction module based on attention; The third module is used to perform target recognition and detection on the preprocessed image data set based on the LCSD-YOLO network model to obtain the target recognition and detection results.

Citation Information

Patent Citations

  • Lightweight YOLO model construction method for inland ship detection

    CN118747849A

  • Lightweight dense pedestrian detection method and system based on YOLOv8

    CN119229374A

  • Lightweight parking detection method based on multi-scale attention mechanism

    CN119314141A

  • Synthetic aperture radar (SAR) image target detection method

    US20230169623A1