Mask target detection method and system based on improved YOLOWorld world model

The improved YOLOWorld model enhances mask detection by addressing inaccuracies in complex backgrounds and lighting conditions, ensuring real-time performance and robust recognition across varied scenarios.

CN120318754APending Publication Date: 2025-07-15SHENZHEN POLYTECHNIC
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510380120.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The existing mask detection methods have high error detection rates in complex contexts, difficult to achieve real-time and accuracy balance, and limited model recognition capabilities, especially when running on embedded devices with limited resources, they have poor detection results.

Method used

The improved YOLOWorld world model is adopted to carry out mask target detection through data preprocessing, image segmentation and transmittance calculation, combined with the improved SPPFCSPC module and feature fusion and output module, including size transformation, normalization, data augmentation, image segmentation, transmittance calculation and feature pyramid network structure design.

Benefits of technology

It improves the model's adaptability and detection accuracy in complex contexts, reduces computing overhead, enhances the recognition ability of masks of different postures and angles, improves generalization ability and computing efficiency, and is suitable for resource-constrained equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318754A_ABST
    Figure CN120318754A_ABST
Patent Text Reader

Abstract

The invention discloses a mask target detection method based on an improved YOLOWorld world model, and the method comprises the following steps: obtaining a to-be-detected mask image, carrying out the data preprocessing of the obtained to-be-detected mask image, so as to obtain a mask image after the data preprocessing, carrying out the image processing of the mask image after the data preprocessing, and obtaining a mask target. And performing image processing on the mask target detection model to obtain a mask image after image processing, and inputting the obtained mask image after image processing into a pre-trained mask target detection model to obtain a final detection result. The technical problem that the false detection problem of an existing color space conversion method under a complex background is prominent can be solved, the technical problem that an existing morphological processing method is difficult to achieve balance between real-time performance and precision can be solved, and the technical problem that an existing template matching method is limited in model recognition capacity can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision, and more specifically, relates to a mask target detection method and system based on an improved YOLOWorld world model. Background Art

[0002] In industrial production, construction, and public place safety supervision, wearing a mask is an important measure to prevent accidental injuries and disease transmission. Whether individuals are wearing masks can be quickly and accurately identified, which is crucial for ensuring personnel safety and improving supervision efficiency. Therefore, an efficient and accurate mask detection method has important theoretical significance and practical application value.

[0003] Existing mask detection methods mainly include three types: color space conversion method, morphological processing method, and template matching method. Color space conversion highlights mask features by changing the color model of the image; morphological processing is used to remove noise and interference in the image to make the mask area clearer; template matching compares the image to be detected with a pre-defined mask template to determine the position and presence of the mask.

[0004] However, the above several mask detection methods all have some non-negligible defects:

[0005] First, the problem of false detection of the existing color space conversion method in complex backgrounds is relatively prominent: when the light is too strong or too weak, the image quality will significantly decline, resulting in difficult accurate identification of mask features; at the same time, objects similar in shape to the mask will also interfere with the detection results, reducing the detection accuracy.

[0006] Second, it is difficult for the existing morphological processing method to balance real-time performance and accuracy. Especially when running on resource-constrained embedded devices, it is impossible to reduce the computational overhead while maintaining high accuracy;

[0007] Third, the model recognition ability of the existing template matching method is limited: in actual application scenarios, the postures and angles of people vary greatly, and the mask may appear in various different postures, which makes the model prone to missed detection and false detection. Summary of the Invention

[0008] Aiming at the above defects or improvement requirements of the existing technology, the present invention provides a mask target detection method and system based on an improved YOLOWorld world model, aiming to solve the technical problem that the problem of false detection of the existing color space conversion method in complex backgrounds is relatively prominent, as well as the technical problem that it is difficult for the existing morphological processing method to balance real-time performance and accuracy, and the technical problem that the model recognition ability of the existing template matching method is limited.

[0009] To achieve the above object, according to one aspect of the present invention, a mask target detection method based on an improved YOLOWorld world model is provided, including the following steps:

[0010] (1) Obtain the mask image to be detected;

[0011] (2) Perform data preprocessing on the mask image to be detected obtained in step (1) to obtain the mask image after data preprocessing;

[0012] (3) Perform image processing on the mask image after data preprocessing in step (2) to obtain the mask image after image processing;

[0013] (4) Input the mask image after image processing obtained in step (3) into a pre-trained mask target detection model to obtain the final detection result.

[0014] Preferably, the data preprocessing operation in step (2) includes sequentially performing size transformation processing, normalization processing, and data augmentation processing on the mask image, where the size transformation processing is to scale the mask image to a fixed size of 512 pixels × 512 pixels; the normalization processing is to map each pixel value in the mask image after size transformation processing to the range of 0 to 1; the data augmentation processing is to perform one or any combination of random rotation, mirroring, and HSV perturbation processing on the mask image after normalization processing.

[0015] Step (3) includes the following sub-steps:

[0016] (3-1) Divide the mask image after data preprocessing in step (2) into the current region and other regions;

[0017] (3-2) Obtain the atmospheric light value of the mask image after data preprocessing in step (2) according to the current region obtained in step (3-1);

[0018] (3-3) Obtain the transmittance of the mask image after data preprocessing in step (2) according to the transmittance of the current region and other regions obtained in step (3-1);

[0019] (3-4) Perform defogging processing on the mask image after data preprocessing according to the transmittance of the mask image after data preprocessing obtained in step (3-3) to obtain the mask image I after image processing * 。

[0020] The final detection result obtained in step (4) exists in the form of a detection box, and each detection box marks the predicted mask target position and mask target category.

[0021] Preferably, step (3-1) includes the following sub-steps:

[0022] (3-1-1) Pre-segment the mask image after data preprocessing in step (2) to obtain multiple image sub-regions, and each image sub-region is used as a superpixel point;

[0023] (3-1-2) Use the K-means clustering algorithm to cluster all the superpixel points obtained in step (3-1-1) to obtain the current region and other regions, obtain the gray values of all the superpixel points obtained in step (3-1-1), and take the average of all the gray values as the adaptive threshold. All the superpixel points with gray values greater than or equal to the adaptive threshold among all the superpixel points obtained in step (3-1-1) are classified into the current region, and all the superpixel points with gray values less than the adaptive threshold among all the superpixel points obtained in step (3-1-1) are classified into other regions.

[0024] Step (3-2) is specifically as follows: First, sort the gray values of all the superpixel points in the current region in descending order. Then, obtain the top k% of the superpixel points from the sorting result, and obtain the corresponding superpixel points of these superpixel points in the mask image after data preprocessing in step (2). Subsequently, take the average of the gray values of all the corresponding superpixel points as the atmospheric light value of the mask image after data preprocessing in step (2). That is, the atmospheric light value is a three-element vector, and each element corresponds to one of the R, G, and B color channels, where the value range of k is from 1 to 5, and preferably 1.

[0025] Step (3-3) is specifically as follows. First, for the current region, the transmittance t of the current region is calculated by the following formula f :

[0026]

[0027] where I f is the pixel value of the current region, A is the atmospheric light value, β is the medium scattering coefficient, and d f is the depth information of the current region.

[0028] Then, for other regions, the transmittance t of the other regions is calculated by the following formula b :

[0029]

[0030] where I b is the pixel value of the other region, and d b is the depth information of the other region.

[0031] Subsequently, for the transmittance t f of the current region and the transmittance t bPerform weighted averaging to obtain the transmittance t of the mask image after data preprocessing in step (2):

[0032]

[0033] where S f is the area of the current region, and S b is the area of other regions.

[0034] Step (3-4) is calculated using the following formula:

[0035]

[0036] where I is the mask image after data preprocessing obtained in step (3-3), and t0 is the minimum value of the preset transmittance, whose value range is from 0 to 0.1, preferably 0.05.

[0037] Preferably, the mask target detection model uses an improved YOLOWorld model, which includes a backbone network, a Spatial Pyramid Pooling SPPFCSPC module, a Cross-Stage Partial Connection branch module, and a Feature Fusion and Output module;

[0038] The specific structure of the backbone network is:

[0039] The first layer is a convolutional layer, whose input is an image with a dimension of 512*512*3. It performs a Convolution+Batch Normalization+SiLU (CBS) operation on this image, where the parameters are: the number of output channels is 64, the convolution kernel size is 3, the stride is 2, the padding value is 1, and the output dimension is a feature map of 256×256×32.

[0040] The second layer is a convolutional layer, whose input is the feature map with a dimension of 256×256×32 output from the first layer. It performs a CBS operation on this feature map, where the parameters are: the number of output channels is 64, the convolution kernel size is 3, the stride is 2, the padding value is 1, and the output dimension is a feature map of 128×128×64.

[0041] The second layer is a convolutional layer, whose input is the feature map with a dimension of 128×128×64 output from the second layer. It performs a CBS operation on this feature map, where the parameters are: the number of output channels is 128, the convolution kernel size is 3, the stride is 1, the padding value is 1, and the output dimension is a feature map of 128×128×128.

[0042] The input of the SPPFCSPC module is the feature map with a dimension of 128×128×64 output by the second layer of the backbone network. It first performs three CBS operations on the input feature map to obtain a feature map x1 with a dimension of 128×128×64. The specific parameters of the CBS operation are: convolution kernel size 3, stride 1, padding value 1, and output channel number 64. Then, it performs two-level max pooling on the feature map x1 to generate a feature map x2 with a dimension of 128×128×64, a feature map x3 with a dimension of 128×128×64, and a feature map m(x3) with a dimension of 128×128×64. Subsequently, it concatenates the obtained feature maps x1, x2, x3, and m(x3) along the channel dimension to obtain a feature map with a dimension of 128×128×256. Finally, it performs a CBS operation on the obtained feature map with a dimension of 128×128×256 to output a feature map y1 with a dimension of 128×128×64. The specific parameters of the CBS operation are: convolution kernel size 3, stride 1, padding value 1, and output channel number 64.

[0043] The input of the cross-stage partial connection branch module is the feature map output by the SPPFCSPC module. It performs a CBS operation on the input feature map to output a feature map y2 with a dimension of 128×128×64. The specific parameters of the CBS operation are: convolution kernel size 3, stride 1, padding value 1, and output channel number 64.

[0044] Preferably, the specific structure of the feature fusion and output module is:

[0045] The fourth layer is the pooling layer. Its input is the feature map with a dimension of 128×128×64 output by the second layer in the backbone network. It performs max pooling on this feature map to output a feature map with a dimension of 64×64×64.

[0046] The fifth layer is the convolutional layer. Its input is the feature map output by the fourth layer. It performs a CBS operation on this feature map to output a feature map with a dimension of 64×64×128. The specific parameters of the CBS operation are: output channel number 128, convolution kernel size 3×3, stride 1, and padding value 1.

[0047] The sixth layer is the upsampling layer. Its input is the feature map output by the fifth layer. It performs upsampling on this feature map to fuse feature information at different levels and outputs a feature map with a dimension of 128×128×128.

[0048] The seventh layer is the downsampling layer. Its input is the feature map output by the sixth layer. It performs downsampling on this feature map to output a feature map with a dimension of 64×64×256.

[0049] The eighth layer is a multi-scale feature extraction layer. Its input is the feature map output by the seventh layer. It performs multi-scale feature extraction on this feature map to output a feature map with a dimension of 64×64×256.

[0050] The ninth layer is a downsampling layer. Its input is the feature map output by the eighth layer. It performs downsampling on this feature map to output a feature map with a dimension of 32×32×512.

[0051] The tenth layer is a multi-scale feature extraction layer. Its input is the feature map output by the ninth layer. It performs multi-scale feature extraction on this feature map to output a feature map with a dimension of 32×32×512.

[0052] The eleventh layer is a semantic fusion layer. Its input is the feature map output by the tenth layer. It performs multi-scale feature extraction on this feature map to output a feature map with a dimension of 64×64×384.

[0053] The twelfth layer is an upsampling layer. Its input is the feature map output by the eleventh layer. It performs upsampling on this feature map to output a feature map with a dimension of 128×128×384.

[0054] The thirteenth layer is a feature fusion layer. Its input is the feature map output by the twelfth layer. It performs feature fusion on this feature map to output a feature map with a dimension of 64×64×576. The specific parameters for feature fusion are: convolution kernel size 3×3, stride 1, padding value 1.

[0055] The fourteenth layer is a convolutional layer. Its input is the feature map output by the thirteenth layer. It performs CBS operations on this feature map to output a feature map with a dimension of 32×32×512. The specific parameters for the CBS operations are: number of output channels 128, convolution kernel size 3×3, stride 1, padding value 1.

[0056] The fifteenth layer is a semantic fusion layer. Its input is the feature map output by the fourteenth layer. It performs multi-scale feature fusion on this feature map. The specific parameters for multi-scale feature fusion are: convolution kernel size 3×3, stride 1, padding value 1.

[0057] The sixteenth layer is a convolutional layer. Its input is the feature map output by the fifteenth layer. It performs CBS operations on this feature map to output a feature map with a dimension of 32×32×256. The specific parameters for the CBS operations are: number of output channels 128, convolution kernel size 3×3, stride 1, padding value 1.

[0058] The seventeenth layer is a convolutional layer. Its input is the feature map output by the sixteenth layer. It performs a CBS operation on this feature map to output a feature map with a dimension of 32×32×1024. The specific parameters of the CBS operation are: the number of output channels is 256, the convolutional kernel size is 3×3, the stride is 1, and the padding value is 1.

[0059] The eighteenth layer is a convolutional layer. Its input is the feature map output by the seventeenth layer. It performs a CBS operation on this feature map to output a feature map with a dimension of 32×32×512. The specific parameters of the CBS operation are: the number of output channels is 128, the convolutional kernel size is 3×3, the stride is 1, and the padding value is 1.

[0060] The nineteenth layer is a detection head layer. Its input is the feature map output by the eighteenth layer. It performs small object detection branch processing on this feature map to output detection boxes and classification results. The specific parameters of the detection head layer are: the convolutional kernel size is 3×3, the stride is 1, and the padding value is 1.

[0061] The twentieth layer is a detection head layer. Its input is the feature map output by the nineteenth layer. It performs small object detection branch processing on this feature map to output detection boxes and classification results. The specific parameters of the detection head layer are: the convolutional kernel size is 3×3, the stride is 1, and the padding value is 1.

[0062] The twenty - first layer is a detection head layer. Its input is the feature map output by the twentieth layer. It performs small object detection branch processing on this feature map to output detection boxes and classification results. The specific parameters of the detection head layer are: the convolutional kernel size is 3×3, the stride is 1, and the padding value is 1.

[0063] The twenty - second layer is a detection head layer. Its input is the feature map output by the twenty - first layer. It performs medium object detection branch processing on this feature map to output detection boxes and classification results. The specific parameters of the detection head layer are: the convolutional kernel size is 3×3, the stride is 1, and the padding value is 1.

[0064] The twenty - third layer is a detection head layer. Its input is the feature map output by the twenty - second layer. It performs medium object detection branch processing on this feature map to output detection boxes and classification results. The specific parameters of the detection head layer are: the convolutional kernel size is 3×3, the stride is 1, and the padding value is 1.

[0065] The twenty - fourth layer is a detection head layer. Its input is the feature maps output by the twenty - second layer and the twenty - third layer. This layer performs detection branch processing on these feature maps to output detection boxes and classification results, and then generates a processed mask image based on these results. The specific parameters of the detection branch processing are: the convolutional kernel size is 3×3, the stride is 1, and the padding value is 1.

[0066] Preferably, the mask object detection model is trained through the following steps:

[0067] (a1) Download the open-source mask dataset and divide the mask dataset into a training set and a test set according to a ratio of 8:2.

[0068] (a2) Perform data preprocessing on the training set obtained in step (a1) to obtain the preprocessed training set.

[0069] (a3) Perform image enhancement processing on the preprocessed training set obtained in the online process step (a2) to obtain the enhanced training set. For each sample in the enhanced training set (whose dimension is 512×512×3), input it into the first layer of the backbone network of the mask object detection model for processing to output a feature map corresponding to the sample with a dimension of 256×256×64.

[0070] (a4) For each sample in the enhanced training set obtained in step (a3), input the feature map corresponding to the sample with a dimension of 256×256×64 obtained in step (a3) into the second layer of the backbone network of the mask object detection model for processing to output a feature map corresponding to the sample with a dimension of 128×128×128.

[0071] (a5) For each sample in the enhanced training set obtained in step (a3), input the feature map corresponding to the sample with a dimension of 128×128×128 obtained in step (a4) into the third layer of the backbone network of the mask object detection model for multi-scale feature extraction to output a feature map corresponding to the sample with a dimension of 128×128×128.

[0072] (a6) For each sample in the enhanced training set obtained in step (a3), input the feature map corresponding to the sample with a dimension of 128×128×128 obtained in step (a5) into the fourth layer of the backbone network of the mask object detection model for downsampling to output a feature map corresponding to the sample with a dimension of 64×64×256.

[0073] (a7) For each sample in the enhanced training set obtained in step (a3), input the feature map corresponding to the sample with a dimension of 64×64×256 obtained in step (a6) into the fifth layer of the backbone network of the mask object detection model for multi-scale feature extraction to output a feature map corresponding to the sample with a dimension of 64×64×256.

[0074] (a8) For each sample in the enhanced training set obtained in step (a3), the feature map corresponding to this sample with a dimension of 64×64×256 obtained in step (a7) is input into the 6th layer of the backbone network of the mask target detection model for downsampling to output a feature map corresponding to this sample with a dimension of 32×32×512.

[0075] (a9) For each sample in the enhanced training set obtained in step (a3), the feature map corresponding to this sample with a dimension of 32×32×512 obtained in step (a8) is input into the 7th layer of the backbone network of the mask target detection model for multi-scale feature extraction to output a feature map corresponding to this sample with a dimension of 32×32×512.

[0076] (a10) For each sample in the enhanced training set obtained in step (a3), the feature map corresponding to this sample with a dimension of 128×128×128 obtained in step (a5), the feature map corresponding to this sample with a dimension of 64×64×256 obtained in step (a7), and the feature map corresponding to this sample with a dimension of 32×32×512 obtained in step (a9) are input into the 8th layer of the backbone network of the mask target detection model for multi-scale feature extraction to output a feature map corresponding to this sample with a dimension of 64×64×384.

[0077] (a11) For each sample in the enhanced training set obtained in step (a3), the feature map corresponding to this sample with a dimension of 64×64×384 obtained in step (a10) is input into the 9th layer of the backbone network of the mask target detection model for processing to output a feature map corresponding to this sample with a dimension of 32×32×256.

[0078] (a12) For each sample in the enhanced training set obtained in step (a3), the feature map corresponding to this sample with a dimension of 32×32×512 obtained in step (a9) and each feature map with a dimension of 32×32×256 obtained in step (a11) are input into the 10th layer of the backbone network of the mask target detection model for processing to output a feature map corresponding to this sample with a dimension of 32×32×768.

[0079] (a13) For each sample in the enhanced training set obtained in step (a3), the feature map corresponding to this sample with a dimension of 32×32×768 obtained in step (a12) is input into the 11th layer of the backbone network of the mask target detection model for processing to output a feature map corresponding to this sample with a dimension of 32×32×512.

[0080] (a14) For each sample in the enhanced training set obtained in step (a3), the feature map corresponding to this sample with a dimension of 64×64×384 obtained in step (a10) is input into the 12th layer of the backbone network of the mask target detection model for upsampling to output a feature map corresponding to this sample with a dimension of 128×128×384.

[0081] (a15) For each sample in the enhanced training set obtained in step (a3), the feature map corresponding to this sample with a dimension of 128×128×128 obtained in step (a5) and each feature map with a dimension of 128×128×384 obtained in step (a14) are input into the 13th layer of the backbone network of the mask target detection model for processing to output a feature map corresponding to this sample with a dimension of 128×128×512.

[0082] (a16) For each sample in the enhanced training set obtained in step (a3), the feature map corresponding to this sample with a dimension of 128×128×512 obtained in step (a15) is input into the 14th layer of the backbone network of the mask target detection model for processing to output a feature map corresponding to this sample with a dimension of 128×128×256.

[0083] (a17) For each sample in the enhanced training set obtained in step (a3), the feature map corresponding to this sample with a dimension of 64×64×384 obtained in step (a10), each feature map with a dimension of 32×32×512 obtained in step (a13), and each feature map with a dimension of 128×128×256 obtained in step (a16) are input into the 15th layer of the backbone network of the mask target detection model for multi-scale feature extraction to output a feature map corresponding to this sample with a dimension of 64×64×576.

[0084] (a18) For each sample in the enhanced training set obtained in step (a3), the feature map corresponding to this sample with a size of 64×64×576 obtained in step (a17) is input into the 16th layer of the backbone network of the mask target detection model for processing to output a feature map corresponding to this sample with a dimension of 32×32×256.

[0085] (a19) For each sample in the augmented training set obtained in step (a3), input the feature map corresponding to this sample with a size of 32×32×256 obtained in step (a11), the feature map corresponding to this sample with a size of 32×32×512 obtained in step (a13), and the feature map corresponding to this sample with a size of 32×32×256 obtained in step (a18) into the 17th layer of the backbone network of the mask object detection model for processing, so as to output a feature map corresponding to this sample with a dimension of 32×32×1024.

[0086] (a20) For each sample in the augmented training set obtained in step (a3), input the feature map corresponding to this sample with a size of 32×32×1024 obtained in step (a19) into the 18th layer of the backbone network of the mask object detection model for processing, so as to output a feature map corresponding to this sample with a dimension of 32×32×512.

[0087] (a21) For each sample in the augmented training set obtained in step (a3), input the feature map corresponding to this sample with a size of 64×64×576 obtained in step (a17) into the 19th layer of the backbone network of the mask object detection model for upsampling, so as to output a feature map corresponding to this sample with a dimension of 128×128×576.

[0088] (a22) For each sample in the augmented training set obtained in step (a3), input the feature map corresponding to this sample with a size of 128×128×384 obtained in step (a14), the feature map corresponding to this sample with a size of 128×128×256 obtained in step (a16), and the feature map corresponding to this sample with a size of 128×128×576 obtained in step (a21) into the 20th layer of the backbone network of the mask object detection model for processing, so as to output a feature map corresponding to this sample with a dimension of 128×128×1216.

[0089] (a23) For each sample in the augmented training set obtained in step (a3), input the feature map corresponding to this sample with a size of 128×128×1216 obtained in step (a22) into the 21st layer of the backbone network of the mask object detection model for processing, so as to output a feature map corresponding to this sample with a dimension of 128×128×256.

[0090] (a24) For each sample in the augmented training set obtained in step (a3), the feature map corresponding to this sample with a size of 128×128×128 obtained in step (a5) is input into the 22nd layer of the backbone network of the mask target detection model for processing, so as to output a feature map corresponding to this sample with a dimension of 128×128×256.

[0091] (a25) For each sample in the augmented training set obtained in step (a3), the feature map corresponding to this sample with a size of 64×64×256 obtained in step (a7) is input into the 23rd layer of the backbone network of the mask target detection model for processing, so as to output a feature map corresponding to this sample with a dimension of 64×64×576.

[0092] (a26) For each sample in the augmented training set obtained in step (a3), the feature map corresponding to this sample with a size of 128×128×256 obtained in step (a23) and the feature map corresponding to this sample with a size of 128×128×256 obtained in step (a24) are input into the 24th layer of the backbone network of the mask target detection model for processing, so as to output a feature map corresponding to this sample with a dimension of 128×128×256.

[0093] (a27) For each sample in the augmented training set obtained in step (a3), the feature map corresponding to this sample with a size of 64×64×576 obtained in step (a17) and the feature map corresponding to this sample with a size of 64×64×576 obtained in step (a25) are input into the 24th layer of the backbone network of the mask target detection model for processing, so as to output a feature map corresponding to this sample with a dimension of 64×64×576.

[0094] (a28) For each sample in the augmented training set obtained in step (a3), the feature map corresponding to this sample with a size of 32×32×512 obtained in step (a9) and the feature map corresponding to this sample with a size of 32×32×512 obtained in step (a20) are input into the 23rd layer of the backbone network of the mask target detection model for processing, so as to output a feature map corresponding to this sample with a dimension of 32×32×512.

[0095] (a29) For each sample in the enhanced training set obtained in step (a3), input the feature map corresponding to this sample with a size of 128×128×256 obtained in step (a26), the feature map corresponding to this sample with a size of 64×64×576 obtained in step (a27), and the feature map corresponding to this sample with a size of 32×32×512 obtained in step (a28) into the 24th layer of the backbone network of the mask target detection model for processing, so as to obtain the bounding box regression loss and classification loss corresponding to this sample.

[0096] (a30) For each sample in the enhanced training set obtained in step (a3), according to the bounding box regression loss L EIou and classification loss L BCE obtained for this sample in step (a29), calculate the total loss L total , where λ1 and λ2 represent loss weights, and the value ranges of all three are positive real numbers. Preferably, they are equal to 0.5 and 1.5 respectively.

[0097] (a31) For each sample in the enhanced training set obtained in step (a3), according to the total loss corresponding to this sample obtained in step (a30), and use gradient descent to iteratively train the mask target detection model until the mask target detection model reaches the preset number of iterations, and obtain the optimal parameters of the mask target detection model at this time, so as to obtain a preliminarily trained mask target detection model.

[0098] (a32) Use the test set obtained in step (a1) to test the mask target detection model preliminarily trained in step (a31) until the obtained detection accuracy reaches the optimum, so as to obtain a finally trained mask target detection model.

[0099] Preferably, the image enhancement processing in step (a3) randomly selects one or any combination of the following 9 data enhancement methods: image HSV enhancement, including hue, saturation, and brightness enhancement, and their enhancement factors are set to 0.015, 0.7, and 0.4 respectively; image translation, and its translation factor is 0.1; image scaling, and its scaling factor is 0.5; image horizontal flipping, and its flipping factor is 0.5; multi-image stitching, and its stitching factor is 1.0; image erasing, and its erasing factor is 0.4; image cropping, and its cropping factor is 1.0.

[0100] Preferably, the bounding box regression loss L EIou adopts the EIoU loss, and its calculation formula is:

[0101]

[0102] Among them, b represents the center point coordinates of the predicted bounding box, b gt represents the center point coordinates of the ground truth bounding box, w represents the width of the predicted bounding box, w gt represents the width of the ground truth bounding box, h represents the height of the predicted bounding box, h gt represents the height of the ground truth bounding box, c represents the diagonal length of the smallest enclosing rectangle that can contain both the predicted bounding box and the ground truth bounding box, C w represents the width of the smallest enclosing rectangle that can contain both the predicted bounding box and the ground truth bounding box, C h represents the height of the smallest enclosing rectangle that can contain both the predicted bounding box and the ground truth bounding box, p(w, w gt ) represents the deviation between the width of the predicted bounding box and the width of the ground truth bounding box, used to calculate the deviation in width between the predicted bounding box and the ground truth bounding box, p(h, h gt ) represents the deviation between the height of the predicted bounding box and the height of the ground truth bounding box, used to calculate the deviation in height between the predicted bounding box and the ground truth bounding box, p(b, b gt ) represents the deviation between the center point of the predicted bounding box and the center point of the ground truth bounding box, used to calculate the deviation in the center position between the predicted bounding box and the ground truth bounding box, L Iou represents the Intersection over Union (IoU) loss, used to measure the overlap degree between the predicted bounding box and the ground truth bounding box, and its calculation formula is:

[0103]

[0104] where Area of Overlap represents the intersection area of the predicted bounding box and the ground truth bounding box, and Area of Union represents the union area of the predicted bounding box and the ground truth bounding box;

[0105] The intersection area Area of Overlap of the predicted bounding box and the ground truth bounding box = max(0, right - left) * max(0, bottom - top), where right and left are the right boundary coordinates and left boundary coordinates of the intersection area of the predicted bounding box and the ground truth bounding box respectively, and bottom and top are the lower boundary coordinates and upper boundary coordinates of the intersection area of the predicted bounding box and the ground truth bounding box respectively;

[0106] The union area Area of Union of the predicted bounding box and the ground truth bounding box = the area of the predicted bounding box + the area of the ground truth bounding box - the intersection area Area of Overlap.

[0107] Preferably, the classification loss L BCE adopts binary cross - entropy loss, and its calculation formula is:

[0108] L BCE = -α * y * log(p) - β * (1 - y) * log(1 - p)

[0109] Among them, α and β are loss weights, and α + β = 1. The value range of α is from 0 to 1, preferably 0.5. The value range of β is from 0 to 1, preferably 0.5. y is the true classification value of the sample, and p is the predicted classification value of the sample.

[0110] The calculation formula for the total loss is:

[0111] L total = λ1 * L EIou + λ2 * L BCE

[0112] Among them, λ1 and λ2 are loss weights used to balance the contributions of different loss terms to the total loss. The value range of λ1 is from 0 to 1, and the value range of λ2 is from 1 to 2.

[0113] According to another aspect of the present invention, a mask target detection system based on an improved YOLOWorld world model is provided, including:

[0114] The first module is used to obtain the mask image to be detected;

[0115] The second module is used to perform data preprocessing on the mask image to be detected obtained by the first module to obtain the mask image after data preprocessing;

[0116] The third module is used to perform image processing on the mask image after data preprocessing by the second module to obtain the mask image after image processing;

[0117] The fourth module is used to input the mask image after image processing obtained by the third module into a pre-trained mask target detection model to obtain the final detection result.

[0118] Generally speaking, compared with the prior art by the above technical solutions conceived by the present invention, the following beneficial effects can be achieved:

[0119] (1) The present invention can solve the prominent technical problem of misdetection of the existing color space conversion method in complex backgrounds: Since the present invention adopts sub-steps (3-1) to (3-4), by introducing an attention mechanism and a data augmentation strategy, it significantly enhances the model's adaptability to complex backgrounds; the attention mechanism enables the model to focus on the key features of the mask (such as contours and edges), thereby effectively reducing the interference of light changes and similar objects; at the same time, the data augmentation strategy further improves the robustness and accuracy of the model by simulating various complex scenarios;

[0120] (2) The present invention can solve the technical problem that it is difficult to achieve a balance between real-time performance and accuracy in existing morphological processing methods: Due to the model structure design in step (4), the present invention adopts a feature pyramid network structure and an improved SPPFCSPC module, which reduces the computational overhead while ensuring high detection accuracy; the FPN structure can fuse feature maps of different scales, enabling effective detection of targets of different sizes and positions, especially suitable for objects with large size differences such as masks; the SPPFCSPC module refines and strengthens feature expressions through convolution operations, further improving the detection accuracy while maintaining the efficiency of the model;

[0121] (3) The present invention can solve the technical problem that the model recognition ability of existing template matching methods is limited: Due to the model structure design in sub-steps (3-1) to (3-4) and step (4), the present invention significantly improves the recognition ability of masks at extreme angles and poses through an innovative model architecture and training method; the model can learn more comprehensive and detailed feature representations, thus effectively reducing the situations of missed detection and false detection; in addition, operations such as random rotation and flipping in the data augmentation strategy also make the model more robust when dealing with masks in different poses;

[0122] (4) The model of the present invention has excellent generalization ability: Due to the dataset adjustment and fusion method and data augmentation technology in steps (a1) to (a32), the present invention ensures the stability and reliability of the model. It not only enriches the diversity of training data but also enables the model to better adapt to various actual application scenarios, improving the generalization ability of the model.

[0123] (5) The model of the present invention has good computational efficiency: The present invention focuses on optimizing computational efficiency in model design. By adopting the model structure design in step (4), using a lightweight network structure and an efficient algorithm, it reduces the number of model parameters and computational complexity, making it more suitable for running on resource-constrained embedded devices while maintaining high detection accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0124] Figure 1 is a flowchart of the mask target detection method based on the improved YOLOWorld world model of the present invention;

[0125] Figure 2 is a schematic structural diagram of the mask target detection model of the present invention;

[0126] Figure 3 is a specific schematic diagram of step (3) in the method of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0127] In order to make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0128] The core idea of the present invention is to significantly improve the object detection performance of masks and safety helmets through three key steps. First, the present invention performs detailed data preprocessing on the object detection images, including size adjustment, normalization, and data augmentation, to enrich the context information of the images and enhance the robustness of the detection network. Secondly, the present invention makes an innovative improvement to the original SPPF module and replaces it with the SPPFCSPC module to improve the accuracy of the object detection network. Even in a complex and changing environment, this module can ensure the high quality and accuracy of the detection results.

[0129] As Figure 1 shown, the present invention provides a mask object detection method based on an improved YOLOWorld world model, including the following steps:

[0130] (1) Obtain a mask image to be detected;

[0131] Specifically, in this step, a mask image to be detected is obtained in real time from a camera or a real-time monitoring platform.

[0132] The advantage of this step (1) is that the images obtained in real time can ensure the diversity and representativeness of the data, providing a rich sample basis for the training of subsequent detection models.

[0133] (2) Perform data preprocessing on the mask image to be detected obtained in step (1) to obtain a mask image after data preprocessing;

[0134] Specifically, the data preprocessing operations include performing size transformation processing, normalization processing, and data augmentation processing on the mask image in sequence. Among them, the size transformation processing is to scale the mask image to a fixed size of 512 pixels × 512 pixels; the normalization processing is to map each pixel value in the mask image after the size transformation processing to the range of 0 to 1; the data augmentation processing is to perform one or any combination of random rotation, mirroring, and HSV perturbation processing on the mask image after the normalization processing.

[0135] The advantage of this step (2) is that through size transformation, normalization, and data augmentation processing, not only the image size is unified, the context information of the mask images is enriched, and the generalization ability of the model is improved, but also various actual application scenarios are simulated through the augmentation operation, further enhancing the robustness and adaptability of the mask target detection network that improves the YOLOWorld world model.

[0136] (3) Perform image processing on the mask images after the data preprocessing in step (2) to obtain the mask images after image processing;

[0137] As Figure 3 shown, step (3) includes the following sub-steps:

[0138] (3-1) Divide the mask images after the data preprocessing in step (2) into the current region and other regions;

[0139] The advantage of this step (3-1) is that through pre-segmentation and the K-means clustering algorithm, the clustering centers of the superpixel points can be adaptively determined, thereby realizing reasonable segmentation of the image and providing an accurate regional division for subsequent calculation of the atmospheric light value and transmittance.

[0140] This step (3-1) includes the following sub-steps:

[0141] (3-1-1) Perform pre-segmentation on the mask images after the data preprocessing in step (2) to obtain multiple image sub-regions, and each image sub-region serves as a superpixel point;

[0142] (3-1-2) Use the K-means clustering algorithm to cluster all the superpixel points obtained in step (3-1-1) to obtain the current region and other regions, obtain the gray values of all the superpixel points obtained in step (3-1-1), and take the average value of all the gray values as the adaptive threshold. All the superpixel points with gray values greater than or equal to this adaptive threshold among all the superpixel points obtained in step (3-1-1) are classified into the current region, and all the superpixel points with gray values less than this adaptive threshold among all the superpixel points obtained in step (3-1-1) are classified into other regions.

[0143] (3-2) Obtain the atmospheric light value of the mask images after the data preprocessing in step (2) according to the current region obtained in step (3-1);

[0144] This step is specifically as follows: First, sort the grayscale values of all superpixels in the current region in descending order. Then, obtain the top k% (where the value range of k is from 1 to 5, preferably 1) of the superpixels from the sorting result, and obtain the corresponding superpixels of these superpixels in the preprocessed mask image in step (2). Subsequently, take the average value of the grayscale values of all corresponding superpixels as the atmospheric light value of the preprocessed mask image in step (2), that is, the atmospheric light value is a three-element vector, and each element corresponds to one of the R, G, and B color channels.

[0145] The advantage of this step (3-2) is that by sorting and selecting the top k% of the superpixels, the bright area in the image can be effectively captured, thereby accurately estimating the atmospheric light value and providing a key parameter for subsequent defogging processing.

[0146] (3-3) Obtain the transmittance of the preprocessed mask image in step (2) according to the transmittances of the current region and other regions obtained in step (3-1);

[0147] This step is specifically as follows. First, for the current region, calculate the transmittance t of this current region through the following formula f :

[0148]

[0149] where I f is the pixel value of the current region, A is the atmospheric light value, β is the medium scattering coefficient, and d f is the depth information of the current region.

[0150] Then, for other regions, calculate the transmittance t of this other region through the following formula b :

[0151]

[0152] where I b is the pixel value of the other region, and d b is the depth information of the other region.

[0153] Subsequently, perform weighted averaging on the transmittance t f of the current region and the transmittance t b of the other regions to obtain the transmittance t of the preprocessed mask image in step (2):

[0154]

[0155] where S f is the area of the current region, and S b is the area of the other region.

[0156] The advantage of this step (3-3) is that the transmittances of the current region and other regions are calculated separately and weighted averaged, fully considering the characteristics of different regions in the image, making the final transmittance more accurate and reasonable.

[0157] (3-4) Perform defogging processing on the preprocessed mask image of the mask according to the transmittance of the preprocessed mask image obtained in step (3-3) to obtain the processed mask image I of the image processing * ;

[0158] This step is calculated using the following formula:

[0159]

[0160] where I is the preprocessed mask image obtained in step (3-3), and t0 is the minimum value of the preset transmittance, and its value range is from 0 to 0.1, preferably 0.05.

[0161] The advantage of this step (3-4) is that it can remove the fog influence in the image, enhance the clarity and contrast of the image, and further improve the accuracy of mask target detection.

[0162] The advantages of the above sub-steps (3-1) to (3-4) are that the above image is segmented into regions and other regions, and the transmittance is calculated separately and then the transmittance of the entire marine image is determined, so as to solve the problem of regional distortion after the dark channel prior image algorithm.

[0163] (4) Input the processed mask image obtained in step (3) into a pre-trained mask target detection model to obtain the final detection result;

[0164] The advantage of this step (4) is that by using a pre-trained mask target detection model, it can quickly and accurately detect the input image, output a detection box containing the mask target position and category, and achieve efficient target detection.

[0165] Specifically, the final detection result obtained in this step exists in the form of a detection box, and each detection box marks the predicted mask target position and mask target category.

[0166] Specifically, such as Figure 2As shown, the mask target detection model in the present invention adopts an improved YOLOWorld model, which includes a backbone network, a Spatial Pyramid Pooling Faster Cross StagePartial Channel (SPPFCSPC) module, a cross-stage partial connection branch module, and a feature fusion and output module.

[0167] The specific structure of the backbone network is as follows:

[0168] The first layer is a convolutional layer. Its input is an image with a dimension of 512*512*3. It performs a Convolution+Batch Normalization+SiLU (CBS) operation on this image. The parameters are as follows: the number of output channels is 64, the convolutional kernel size is 3, the stride is 2, the padding value is 1, and the output is a feature map with a dimension of 256×256×32.

[0169] The second layer is a convolutional layer. Its input is the feature map with a dimension of 256×256×32 output by the first layer. It performs a CBS operation on this feature map. The parameters are as follows: the number of output channels is 64, the convolutional kernel size is 3, the stride is 2, the padding value is 1, and the output is a feature map with a dimension of 128×128×64.

[0170] The input of the SPPFCSPC module is the feature map with a dimension of 128×128×64 output by the second layer of the backbone network. It first performs three CBS operations (cv1, cv3, cv4) on the input feature map to obtain a feature map x1 with a dimension of 128×128×64. The specific parameters of the CBS operation are as follows: the convolutional kernel size is 3, the stride is 1, the padding value is 1, and the number of output channels is 64. Then, it performs two-level maximum pooling processing (pooling kernel 5×5, stride 1, padding 2) on the feature map x1 to generate a feature map x2 with a dimension of 128×128×64 (denoted as x2-1, used to capture a wider context), a feature map x3 with a dimension of 128×128×64 (denoted as x3-1, used to enhance local feature expression), and a feature map m(x3) with a dimension of 128×128×64 (denoted as m(x3)-1, used to further highlight important features). Subsequently, the obtained feature maps x1, x2, x3, and m(x3) are concatenated along the channel dimension to obtain a feature map with a dimension of 128×128×256. Finally, it performs a CBS operation on the obtained feature map with a dimension of 128×128×256 to output a feature map y1 with a dimension of 128×128×64. The specific parameters of the CBS operation are as follows: the convolutional kernel size is 3, the stride is 1, the padding value is 1, and the number of output channels is 64.

[0171] The input of the cross-stage partial connection branch module is the feature map output by the SPPFCSPC module. It performs a CBS operation on the input feature map to output a feature map y2 with dimensions of 128×128×64. The specific parameters of the CBS operation are: convolution kernel size 3, stride 1, padding value 1, and number of output channels 64.

[0172] The specific structure of the feature fusion and output module is as follows:

[0173] The fourth layer is a pooling layer. Its input is the feature map with dimensions of 128×128×64 output by the second layer in the backbone network. It performs max-pooling on this feature map (where the pooling kernel is 2×2 and the stride is 2) to output a feature map with dimensions of 64×64×64.

[0174] The fifth layer is a convolutional layer. Its input is the feature map output by the fourth layer. It performs a CBS operation on this feature map to output a feature map with dimensions of 64×64×128. The specific parameters of the CBS operation are: number of output channels 128, convolution kernel size 3×3, stride 1, and padding value 1.

[0175] The sixth layer is an upsampling layer. Its input is the feature map output by the fifth layer. It performs upsampling on this feature map (where the scaling factor is 2) to fuse feature information at different levels and outputs a feature map with dimensions of 128×128×128.

[0176] The seventh layer is a downsampling layer. Its input is the feature map output by the sixth layer. It performs downsampling on this feature map (where the pooling kernel is 2×2 and the stride is 2) to output a feature map with dimensions of 64×64×256.

[0177] The eighth layer is a multi-scale feature extraction layer. Its input is the feature map output by the seventh layer. It performs multi-scale feature extraction on this feature map to output a feature map with dimensions of 64×64×256.

[0178] The ninth layer is a downsampling layer. Its input is the feature map output by the eighth layer. It performs downsampling on this feature map (where the pooling kernel is 2×2 and the stride is 2) to output a feature map with dimensions of 32×32×512.

[0179] The tenth layer is a multi-scale feature extraction layer. Its input is the feature map output by the ninth layer. It performs multi-scale feature extraction on this feature map to output a feature map with dimensions of 32×32×512.

[0180] The eleventh layer is a semantic fusion layer. Its input is the feature map output by the tenth layer. It performs multi-scale feature extraction on this feature map to output a feature map with dimensions of 64×64×384.

[0181] The twelfth layer is an upsampling layer. Its input is the feature map output by the eleventh layer. It performs upsampling on this feature map (where the scaling factor is 2) to output a feature map with a dimension of 128×128×384.

[0182] The thirteenth layer is a feature fusion layer. Its input is the feature map output by the twelfth layer. It performs feature fusion on this feature map to output a feature map with a dimension of 64×64×576. The specific parameters for feature fusion are: convolution kernel size 3×3, stride 1, padding value 1.

[0183] The fourteenth layer is a convolutional layer. Its input is the feature map output by the thirteenth layer. It performs a CBS operation on this feature map to output a feature map with a dimension of 32×32×512. The specific parameters for the CBS operation are: number of output channels 128, convolution kernel size 3×3, stride 1, padding value 1.

[0184] The fifteenth layer is a semantic fusion layer. Its input is the feature map output by the fourteenth layer. It performs multi-scale feature fusion on this feature map. The specific parameters for multi-scale feature fusion are: convolution kernel size 3×3, stride 1, padding value 1.

[0185] The sixteenth layer is a convolutional layer. Its input is the feature map output by the fifteenth layer. It performs a CBS operation on this feature map to output a feature map with a dimension of 32×32×256. The specific parameters for the CBS operation are: number of output channels 128, convolution kernel size 3×3, stride 1, padding value 1.

[0186] The seventeenth layer is a convolutional layer. Its input is the feature map output by the sixteenth layer. It performs a CBS operation on this feature map to output a feature map with a dimension of 32×32×1024. The specific parameters for the CBS operation are: number of output channels 256, convolution kernel size 3×3, stride 1, padding value 1.

[0187] The eighteenth layer is a convolutional layer. Its input is the feature map output by the seventeenth layer. It performs a CBS operation on this feature map to output a feature map with a dimension of 32×32×512. The specific parameters for the CBS operation are: number of output channels 128, convolution kernel size 3×3, stride 1, padding value 1.

[0188] The nineteenth layer is a detection head layer. Its input is the feature map output by the eighteenth layer. It performs small object detection branch processing on this feature map to output detection boxes and classification results. The specific parameters for the detection head layer are: convolution kernel size 3×3, stride 1, padding value 1.

[0189] The twentieth layer is the detection head layer. Its input is the feature map output by the nineteenth layer. It performs small object detection branch processing on this feature map to output detection boxes and classification results. The specific parameters of the detection head layer are: convolution kernel size 3×3, stride 1, padding value 1.

[0190] The twenty - first layer is the detection head layer. Its input is the feature map output by the twentieth layer. It performs small object detection branch processing on this feature map to output detection boxes and classification results. The specific parameters of the detection head layer are: convolution kernel size 3×3, stride 1, padding value 1.

[0191] The twenty - second layer is the detection head layer. Its input is the feature map output by the twenty - first layer. It performs medium object detection branch processing on this feature map to output detection boxes and classification results. The specific parameters of the detection head layer are: convolution kernel size 3×3, stride 1, padding value 1.

[0192] The twenty - third layer is the detection head layer. Its input is the feature map output by the twenty - second layer. It performs medium object detection branch processing on this feature map to output detection boxes and classification results. The specific parameters of the detection head layer are: convolution kernel size 3×3, stride 1, padding value 1.

[0193] The twenty - fourth layer is the detection head layer. Its inputs are the feature maps output by the twenty - second and twenty - third layers. This layer performs detection branch processing on these feature maps to output detection boxes and classification results, and then generates processed mask images based on these results. The specific parameters of the detection branch are: convolution kernel size 3×3, stride 1, padding value 1.

[0194] The mask object detection model of the present invention is obtained through the following steps:

[0195] (a1) Download an open - source mask dataset (which includes 100,000 labeled images), and divide this mask dataset into a training set and a test set according to a ratio of 8:2.

[0196] Specifically, in this step, mask images to be detected are obtained in real - time from a camera or a real - time monitoring platform.

[0197] The advantage of this step (1) is that the images obtained in real - time can ensure the diversity and representativeness of the data, providing a rich sample basis for the training of the subsequent detection model.

[0198] (a2) Perform data pre - processing on the training set obtained in step (a1) to obtain a pre - processed training set, where each sample has a dimension of 512×512×3;

[0199] Specifically, the data preprocessing process in this step is basically the same as that in step (2) above. The scale transformation process therein is to adjust the size of each image in the training set to 512×512×3 to ensure that all samples in the training set have consistent dimensions.

[0200] By adjusting the size of the images, a unified input format is provided for the subsequent training of the detection model.

[0201] (a3) Perform image enhancement processing on the preprocessed training set obtained in step (a2) of the online process to obtain an enhanced training set. For each sample in the enhanced training set (whose dimension is 512×512×3), input it into the first layer of the backbone network of the mask target detection model for processing to output a feature map corresponding to this sample with a dimension of 256×256×64.

[0202] Specifically, in this step, one or any combination of the following 9 data enhancement methods can be randomly selected for processing: image HSV enhancement, including hue, saturation, and brightness enhancement, with their enhancement factors set to 0.015, 0.7, and 0.4 respectively; image translation, with a translation factor of 0.1; image scaling, with a scaling factor of 0.5; image horizontal flipping, with a flipping factor of 0.5; multi-image stitching, with a stitching factor of 1.0; image erasing, with an erasing factor of 0.4; image cropping, with a cropping factor of 1.0.

[0203] (a4) For each sample in the enhanced training set obtained in step (a3), input the feature map corresponding to this sample with a dimension of 256×256×64 obtained in step (a3) into the second layer of the backbone network of the mask target detection model for processing to output a feature map corresponding to this sample with a dimension of 128×128×128.

[0204] (a5) For each sample in the enhanced training set obtained in step (a3), input the feature map corresponding to this sample with a dimension of 128×128×128 obtained in step (a4) into the third layer of the backbone network of the mask target detection model for multi-scale feature extraction to output a feature map corresponding to this sample with a dimension of 128×128×128.

[0205] (a6) For each sample in the enhanced training set obtained in step (a3), input the feature map corresponding to this sample with a dimension of 128×128×128 obtained in step (a5) into the fourth layer of the backbone network of the mask target detection model for downsampling to output a feature map corresponding to this sample with a dimension of 64×64×256.

[0206] (a7) For each sample in the enhanced training set obtained in step (a3), the feature map corresponding to this sample with a dimension of 64×64×256 obtained in step (a6) is input into the 5th layer of the backbone network of the mask target detection model for multi-scale feature extraction, so as to output the feature map corresponding to this sample with a dimension of 64×64×256.

[0207] (a8) For each sample in the enhanced training set obtained in step (a3), the feature map corresponding to this sample with a dimension of 64×64×256 obtained in step (a7) is input into the 6th layer of the backbone network of the mask target detection model for downsampling, so as to output the feature map corresponding to this sample with a dimension of 32×32×512.

[0208] (a9) For each sample in the enhanced training set obtained in step (a3), the feature map corresponding to this sample with a dimension of 32×32×512 obtained in step (a8) is input into the 7th layer of the backbone network of the mask target detection model for multi-scale feature extraction, so as to output the feature map corresponding to this sample with a dimension of 32×32×512.

[0209] (a10) For each sample in the enhanced training set obtained in step (a3), the feature map corresponding to this sample with a dimension of 128×128×128 obtained in step (a5), the feature map corresponding to this sample with a dimension of 64×64×256 obtained in step (a7), and the feature map corresponding to this sample with a dimension of 32×32×512 obtained in step (a9) are input into the 8th layer of the backbone network of the mask target detection model for multi-scale feature extraction, so as to output the feature map corresponding to this sample with a dimension of 64×64×384.

[0210] (a11) For each sample in the enhanced training set obtained in step (a3), the feature map corresponding to this sample with a dimension of 64×64×384 obtained in step (a10) is input into the 9th layer of the backbone network of the mask target detection model for processing, so as to output the feature map corresponding to this sample with a dimension of 32×32×256.

[0211] (a12) For each sample in the enhanced training set obtained in step (a3), the feature map corresponding to this sample with a dimension of 32×32×512 obtained in step (a9) and each feature map with a dimension of 32×32×256 obtained in step (a11) are input into the 10th layer of the backbone network of the mask target detection model for processing, so as to output the feature map corresponding to this sample with a dimension of 32×32×768.

[0212] (a13) For each sample in the enhanced training set obtained in step (a3), the feature map corresponding to this sample with a dimension of 32×32×768 obtained in step (a12) is input into the 11th layer of the backbone network of the mask target detection model for processing, so as to output the feature map corresponding to this sample with a dimension of 32×32×512.

[0213] (a14) For each sample in the enhanced training set obtained in step (a3), the feature map corresponding to this sample with a dimension of 64×64×384 obtained in step (a10) is input into the 12th layer of the backbone network of the mask target detection model for upsampling, so as to output the feature map corresponding to this sample with a dimension of 128×128×384.

[0214] (a15) For each sample in the enhanced training set obtained in step (a3), the feature map corresponding to this sample with a dimension of 128×128×128 obtained in step (a5) and each feature map with a dimension of 128×128×384 obtained in step (a14) are input into the 13th layer of the backbone network of the mask target detection model for processing, so as to output the feature map corresponding to this sample with a dimension of 128×128×512.

[0215] (a16) For each sample in the enhanced training set obtained in step (a3), the feature map corresponding to this sample with a dimension of 128×128×512 obtained in step (a15) is input into the 14th layer of the backbone network of the mask target detection model for processing, so as to output the feature map corresponding to this sample with a dimension of 128×128×256.

[0216] (a17) For each sample in the enhanced training set obtained in step (a3), the feature map corresponding to this sample with a dimension of 64×64×384 obtained in step (a10), each feature map with a dimension of 32×32×512 obtained in step (a13), and each feature map with a dimension of 128×128×256 obtained in step (a16) are input into the 15th layer of the backbone network of the mask target detection model for multi-scale feature extraction, so as to output the feature map corresponding to this sample with a dimension of 64×64×576.

[0217] (a18) For each sample in the enhanced training set obtained in step (a3), the feature map corresponding to this sample with a size of 64×64×576 obtained in step (a17) is input into the 16th layer of the backbone network of the mask target detection model for processing, so as to output the feature map corresponding to this sample with a dimension of 32×32×256.

[0218] (a19) For each sample in the augmented training set obtained in step (a3), the feature map corresponding to this sample with a size of 32×32×256 obtained in step (a11), the feature map corresponding to this sample with a size of 32×32×512 obtained in step (a13), and the feature map corresponding to this sample with a size of 32×32×256 obtained in step (a18) are input into the 17th layer of the backbone network of the mask target detection model for processing, so as to output a feature map corresponding to this sample with a dimension of 32×32×1024.

[0219] (a20) For each sample in the augmented training set obtained in step (a3), the feature map corresponding to this sample with a size of 32×32×1024 obtained in step (a19) is input into the 18th layer of the backbone network of the mask target detection model for processing, so as to output a feature map corresponding to this sample with a dimension of 32×32×512.

[0220] (a21) For each sample in the augmented training set obtained in step (a3), the feature map corresponding to this sample with a size of 64×64×576 obtained in step (a17) is input into the 19th layer of the backbone network of the mask target detection model for upsampling, so as to output a feature map corresponding to this sample with a dimension of 128×128×576.

[0221] (a22) For each sample in the augmented training set obtained in step (a3), the feature map corresponding to this sample with a size of 128×128×384 obtained in step (a14), the feature map corresponding to this sample with a size of 128×128×256 obtained in step (a16), and the feature map corresponding to this sample with a size of 128×128×576 obtained in step (a21) are input into the 20th layer of the backbone network of the mask target detection model for processing, so as to output a feature map corresponding to this sample with a dimension of 128×128×1216.

[0222] (a23) For each sample in the augmented training set obtained in step (a3), the feature map corresponding to this sample with a size of 128×128×1216 obtained in step (a22) is input into the 21st layer of the backbone network of the mask target detection model for processing, so as to output a feature map corresponding to this sample with a dimension of 128×128×256.

[0223] (a24) For each sample in the augmented training set obtained in step (a3), the feature map of size 128×128×128 corresponding to this sample obtained in step (a5) is input into the 22nd layer of the backbone network of the mask target detection model for processing, so as to output a feature map corresponding to this sample with a dimension of 128×128×256.

[0224] (a25) For each sample in the augmented training set obtained in step (a3), the feature map of size 64×64×256 corresponding to this sample obtained in step (a7) is input into the 23rd layer of the backbone network of the mask target detection model for processing, so as to output a feature map corresponding to this sample with a dimension of 64×64×576.

[0225] (a26) For each sample in the augmented training set obtained in step (a3), the feature map of size 128×128×256 corresponding to this sample obtained in step (a23) and the feature map of size 128×128×256 corresponding to this sample obtained in step (a24) are input into the 24th layer of the backbone network of the mask target detection model for processing, so as to output a feature map corresponding to this sample with a dimension of 128×128×256.

[0226] (a27) For each sample in the augmented training set obtained in step (a3), the feature map of size 64×64×576 corresponding to this sample obtained in step (a17) and the feature map of size 64×64×576 corresponding to this sample obtained in step (a25) are input into the 24th layer of the backbone network of the mask target detection model for processing, so as to output a feature map corresponding to this sample with a dimension of 64×64×576.

[0227] (a28) For each sample in the augmented training set obtained in step (a3), the feature map of size 32×32×512 corresponding to this sample obtained in step (a9) and the feature map of size 32×32×512 corresponding to this sample obtained in step (a20) are input into the 23rd layer of the backbone network of the mask target detection model for processing, so as to output a feature map corresponding to this sample with a dimension of 32×32×512.

[0228] (a29) For each sample in the enhanced training set obtained in step (a3), input the feature map corresponding to this sample with a size of 128×128×256 obtained in step (a26), the feature map corresponding to this sample with a size of 64×64×576 obtained in step (a27), and the feature map corresponding to this sample with a size of 32×32×512 obtained in step (a28) into the 24th layer of the backbone network of the mask target detection model for processing, so as to obtain the bounding box regression loss and classification loss corresponding to this sample.

[0229] Specifically, the bounding box regression loss L EIou Adopts the EIoU loss, and its calculation formula is:

[0230]

[0231] Where b represents the center point coordinates of the predicted box, b gt Represents the center point coordinates of the ground truth box, w represents the width of the predicted box, w gt Represents the width of the ground truth box, h represents the height of the predicted box, h gt Represents the height of the ground truth box, c represents the diagonal length of the smallest circumscribed rectangle that can contain both the predicted box and the ground truth box, C w Represents the width of the smallest circumscribed rectangle that can contain both the predicted box and the ground truth box, C h Represents the height of the smallest circumscribed rectangle that can contain both the predicted box and the ground truth box, p(w, w gt ) represents the deviation between the width of the predicted box and the width of the ground truth box, and is used to calculate the deviation between the predicted box and the ground truth box in terms of width, p(h, h gt ) represents the deviation between the height of the predicted box and the height of the ground truth box, and is used to calculate the deviation between the predicted box and the ground truth box in terms of height, p(b, b gt ) represents the deviation between the center point of the predicted box and the center point of the ground truth box, and is used to calculate the deviation between the predicted box and the ground truth box in terms of the center position, L Iou Represents the Intersection over Union (IoU) loss, which is used to measure the overlap degree between the predicted box and the ground truth box, and its calculation formula is:

[0232]

[0233] Among them, Area of Overlap represents the intersection area of the predicted bounding box and the ground truth bounding box, and Area of Union represents the union area of the predicted bounding box and the ground truth bounding box. The intersection area of the predicted bounding box and the ground truth bounding box, Area of Overlap = max(0, right - left) * max(0, bottom - top), where right and left are the right and left boundary coordinates of the intersection region of the predicted bounding box and the ground truth bounding box respectively, and bottom and top are the bottom and top boundary coordinates of the intersection region of the predicted bounding box and the ground truth bounding box respectively; the union area of the predicted bounding box and the ground truth bounding box, Area of Union = the area of the predicted bounding box + the area of the ground truth bounding box - the intersection area Area of Overlap;

[0234] Classification loss L BCE Binary cross-entropy loss is adopted, and its calculation formula is:

[0235] L BCE = -α * y * log(p) - β * (1 - y) * log(1 - p)

[0236] Among them, α and β are loss weights, and α + β = 1. The value range of α is from 0 to 1, preferably 0.5. The value range of β is from 0 to 1, preferably 0.5. y is the true classification value of the sample, and p is the predicted classification value of the sample.

[0237] (a30) For each sample in the enhanced training set obtained in step (a3), according to the bounding box regression loss L EIou and classification loss L BCE corresponding to this sample obtained in step (a29), calculate the total loss L total , where λ1 and λ2 represent loss weights, and the value ranges of all three are positive real numbers. Preferably, they are equal to 0.5 and 1.5 respectively.

[0238] The calculation formula of the total loss is:

[0239] L total = λ1 * L EIou + λ2 * L BCE

[0240] Among them, λ1 and λ2 are loss weights, which are used to balance the contributions of different loss terms to the total loss. In the present invention, the value range of λ1 is from 0 to 1, and the value range of λ2 is from 1 to 2. Preferably, λ1 = 0.5 and λ2 = 1.5.

[0241] For each sample in the enhanced training set obtained in step (a3), based on the total loss corresponding to this sample obtained in step (a30), use gradient descent to iteratively train the mask object detection model until the mask object detection model reaches a preset number of iterations (which is 100 times in the present invention), and obtain the optimal parameters of the mask object detection model at this time, so as to obtain a preliminarily trained mask object detection model.

[0242] Use the test set obtained in step (a1) to test the mask object detection model preliminarily trained in step (a31) until the obtained detection accuracy reaches the optimum, so as to obtain a finally trained mask object detection model.

[0243] Performance comparison

[0244] To comprehensively evaluate the performance of the present invention, multi-dimensional evaluation metrics are adopted. In terms of object detection performance, average precision (AP for short) and mean average precision (mAP for short) are introduced. AP quantifies the detection performance of the model on a single category by calculating the area under the precision-recall curve; mAP evaluates the comprehensive detection performance of the model in a multi-category scenario by taking the average of the APs of all categories, reflecting its generalization ability and overall performance.

[0245] In terms of evaluating the practicality of the model, three key metrics are introduced: frames per second (FPS for short), giga floating-point operations per second (GFLOPs for short), and the number of model parameters (Parameters). FPS measures the real-time processing ability of the model, reflecting the number of images processed per unit time; GFLOPs measures the computational complexity of the model, reflecting the amount of floating-point operations required for a single forward propagation; Parameters measures the scale and storage requirements of the model, reflecting the total number of trainable parameters. These metrics jointly evaluate the real-time performance, computational resource requirements, and storage efficiency of the model.

[0246] Test case 1

[0247] To verify the performance advantages of the method of the present invention in the mask target detection task, the mask target detection model of the present invention was compared and tested with 5 other related models on the Face-Mask-Detector-YOLO-Faster-R-CNN dataset. These 6 models include the mask detection model of the present invention, the Spatial Pyramid Pooling (SPP for short) model, the improved Spatial Pyramid Pooling-F (SPPF for short) model, the improved Simulated Spatial Pyramid Pooling-F (SimSPPF for short) model, the Basic Deformable Convolutional Module (BasicRFB for short) model, and the Spatial Pyramid Pooling-Cross Stage Partial Connections (SPPCSPC for short) model. The test metrics include FPS, GFLOPs, the number of model parameters, AP, and mAP, and a comprehensive evaluation was carried out from aspects such as real-time performance, computational complexity, model scale, and detection accuracy. The experimental results are shown in Table 1:

[0248] Table 1 Performance Comparison Table

[0249]

[0250] As can be seen from Table 1 above, in terms of enhancing feature expression:

[0251] SPP: It can effectively capture features of different scales, but its ability to extract fine-grained features in complex scenarios is limited. Due to the lack of additional convolutional processing, it may not be able to fully utilize the multi-scale information after pooling.

[0252] SPPF: Similar to SPP, it can effectively capture features of different scales, but its ability to extract fine-grained features in complex scenarios is still limited. Due to the lack of additional convolutional processing, it may not be able to fully utilize the multi-scale information after pooling.

[0253] SimSPPF: Similar to SPPF, it can effectively capture features of different scales, but its ability to extract fine-grained features in complex scenarios is still limited. Due to the lack of additional convolutional processing, it may not be able to fully utilize the multi-scale information after pooling.

[0254] BasicRFB: It can capture multi-scale features, and through its unique structural design, it has richer feature expressions compared with SPP and its variants, which helps to improve the detection performance of the model.

[0255] SPPCSPC: It can not only capture multi-scale features, but also further refine and strengthen these features through convolution operations. The CSP structure helps to retain more information of the original input, thus improving the quality of feature expressions, which is particularly beneficial for object detection in complex backgrounds.

[0256] The mask object detection model of the present invention: It can not only capture multi-scale features, but also further refine and strengthen these features through convolution operations. The CSP structure helps to retain more information of the original input, thus improving the quality of feature expressions, which is particularly beneficial for object detection in complex backgrounds.

[0257] As can be seen from Table 1 above, in terms of improving the detection performance of small objects:

[0258] SPP: It performs mediocre in small object detection because the fixed-size filters may not be sufficient to effectively capture the details of small targets. For very small targets, other methods may be needed to assist in detection.

[0259] SPPF: It performs mediocre in small object detection because the fixed-size filters may not be sufficient to effectively capture the details of small targets. For very small targets, other methods may be needed to assist in detection.

[0260] SimSPPF: It performs mediocre in small object detection because the fixed-size filters may not be sufficient to effectively capture the details of small targets. For very small targets, other methods may be needed to assist in detection.

[0261] BasicRFB: Due to its structural design, it can better integrate multi-scale information, so it usually performs better in small object detection. Its better feature expression ability and robustness make it more suitable for handling small-sized targets.

[0262] SPPCSPC: Since it can better integrate multi-scale information, and the CSP structure can help retain more fine-grained features, SPPCSPC usually performs better in small object detection. Its better feature expression ability and robustness make it more suitable for handling small-sized targets.

[0263] SPPFCSPC (the mask object detection model of the present invention): Since it can better integrate multi-scale information, and the CSP structure can help retain more fine-grained features, SPPFCSPC usually performs better in small object detection. Its better feature expression ability and robustness make it more suitable for handling small-sized targets.

[0264] As can be seen from Table 1 above, in terms of reducing the risk of overfitting:

[0265] SPP: Due to its relatively simple structure, more regularization means may be needed to prevent overfitting.

[0266] SPPF: Due to its relatively simple structure, more regularization means may be needed to prevent overfitting.

[0267] SimSPPF: Due to its relatively simple structure, more regularization means may be needed to prevent overfitting.

[0268] BasicRFB: Through its unique structural design, it can utilize data more effectively, reducing the risk of overfitting, especially in the case of limited data volume.

[0269] SPPCSPC: Through the CSP structure and convolution operations, SPPCSPC can utilize data more effectively, reducing the risk of overfitting, especially in the case of limited data volume.

[0270] The mask target detection model of the present invention: Through the CSP structure and convolution operations, SPPFCSPC can utilize data more effectively, reducing the risk of overfitting, especially in the case of limited data volume.

[0271] Those skilled in the art can easily understand that the above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A mask target detection method based on an improved YOLOWorld world model, characterized in that, It includes the following steps: (1) Obtain the mask image to be detected; (2) Perform data preprocessing on the mask image to be detected obtained in step (1) to obtain the mask image after data preprocessing; (3) Perform image processing on the mask image after data preprocessing in step (2) to obtain the mask image after image processing; (4) Input the mask image after image processing obtained in step (3) into a pre-trained mask target detection model to obtain the final detection result.

2. The mask target detection method based on the improved YOLOWorld world model according to claim 1, wherein the data preprocessing operation in step (2) includes performing size transformation processing, normalization processing, and data augmentation processing on the mask image in sequence, where the size transformation processing is to scale the mask image to a fixed size of 512 pixels × 512 pixels; the normalization processing is to map each pixel value in the mask image after size transformation processing to the range of 0 to 1; the data augmentation processing is to perform one or any combination of random rotation, mirroring, and HSV perturbation processing on the mask image after normalization processing. Step (3) includes the following sub-steps: (3-1) Divide the mask image after data preprocessing in step (2) into the current region and other regions; (3-2) Obtain the atmospheric light value of the mask image after data preprocessing in step (2) according to the current region obtained in step (3-1); (3-3) Obtain the transmittance of the mask image after data preprocessing in step (2) according to the transmittance of the current region and other regions obtained in step (3-1). (3-4) Dehaze the preprocessed mask image according to the transmittance of the preprocessed mask image obtained in step (3-3) to obtain the processed mask image I * . The final detection result obtained in step (4) exists in the form of a detection box, and each detection box marks the predicted mask target position and mask target category.

3. The mask target detection method based on the improved YOLOWorld world model according to claim 1 or 2, wherein step (3-1) includes the following sub-steps: (3-1-1) Perform pre-segmentation on the mask image after data preprocessing in step (2) to obtain multiple image sub-regions, and each image sub-region is used as a superpixel point; (3-1-2) Use the K-means clustering algorithm to cluster all the superpixel points obtained in step (3-1-1) to obtain the current region and other regions, obtain the gray values of all the superpixel points obtained in step (3-1-1), and take the average value of all the gray values as the adaptive threshold. All the superpixel points with gray values greater than or equal to the adaptive threshold among all the superpixel points obtained in step (3-1-1) are classified into the current region, and all the superpixel points with gray values less than the adaptive threshold among all the superpixel points obtained in step (3-1-1) are classified into other regions. Step (3-2) is specifically as follows: First, sort the grayscale values of all superpixels in the current region in descending order. Then, obtain the top k% of the superpixels from the sorting result, and obtain the corresponding superpixels of these superpixels in the mask image after data preprocessing in step (2). Subsequently, take the average value of the grayscale values of all corresponding superpixels as the atmospheric light value of the mask image after data preprocessing in step (2), that is, the atmospheric light value is a three-element vector, and each element corresponds to one of the R, G, and B color channels, where the value range of k is from 1 to 5, and preferably 1. Step (3-3) specifically is as follows. First, for the current area, the transmittance t of the current area is calculated using the following formula f :[[]] Where I f is the pixel value of the current area, A is the atmospheric light value, β is the medium scattering coefficient, and d f is the depth information of the current area. Then, for other regions, the transmittance t of the other regions is calculated by the following formula b :[[]]END]] where I b is the pixel value of other regions, and d b is the depth information of other regions. Subsequently, the transmittance t of the current region f and the transmittance t of other regions b are weighted and averaged to obtain the transmittance t of the mask image after data preprocessing in step (2): Among them, S f is the area of the current region, and S b is the area of other regions. Step (3-4) is calculated using the following formula: where I is the mask image after data preprocessing obtained in step (3-3), and t0 is the minimum value of the preset transmittance, and its value range is from 0 to 0.1, and preferably 0.

05.

4. The mask target detection method based on the improved YOLOWorld world model according to any one of claims 1 to 3, characterized in that the mask target detection model uses an improved YOLOWorld model, which includes a backbone network, a spatial pyramid pooling SPPFCSPC module, a cross-stage partial connection branch module, and a feature fusion and output module; The specific structure of the backbone network is: The first layer is a convolutional layer, whose input is an image with a dimension of 512*512*3. It performs a convolution normalization activation (Convolution+Batch Normalization+SiLU, convolution+batch normalization+SiLU activation, abbreviated as CBS) operation on this image, where the parameters are: the number of output channels is 64, the convolution kernel size is 3, the stride is 2, the padding value is 1, and the output dimension is a feature map of 256×256×32. The second layer is a convolutional layer, whose input is the feature map with a dimension of 256×256×32 output by the first layer. It performs a CBS operation on this feature map, where the parameters are: the number of output channels is 64, the convolution kernel size is 3, the stride is 2, the padding value is 1, and the output dimension is a feature map of 128×128×128. The third layer is a convolutional layer, whose input is the feature map with a dimension of 128×128×64 output by the second layer. It performs a CBS operation on this feature map, where the parameters are: the number of output channels is 128, the convolution kernel size is 3, the stride is 1, the padding value is 1, and the output dimension is a feature map of 128×128×64. The input of the SPPFCSPC module is the feature map with a dimension of 128×128×64 output by the second layer of the backbone network. It first performs three CBS operations on the input feature map to obtain a feature map x1 with a dimension of 128×128×64. The specific parameters of the CBS operation are: convolution kernel size 3, stride 1, padding value 1, and number of output channels 64. Then, it performs two-level max pooling on the feature map x1 to generate a feature map x2 with a dimension of 128×128×64, a feature map x3 with a dimension of 128×128×64, and a feature map m(x3) with a dimension of 128×128×64. Subsequently, it concatenates the obtained feature maps x1, x2, x3, and m(x3) along the channel dimension to obtain a feature map with a dimension of 128×128×256. Finally, it performs a CBS operation on the obtained feature map with a dimension of 128×128×256 to output a feature map y1 with a dimension of 128×128×64. The specific parameters of the CBS operation are: convolution kernel size 3, stride 1, padding value 1, and number of output channels 64. The input of the cross-stage partial connection branch module is the feature map output by the SPPFCSPC module. It performs a CBS operation on the input feature map to output a feature map y2 with a dimension of 128×128×64. The specific parameters of the CBS operation are: convolution kernel size 3, stride 1, padding value 1, and number of output channels 64.

5. The method for mask target detection based on the improved YOLOWorld world model according to claim 4, wherein The specific structure of the feature fusion and output module is as follows: The fourth layer is a pooling layer. Its input is the feature map with a dimension of 128×128×64 output by the second layer in the backbone network. It performs max pooling on this feature map to output a feature map with a dimension of 64×64×64. The fifth layer is a convolutional layer. Its input is the feature map output by the fourth layer. It performs a CBS operation on this feature map to output a feature map with a dimension of 64×64×128. The specific parameters of the CBS operation are: number of output channels 128, convolution kernel size 3×3, stride 1, padding value 1. The sixth layer is an upsampling layer. Its input is the feature map output by the fifth layer. It performs upsampling on this feature map to fuse feature information at different levels and outputs a feature map with a dimension of 128×128×128. The seventh layer is a downsampling layer. Its input is the feature map output by the sixth layer. It performs downsampling on this feature map to output a feature map with a dimension of 64×64×256. The eighth layer is a multi-scale feature extraction layer. Its input is the feature map output by the seventh layer. It performs multi-scale feature extraction on this feature map to output a feature map with a dimension of 64×64×256. The ninth layer is a downsampling layer. Its input is the feature map output by the eighth layer. It performs downsampling on this feature map to output a feature map with a dimension of 32×32×512. The tenth layer is a multi-scale feature extraction layer. Its input is the feature map output by the ninth layer. It performs multi-scale feature extraction on this feature map to output a feature map with a dimension of 32×32×512. The eleventh layer is the semantic fusion layer. Its input is the feature map output by the tenth layer. It performs multi-scale feature extraction on this feature map to output a feature map with a dimension of 64×64×384. The twelfth layer is the upsampling layer. Its input is the feature map output by the eleventh layer. It performs upsampling on this feature map to output a feature map with a dimension of 128×128×384. The thirteenth layer is the feature fusion layer. Its input is the feature map output by the twelfth layer. It performs feature fusion on this feature map and outputs a feature map with a dimension of 64×64×576. The specific parameters for feature fusion are: convolution kernel size 3×3, stride 1, padding value 1. The fourteenth layer is the convolutional layer. Its input is the feature map output by the thirteenth layer. It performs a CBS operation on this feature map to output a feature map with a dimension of 32×32×512. The specific parameters for the CBS operation are: number of output channels 128, convolution kernel size 3×3, stride 1, padding value 1. The fifteenth layer is the semantic fusion layer. Its input is the feature map output by the fourteenth layer. It performs multi-scale feature fusion on this feature map. The specific parameters for multi-scale feature fusion are: convolution kernel size 3×3, stride 1, padding value 1. The sixteenth layer is the convolutional layer. Its input is the feature map output by the fifteenth layer. It performs a CBS operation on this feature map to output a feature map with a dimension of 32×32×256. The specific parameters for the CBS operation are: number of output channels 128, convolution kernel size 3×3, stride 1, padding value 1. The seventeenth layer is the convolutional layer. Its input is the feature map output by the sixteenth layer. It performs a CBS operation on this feature map to output a feature map with a dimension of 32×32×1024. The specific parameters for the CBS operation are: number of output channels 256, convolution kernel size 3×3, stride 1, padding value 1. The eighteenth layer is the convolutional layer. Its input is the feature map output by the seventeenth layer. It performs a CBS operation on this feature map to output a feature map with a dimension of 32×32×512. The specific parameters for the CBS operation are: number of output channels 128, convolution kernel size 3×3, stride 1, padding value 1. The nineteenth layer is the detection head layer. Its input is the feature map output by the eighteenth layer. It performs small target detection branch processing on this feature map to output detection boxes and classification results. The specific parameters for the detection head layer are: convolution kernel size 3×3, stride 1, padding value 1. The twentieth layer is the detection head layer. Its input is the feature map output by the nineteenth layer. It performs small target detection branch processing on this feature map to output detection boxes and classification results. The specific parameters for the detection head layer are: convolution kernel size 3×3, stride 1, padding value 1. The twenty-first layer is the detection head layer. Its input is the feature map output by the twentieth layer. It performs small target detection branch processing on this feature map to output detection boxes and classification results. The specific parameters for the detection head layer are: convolution kernel size 3×3, stride 1, padding value 1. The twenty-second layer is the detection head layer. Its input is the feature map output by the twenty-first layer. It performs object detection branch processing on this feature map to output detection boxes and classification results. The specific parameters of the detection head layer are: convolution kernel size 3×3, stride 1, padding value 1. The twenty-third layer is the detection head layer. Its input is the feature map output by the twenty-second layer. It performs object detection branch processing on this feature map to output detection boxes and classification results. The specific parameters of the detection head layer are: convolution kernel size 3×3, stride 1, padding value 1. The twenty-fourth layer is the detection head layer. Its inputs are the feature maps output by the twenty-second and twenty-third layers. This layer performs detection branch processing on these feature maps to output detection boxes and classification results, and then generates a processed mask image based on these results. The specific parameters of the detection branch processing are: convolution kernel size 3×3, stride 1, padding value 1.

6. The method for mask target detection based on the improved YOLOWorld world model according to claim 5, wherein, The mask object detection model is trained through the following steps: (a1) Download an open-source mask dataset and divide this mask dataset into a training set and a test set according to a ratio of 8:

2. (a2) Perform data preprocessing on the training set obtained in step (a1) to obtain a preprocessed training set. (a3) Perform image enhancement processing on the preprocessed training set obtained in step (a2) to obtain an enhanced training set. For each sample in the enhanced training set (whose dimension is 512×512×3), input it into the first layer of the backbone network of the mask object detection model for processing to output a feature map corresponding to this sample with a dimension of 256×256×64. (a4) For each sample in the enhanced training set obtained in step (a3), input the feature map corresponding to this sample with a dimension of 256×256×64 obtained in step (a3) into the second layer of the backbone network of the mask object detection model for processing to output a feature map corresponding to this sample with a dimension of 128×128×128. (a5) For each sample in the enhanced training set obtained in step (a3), input the feature map corresponding to this sample with a dimension of 128×128×128 obtained in step (a4) into the third layer of the backbone network of the mask object detection model for multi-scale feature extraction to output a feature map corresponding to this sample with a dimension of 128×128×128. (a6) For each sample in the enhanced training set obtained in step (a3), input the feature map corresponding to this sample with a dimension of 128×128×128 obtained in step (a5) into the fourth layer of the backbone network of the mask object detection model for downsampling to output a feature map corresponding to this sample with a dimension of 64×64×256. (a7) For each sample in the augmented training set obtained in step (a3), the feature map corresponding to this sample with a dimension of 64×64×256 obtained in step (a6) is input into the 5th layer of the backbone network of the mask target detection model for multi-scale feature extraction, so as to output the feature map corresponding to this sample with a dimension of 64×64×256. (a8) For each sample in the augmented training set obtained in step (a3), the feature map corresponding to this sample with a dimension of 64×64×256 obtained in step (a7) is input into the 6th layer of the backbone network of the mask target detection model for downsampling, so as to output the feature map corresponding to this sample with a dimension of 32×32×512. (a9) For each sample in the augmented training set obtained in step (a3), the feature map corresponding to this sample with a dimension of 32×32×512 obtained in step (a8) is input into the 7th layer of the backbone network of the mask target detection model for multi-scale feature extraction, so as to output the feature map corresponding to this sample with a dimension of 32×32×512. (a10) For each sample in the augmented training set obtained in step (a3), the feature map corresponding to this sample with a dimension of 128×128×128 obtained in step (a5), the feature map corresponding to this sample with a dimension of 64×64×256 obtained in step (a7), and the feature map corresponding to this sample with a dimension of 32×32×512 obtained in step (a9) are input into the 8th layer of the backbone network of the mask target detection model for multi-scale feature extraction, so as to output the feature map corresponding to this sample with a dimension of 64×64×384. (a11) For each sample in the augmented training set obtained in step (a3), the feature map corresponding to this sample with a dimension of 64×64×384 obtained in step (a10) is input into the 9th layer of the backbone network of the mask target detection model for processing, so as to output the feature map corresponding to this sample with a dimension of 32×32×256. (a12) For each sample in the augmented training set obtained in step (a3), the feature map corresponding to this sample with a dimension of 32×32×512 obtained in step (a9) and each feature map with a dimension of 32×32×256 obtained in step (a11) are input into the 10th layer of the backbone network of the mask target detection model for processing, so as to output the feature map corresponding to this sample with a dimension of 32×32×768. (a13) For each sample in the augmented training set obtained in step (a3), the feature map corresponding to this sample with a dimension of 32×32×768 obtained in step (a12) is input into the 11th layer of the backbone network of the mask target detection model for processing, so as to output the feature map corresponding to this sample with a dimension of 32×32×512. (a14) For each sample in the enhanced training set obtained in step (a3), the feature map corresponding to this sample with a dimension of 64×64×384 obtained in step (a10) is input into the 12th layer of the backbone network of the mask target detection model for upsampling to output a feature map corresponding to this sample with a dimension of 128×128×384. (a15) For each sample in the enhanced training set obtained in step (a3), the feature map corresponding to this sample with a dimension of 128×128×128 obtained in step (a5) and each feature map with a dimension of 128×128×384 obtained in step (a14) are input into the 13th layer of the backbone network of the mask target detection model for processing to output a feature map corresponding to this sample with a dimension of 128×128×512. (a16) For each sample in the enhanced training set obtained in step (a3), the feature map corresponding to this sample with a dimension of 128×128×512 obtained in step (a15) is input into the 14th layer of the backbone network of the mask target detection model for processing to output a feature map corresponding to this sample with a dimension of 128×128×256. (a17) For each sample in the enhanced training set obtained in step (a3), the feature map corresponding to this sample with a dimension of 64×64×384 obtained in step (a10), each feature map with a dimension of 32×32×512 obtained in step (a13), and each feature map with a dimension of 128×128×256 obtained in step (a16) are input into the 15th layer of the backbone network of the mask target detection model for multi-scale feature extraction to output a feature map corresponding to this sample with a dimension of 64×64×576. (a18) For each sample in the enhanced training set obtained in step (a3), the feature map corresponding to this sample with a size of 64×64×576 obtained in step (a17) is input into the 16th layer of the backbone network of the mask target detection model for processing to output a feature map corresponding to this sample with a dimension of 32×32×256. (a19) For each sample in the enhanced training set obtained in step (a3), the feature map corresponding to this sample with a size of 32×32×256 obtained in step (a11), the feature map corresponding to this sample with a size of 32×32×512 obtained in step (a13), and the feature map corresponding to this sample with a size of 32×32×256 obtained in step (a18) are input into the 17th layer of the backbone network of the mask target detection model for processing to output a feature map corresponding to this sample with a dimension of 32×32×1024. (a20)For each sample in the augmented training set obtained in step (a3), the feature map of size 32×32×1024 corresponding to this sample obtained in step (a19) is input into the 18th layer of the backbone network of the mask target detection model for processing, so as to output a feature map of dimension 32×32×512 corresponding to this sample. (a21)For each sample in the augmented training set obtained in step (a3), the feature map of size 64×64×576 corresponding to this sample obtained in step (a17) is input into the 19th layer of the backbone network of the mask target detection model for upsampling, so as to output a feature map of dimension 128×128×576 corresponding to this sample. (a22)For each sample in the augmented training set obtained in step (a3), the feature map of size 128×128×384 corresponding to this sample obtained in step (a14), the feature map of size 128×128×256 corresponding to this sample obtained in step (a16), and the feature map of size 128×128×576 corresponding to this sample obtained in step (a21) are input into the 20th layer of the backbone network of the mask target detection model for processing, so as to output a feature map of dimension 128×128×1216 corresponding to this sample. (a23)For each sample in the augmented training set obtained in step (a3), the feature map of size 128×128×1216 corresponding to this sample obtained in step (a22) is input into the 21st layer of the backbone network of the mask target detection model for processing, so as to output a feature map of dimension 128×128×256 corresponding to this sample. (a24)For each sample in the augmented training set obtained in step (a3), the feature map of size 128×128×128 corresponding to this sample obtained in step (a5) is input into the 22nd layer of the backbone network of the mask target detection model for processing, so as to output a feature map of dimension 128×128×256 corresponding to this sample. (a25)For each sample in the augmented training set obtained in step (a3), the feature map of size 64×64×256 corresponding to this sample obtained in step (a7) is input into the 23rd layer of the backbone network of the mask target detection model for processing, so as to output a feature map of dimension 64×64×576 corresponding to this sample. (a26)For each sample in the augmented training set obtained in step (a3), the feature map of size 128×128×256 corresponding to this sample obtained in step (a23) and the feature map of size 128×128×256 corresponding to this sample obtained in step (a24) are input into the 24th layer of the backbone network of the mask target detection model for processing, so as to output a feature map of dimension 128×128×256 corresponding to this sample. (a27)For each sample in the enhanced training set obtained in step (a3), the feature map corresponding to this sample with a size of 64×64×576 obtained in step (a17) and the feature map corresponding to this sample with a size of 64×64×576 obtained in step (a25) are input into the 24th layer of the backbone network of the mask target detection model for processing, so as to output a feature map corresponding to this sample with a dimension of 64×64×576. (a28)For each sample in the enhanced training set obtained in step (a3), the feature map corresponding to this sample with a size of 32×32×512 obtained in step (a9) and the feature map corresponding to this sample with a size of 32×32×512 obtained in step (a20) are input into the 23rd layer of the backbone network of the mask target detection model for processing, so as to output a feature map corresponding to this sample with a dimension of 32×32×512. (a29)For each sample in the enhanced training set obtained in step (a3), the feature map corresponding to this sample with a size of 128×128×256 obtained in step (a26), the feature map corresponding to this sample with a size of 64×64×576 obtained in step (a27), and the feature map corresponding to this sample with a size of 32×32×512 obtained in step (a28) are input into the 24th layer of the backbone network of the mask target detection model for processing, so as to obtain the bounding box regression loss and classification loss corresponding to this sample. For each sample in the enhanced training set obtained in step (a3), according to the bounding box regression loss L corresponding to this sample obtained in step (a29) EIou and the classification loss L BCE , calculate the total loss L total , where λ1 and λ2 represent loss weights, and the value ranges of all three are positive real numbers. Preferably, they are equal to 0.5 and 1.5 respectively. (a31)For each sample in the enhanced training set obtained in step (a3), according to the total loss corresponding to this sample obtained in step (a30) and using gradient descent to iteratively train the mask target detection model until the mask target detection model reaches the preset number of iterations, and obtain the optimal parameters of the mask target detection model at this time, so as to obtain a preliminarily trained mask target detection model. (a32)Use the test set obtained in step (a1) to test the mask target detection model preliminarily trained in step (a31) until the obtained detection accuracy reaches the optimum, so as to obtain a finally trained mask target detection model.

7. The method for mask target detection based on the improved YOLOWorld world model according to claim 6, characterized in that, The image enhancement processing in step (a3) randomly selects one or any combination of the following 9 data enhancement methods: Image HSV enhancement, including hue, saturation, and brightness enhancement, and their enhancement factors are set to 0.015, 0.7, and 0.4 respectively; Image translation, and its translation factor is 0.1; Image scaling, and its scaling factor is 0.5; Image left - right flipping, and its flipping factor is 0.5; Multi - image stitching, and its stitching factor is 1.0; Image erasing, and its erasing factor is 0.4; Image cropping, and its cropping factor is 1.

0.

8. The mask target detection method based on the improved YOLOWorld world model according to claim 7, characterized in that Bounding box regression loss L EIou The EIoU loss is adopted, and its calculation formula is as follows: where b represents the center point coordinates of the predicted bounding box, b gt represents the center point coordinates of the ground truth bounding box, w represents the width of the predicted bounding box, w gt represents the width of the ground truth bounding box, h represents the height of the predicted bounding box, h gt represents the height of the ground truth bounding box, c represents the diagonal length of the smallest bounding rectangle that can contain both the predicted bounding box and the ground truth bounding box, C w represents the width of the smallest bounding rectangle that can contain both the predicted bounding box and the ground truth bounding box, C h represents the height of the smallest bounding rectangle that can contain both the predicted bounding box and the ground truth bounding box, p(w, w gt ) represents the deviation between the width of the predicted bounding box and the width of the ground truth bounding box, used to calculate the deviation in width between the predicted bounding box and the ground truth bounding box, p(h, h gt ) represents the deviation between the height of the predicted bounding box and the height of the ground truth bounding box, used to calculate the deviation in height between the predicted bounding box and the ground truth bounding box, p(b, b gt ) represents the deviation between the center point of the predicted bounding box and the center point of the ground truth bounding box, used to calculate the deviation in the center position between the predicted bounding box and the ground truth bounding box, L Iou represents the Intersection over Union (IoU) loss, used to measure the overlap degree between the predicted bounding box and the ground truth bounding box, and its calculation formula is: Among them, Area of Overlap represents the intersection area between the predicted bounding box and the ground truth bounding box, and Area of Union represents the union area between the predicted bounding box and the ground truth bounding box; The intersection area between the predicted bounding box and the ground truth bounding box, Area of Overlap = max(0, right - left) * max(0, bottom - top), where right and left are the right and left boundary coordinates of the intersection area between the predicted bounding box and the ground truth bounding box respectively, and bottom and top are the bottom and top boundary coordinates of the intersection area between the predicted bounding box and the ground truth bounding box respectively; The union area between the predicted bounding box and the ground truth bounding box, Area of Union = the area of the predicted bounding box + the area of the ground truth bounding box - the intersection area Area of Overlap.

9. The method for mask object detection based on the improved YOLOWorld world model according to claim 8, characterized in that Classification loss L BCE The binary cross-entropy loss is adopted, and its calculation formula is as follows: L BCE = -α * y * log(p) - β * (1 - y) * log(1 - p) where α and β are loss weights, and α + β = 1, the value range of α is from 0 to 1, preferably 0.5, the value range of β is from 0 to 1, preferably 0.5, y is the true classification value of the sample, and p is the predicted classification value of the sample. The calculation formula for the total loss is: L total = λ1 * L EIou + λ2 * L BCE where λ1 and λ2 are loss weights, used to balance the contribution of different loss terms to the total loss, the value range of λ1 is from 0 to 1, and the value range of λ2 is from 1 to 2.

10. A mask target detection system based on an improved YOLOWorld world model, characterized in that, Including: The first module is used to obtain the mask image to be detected; The second module is used to perform data preprocessing on the mask image to be detected obtained by the first module to obtain the mask image after data preprocessing; The third module is used to perform image processing on the mask image after data preprocessing by the second module to obtain the mask image after image processing; The fourth module is used to input the mask image after image processing obtained by the third module into the pre-trained mask object detection model to obtain the final detection result.

Citation Information

Cited By

  • Aero-engine gear damage analysis method and system based on image recognition

    CN120526098A

  • Traditional Chinese medicinal material target detection method and system based on improved YOLOV11

    CN121074358A