Two-dimensional code accurate detection method in complex scene based on GSSYOLO algorithm

By improving the network structure of the YOLOv8 algorithm and introducing the GEIT, SMPCGLU, and SEAM modules, the problem of inaccurate QR code boundary recognition in complex scenarios was solved, achieving higher detection accuracy and robustness.

CN121052269APending Publication Date: 2025-12-02NANJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511171246.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-20
Publication Date
2025-12-02

AI Technical Summary

Technical Problem

The existing YOLOv8 algorithm has difficulty accurately identifying the boundaries of QR codes in complex scenarios, especially under conditions of occlusion, damage, distortion, and background interference, which leads to a decrease in detection accuracy.

Method used

The GSS_YOLO algorithm is adopted, and the edge information perception capability is enhanced by introducing the GEIT module, the local feature capture capability is enhanced by using the SMPCGLU module, and the detection head of YOLOv8 is replaced with the SEAM module to improve the network structure to handle occluded scenes.

Benefits of technology

It significantly improves the accuracy and robustness of the model in QR code boundary recognition in complex scenarios, and enhances the accuracy and efficiency of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121052269A_ABST
    Figure CN121052269A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer vision and image recognition, and provides a method for accurately detecting a two-dimensional code in a complex scene based on a GSSYOLO algorithm. Comprising the following steps: collecting a two-dimensional code image, and constructing a data set; performing denoising preprocessing on the data set; marking the preprocessed image, marking the edge of the two-dimensional code, converting a marking result into a data set in a YOLO format, and dividing a training set, a verification set and a test set; random CoarseDropout and random Grid Distortion online data enhancement operation is carried out on the data set, and the random CoarseDropout and the random Grid Distortion online data enhancement operation is carried out; a GEIT module, an SMPCGLU module and an SEAM module are used for improving the YOLOv8, and a GSSYOLO model is obtained; training the improved GSSYOLO model by using a training set, and verifying the trained model by using a verification set; and testing the finally determined model by using the test set. Through training verification, the mAP50 of the GSSYOLO model on a test set reaches 97.9%, the boxloss is reduced to 0.3016, and the detection precision and robustness of the two-dimensional code in complex scenes such as dense overlapping, background interference and shielding are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision and image recognition technology, and proposes a method for accurate QR code detection in complex scenes based on the GSS_YOLO algorithm, which is particularly suitable for QR code detection in complex scenes with occlusion, damage, distortion, background interference, etc. Background Technology

[0002] In recent years, with the rapid development of the digital economy, QR code recognition technology has played a crucial role in industrial production, logistics management, intelligent warehousing, and equipment monitoring due to its advantages such as large information capacity, low cost, and convenient generation and recognition. Quickly and accurately detecting the position of the QR code in an image is the first and most critical step in the recognition process, directly determining the feasibility and efficiency of subsequent recognition. However, complex application scenarios in reality can easily affect the integrity and boundary areas of QR codes, making it difficult for related algorithms to accurately identify the boundaries of QR codes, posing a serious challenge to QR code target detection. The types of complex scenarios are as follows:

[0003] (1) Complex lighting environment and blur: strong light overexposure, uneven lighting (such as shadows), reflection and motion blur, etc.

[0004] (2) Geometric deformation: stretching, twisting, and wrinkling caused by curved surface labeling or flexible materials (such as cloth and rubber);

[0005] (3) Partial obstruction and damage: Physical obstruction or damage to the QR code itself;

[0006] (4) Multiple codes coexist: Multiple codes are adjacent, overlapping and sticking together in scenarios such as logistics stacking and dense shelves;

[0007] (5) Background interference: There are highly similar patterns (such as barcodes and text).

[0008] As a highly efficient object detection framework in the field of deep learning, the YOLO algorithm is widely used in various object detection fields due to its unique real-time processing capabilities and high accuracy. YOLOv8, in particular, boasts a superior speed-accuracy balance, a concise and efficient network architecture (such as the CSPDarknet backbone and SPPF module), and a more flexible training strategy, performing exceptionally well in multi-scale QR code object detection scenarios requiring rapid response. However, for QR code object detection in complex scenes, when key structural features in the image are missing or occluded, YOLOv8 may struggle to extract crucial features, leading to inaccurate edge localization. Furthermore, in situations with densely overlapping QR codes or similar background interference, the model's accuracy in recognizing QR code boundaries decreases. Summary of the Invention

[0009] To address the shortcomings of existing technical solutions, this invention provides a QR code target detection method based on the GSS_YOLO algorithm, which enables accurate QR code detection in complex backgrounds.

[0010] The objective of this invention is achieved through the following technical solution:

[0011] A method for accurate QR code detection in complex scenarios based on the GSS_YOLO algorithm includes the following steps:

[0012] S1: Collect QR code images and construct a dataset;

[0013] S2: Perform data preprocessing on the dataset;

[0014] S3: Label the preprocessed image, mark the edges of the QR code area, convert the labeling results into a YOLO format dataset, and divide it into training set, validation set and test set;

[0015] S4: Perform online data augmentation on the dataset;

[0016] S5: Establish a GSS_YOLO object detection model based on the YOLOv8 framework. The GSS_YOLO object detection model includes Backbone, Neck, and Head modules. The Backbone includes GEIT (Global Edge Information Transfer) and SMPCGLU (Self-moving Point Convolutional Gated Linear Unit) modules, and the Head includes SEAM (Separated and Enhancement Attention Module) module.

[0017] S6: Train the improved GSS_YOLO model using the training set, and validate the trained model using the validation set;

[0018] S7: Test the finalized model using the test set.

[0019] According to the present invention, a method for accurate QR code detection in complex scenarios based on the GSS_YOLO algorithm is provided, wherein the SMPCGLU module is used to replace the C2f module in the Backbone.

[0020] According to the present invention, a method for accurate QR code detection in complex scenarios based on the GSS_YOLO algorithm is provided, wherein the operation steps of the SMPCGLU module are as follows:

[0021] S11: Perform SMPConv convolution on the SMPCGLU input data to obtain the output data of the first module;

[0022] S12: Use the SiLU function to activate the output data of the first module to obtain the output data of the second module;

[0023] S13: Perform CGLU processing on the output data of the second module to obtain the output data of the third module;

[0024] S14: Perform Dorp Path regularization on the output data of the third module to obtain the output data of the fourth module;

[0025] S15: The output data of the fourth module is added to the input data through residual connection to obtain the output data of the SMPCGLU module.

[0026] According to the present invention, a method for accurate QR code detection in complex scenarios based on the GSS_YOLO algorithm is provided. The GEIT module is integrated into the Backbone. The GEIT module includes two parts: a multi-scale edge information generator and a convolutional edge fusion module.

[0027] According to the present invention, a method for accurate QR code detection in complex scenarios based on the GSS_YOLO algorithm is provided, wherein the operation steps of the GEIT module are as follows:

[0028] S21: The input data is subjected to SobelConv convolution in the Multi-Scale Edge Info Generator module to extract the edge features of the image; the SobelConv convolution extracts the edge features of the input data in the x and y directions respectively, and then fuses and superimposes them through addition operation to obtain a complete edge feature map;

[0029] S22: Max pool the complete edge feature map to obtain a double-downsampled edge feature map, perform a 1×1 convolution on the double-downsampled edge feature map to adjust the number of channels, and obtain the first output data of the Multi-Scale EdgeInfo Generator module;

[0030] S23: Max pool the double-downsampled edge feature map to obtain a quadruple-downsampled edge feature map, and perform a 1×1 convolution on the quadruple-downsampled edge feature map to adjust the number of channels, thereby obtaining the second output data of the Multi-Scale Edge Info Generator module;

[0031] S24: Max pool the four-fold downsampled edge feature map to obtain an eight-fold downsampled edge feature map, and perform a 1×1 convolution on the eight-fold downsampled edge feature map to adjust the number of channels, thereby obtaining the third output data of the Multi-Scale Edge Info Generator module;

[0032] S25: The three output data of the Multi-Scale Edge Info Generator module, together with the output data of the SMPCGLU module located in different network layers, are input into the Conv Edge Fusion module for cross-channel fusion;

[0033] S26: The Conv Edge Fusion module performs 1×1 convolution, 3×3 convolution, and 1×1 convolution on the input data in sequence to obtain the output data of the Conv Edge Fusion module.

[0034] According to the present invention, a method for accurate QR code detection in complex scenarios based on the GSS_YOLO algorithm is provided, wherein the SEAM module is used to replace the Head module.

[0035] According to the present invention, a method for accurate QR code detection in complex scenarios based on the GSS_YOLO algorithm is provided, wherein the operation steps of the SEAM module are as follows:

[0036] S31: Perform three independent CSMM depthwise separable convolutions in parallel on the input data of the SEAM module to generate feature maps of three channels; fuse the feature maps of the three channels and add them to the input data through residual connections to obtain the output data of the fifth module;

[0037] S32: Perform global average pooling on the output data of the fifth module to obtain the output data of the sixth module;

[0038] S33: Input the output data of the sixth module into a two-layer fully connected network to obtain the output data of the seventh module;

[0039] S34: Perform an exponential transformation on the output data of the seventh module to obtain the output data of the eighth module;

[0040] S35: Multiply the output data of the eighth module with the input data channel by channel to obtain the output data of the SEAM module.

[0041] The above-described one or more technical solutions in the embodiments of the present invention have at least one of the following technical effects:

[0042] This invention provides a method for accurate QR code detection in complex scenarios based on the GSS_YOLO algorithm. By introducing the GEIT module, SMPCGLU module, and SEAM module, it achieves the following technical advantages:

[0043] (1) The SMPCGLU module is used to replace the original C2f module, which enhances the network’s ability to capture local features in complex backgrounds and improves the detection rate. This significantly improves the problem of low accuracy of QR code boundary recognition under dense overlapping or similar background interference.

[0044] (2) The introduction of the GEIT module enhances the network’s ability to perceive edge information, effectively solving the problem of inaccurate QR code edge localization in complex scenarios;

[0045] (3) The detection head is improved by using the SEAM module, which enables the model to effectively handle occluded scenes and significantly improves the robustness of the model for QR code detection in occluded scenes. Attached Figure Description

[0046] Figure 1 The present invention provides a flowchart of a method for accurate QR code detection in complex scenarios based on the GSS_YOLO algorithm.

[0047] Figure 2 The network architecture diagram of the GSS_YOLO algorithm proposed in this invention.

[0048] Figure 3 Network structure diagram of the SMPCGLU module.

[0049] Figure 4 Network structure diagram of the GEIT module.

[0050] Figure 5 Network structure diagram of the SEAM module.

[0051] Figure 6 : Parameter information diagram of the GSS_YOLO model during training.

[0052] Figure 7 The results of this invention are compared with those of the original YOLOv8 and some mainstream improved models. Detailed Implementation

[0053] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. These embodiments are based on the technical solution of the present invention and provide detailed implementation methods and specific operating procedures. However, the scope of protection of the present invention is not limited to the following embodiments.

[0054] Example 1: This invention provides a method for accurate QR code detection in complex scenarios based on the GSS_YOLO algorithm, such as... Figure 1 As shown, the method includes the following steps:

[0055] S1: Collect QR code images and construct a dataset;

[0056] S2: Perform data preprocessing on the dataset;

[0057] S3: Label the preprocessed image, mark the edges of the QR code area, convert the labeling results into a YOLO format dataset, and divide it into training set, validation set and test set;

[0058] S4: Perform online data augmentation on the dataset;

[0059] S5: Establish a GSS_YOLO object detection model based on the YOLOv8 framework. The GSS_YOLO object detection model includes Backbone, Neck, and Head modules. The Backbone includes GEIT (Global Edge Information Transfer) and SMPCGLU (Self-moving Point Convolutional Gated Linear Unit) modules, and the Head includes SEAM (Separated and Enhancement Attention Module) module.

[0060] S6: Train the improved GSS_YOLO model using the training set, and validate the trained model using the validation set;

[0061] S7: Test the finalized model using the test set.

[0062] To improve recognition accuracy, further, in step S2, median filtering is used for data preprocessing to remove noise from the image.

[0063] To improve recognition accuracy, in step S3, labelme is used for annotation, and the ratio of the training set, validation set and test set is 7:2:1.

[0064] To improve the model's generalization performance, in step S4, random CoarseDropout and random GridDistortion online data augmentation operations are used to simulate occlusion, damage, distortion, and stretching scenarios of the QR code image, respectively.

[0065] To further improve recognition accuracy, in step S5, the GEIT module is introduced into the Backbone section to enhance the network's ability to perceive edge information; the SMPCGLU module is used to replace the C2f module of the Backbone, enhancing the network's ability to capture local features in complex backgrounds; and SEAM is used to replace the detection head in YOLOv8, enabling the model to effectively handle occluded scenes.

[0066] Specifically, the operating steps of the SMPCGLU module are as follows:

[0067] S11: Perform SMPConv convolution on the SMPCGLU input data to obtain the output data of the first module;

[0068] S12: Use the SiLU function to activate the output data of the first module to obtain the output data of the second module;

[0069] S13: Perform CGLU processing on the output data of the second module to obtain the output data of the third module;

[0070] S14: Perform Dorp Path regularization on the output data of the third module to obtain the output data of the fourth module;

[0071] S15: The output data of the fourth module is added to the input data through residual connection to obtain the output data of the SMPCGLU module.

[0072] Specifically, the GEIT module operates as follows:

[0073] S21: The input data is subjected to SobelConv convolution in the Multi-Scale Edge Info Generator module to extract the edge features of the image; the SobelConv convolution extracts the edge features of the input data in the x and y directions respectively, and then fuses and superimposes them through addition operation to obtain a complete edge feature map;

[0074] S22: Max pooling is performed on the complete edge feature map to obtain a double-downsampled edge feature map. A 1×1 convolution is performed on the double-downsampled edge feature map to adjust the number of channels, and the first output data of the Multi-Scale EdgeInfo Generator module is obtained.

[0075] S23: Max pool the double-downsampled edge feature map to obtain a quadruple-downsampled edge feature map, and perform a 1×1 convolution on the quadruple-downsampled edge feature map to adjust the number of channels, thereby obtaining the second output data of the Multi-Scale Edge Info Generator module;

[0076] S24: Max pool the four-fold downsampled edge feature map to obtain an eight-fold downsampled edge feature map, and perform a 1×1 convolution on the eight-fold downsampled edge feature map to adjust the number of channels, thereby obtaining the third output data of the Multi-Scale Edge Info Generator module;

[0077] S25: The three output data of the Multi-Scale Edge Info Generator module, together with the output data of the SMPCGLU module located in different network layers, are input into the Conv Edge Fusion module for cross-channel fusion;

[0078] S26: The Conv Edge Fusion module performs 1×1 convolution, 3×3 convolution, and 1×1 convolution on the input data in sequence to obtain the output data of the Conv Edge Fusion module.

[0079] Specifically, the steps for running the SEAM module are as follows:

[0080] S31: Perform three independent CSMM depthwise separable convolutions in parallel on the input data of the SEAM module to generate feature maps of three channels; fuse the feature maps of the three channels and add them to the input data through residual connections to obtain the output data of the fifth module;

[0081] S32: Perform global average pooling on the output data of the fifth module to obtain the output data of the sixth module;

[0082] S33: Input the output data of the sixth module into a two-layer fully connected network to obtain the output data of the seventh module;

[0083] S34: Perform an exponential transformation on the output data of the seventh module to obtain the output data of the eighth module;

[0084] S35: Multiply the output data of the eighth module with the input data channel by channel to obtain the output data of the SEAM module.

[0085] To improve recognition accuracy, further, in the model training step S6, the gradient is calculated using the backpropagation algorithm, and based on the set hyperparameters, the model's weights and biases are updated using a momentum-driven stochastic gradient descent (SGD) optimizer. Stable convergence is achieved by minimizing the loss function. The hyperparameters include the base learning rate, momentum coefficient, and number of iterations.

[0086] To improve recognition accuracy, further, in the model verification of step S6, the performance of the model is evaluated. If the performance does not meet expectations, the parameters of the model are adjusted, including the batch size and the number of iterations.

[0087] Example 2:

[0088] An application example of a method for accurate QR code detection in complex scenarios based on the GSS_YOLO algorithm according to Embodiment 1 of the present invention.

[0089] In this application example, such as the process Figure 1 As shown, it includes the following steps:

[0090] S1: Use a camera to capture QR code images and build a dataset containing a total of 1994 images.

[0091] S2: Preprocess the dataset by applying median filtering to denoise the images, effectively removing salt-and-pepper noise while preserving edge information. For two-dimensional images, the core principle of median filtering is to replace the original value of a pixel with the median of all pixel values ​​within its neighborhood window. The square neighborhood window W of pixel (i, j) is... i,j for:

[0092]

[0093] Where k is the side length of the square neighborhood window.

[0094] g(i,j)=med{f(ix,jy), (x,y∈W i,j )}

[0095] Where f(i,j) is the pixel value of the original pixel (i,j), g(i,j) is the pixel value after median filtering of the original pixel, and med is the median operation.

[0096] S3: Label the preprocessed images using labelme to accurately annotate the edges of the QR code images. Finally, convert the annotation results into a YOLO format dataset and divide it into a training set, a validation set, and a test set in a ratio of 7:2:1.

[0097] S4: To enhance the model's ability to recognize occluded, damaged, distorted, and stretched images, online random data augmentation CoarseDropout and GridDistortion operations are used during model training.

[0098] Online data augmentation refers to random augmentation operations performed in real time with specific probabilities each time data is read from the dataset during model training. This method ensures that the samples input to the model in each training epoch may be different augmented samples, thereby improving the model's generalization performance.

[0099] The CoarseDropout operation randomly removes rectangular regions of a specified size from an image to simulate occlusion and damage in QR code images. Dropout is a regularization technique that, during the training of a neural network, randomly sets the output of certain neurons to 0 with a certain probability (1-p). Neurons with outputs set to 0 do not participate in computation or update their weights during the current forward and backward propagation.

[0100] The feedforward operation of a standard neural network is as follows:

[0101] z i (l+1) =w i (l+1) y (l) +b i (l+1)

[0102] y i (l+1) =f(z) i (l+1) )

[0103] Where l∈{1,...,L} is the index of the hidden layer of the network, i is the hidden layer neuron, and z is the index of the hidden layer neuron. (l) Let y be the input vector of layer l. (l) w is the output vector of layer l. (l) and b (l) Let f be the weights and biases of layer l, and f be the activation function.

[0104] After Dropout, the feedforward operation is as follows:

[0105] r i (l) ~Bernoulli(p)

[0106]

[0107] y i (l+1) =f(z) i (l+1) )

[0108] p is the retention probability, r i (l) Let r be an independent and identically distributed Bernoulli random variable, with a probability of p that is 1 and a probability of 1-p that is 0. i(l) When r is 1, it means the i-th neuron is retained; when r is 1, it means the i-th neuron is retained. i (l) When r is 0, it indicates that the i-th neuron has been discarded. (l) Let be the mask vector of the l-th layer, and * denotes element-wise multiplication. This is a scaling factor used to keep the expectation constant. This is the output vector of layer l after the Dropout operation.

[0109] Dropout uses a binary (0 / 1) mask vector to randomly mask individual neurons. CoarseDropout builds upon this by using a binary (0 / 1) mask matrix to randomly mask consecutive rectangular regions. The size and number of masked rectangular regions can be controlled by parameters including num_holes_range, hole_height_range, and hole_width_range.

[0110] The GridDistortion operation distorts the image by moving grid nodes placed on it, simulating the distortion and stretching of a QR code image. First, the image is divided into a uniform grid, with grid intersections called control points. Then, a random horizontal or vertical displacement is applied to each control point, forming an irregular grid. Finally, based on the deformed grid, the pixel displacement of each non-control point is calculated using the following formula:

[0111]

[0112] Among them, C i Let ΔC be the four neighboring control points of pixel P. i w represents the displacement value of the nearest control point. i ΔP(x,y) is the inverse weight of the distance between pixel P and its neighboring control points, and ΔP(x,y) is the displacement value of pixel P.

[0113] S5: In the task of object detection, in order to overcome the limitations of the YOLOv8 model in QR code detection in complex backgrounds, this patent proposes the GSS_YOLO improved model, which significantly improves the detection accuracy and efficiency of QR codes in complex scenarios through multi-dimensional collaborative optimization.

[0114] Figure 2 The network structure diagram of the GSS_YOLO model is divided into three parts: Backbone, Neck (multi-scale feature fusion module), and Head (dynamic detection head).

[0115] Figure 3This is the structure diagram of the SMPCGLU (Self-moving Point Convolutional Gated Linear Unit) module. First, SMPConv represents continuous convolutional kernels using dynamically moving points, allowing these points to autonomously adjust their positions to efficiently capture the spatial features of the input. Next, the SiLU activation function is used to perform a non-linear transformation on the feature network. Then, the gating mechanism of CGLU dynamically adjusts the information flow of the convolutional features, thereby enhancing the model's ability to extract local features and improving computational efficiency. Drop Path regularization can randomly simplify the network structure and prevent overfitting. Finally, a residual connection is used to enhance feature representation while preserving the original input features.

[0116] Figure 4 This is a structural diagram of the GEIT (Global Edge Information Transfer) module, consisting of two parts: Multi-Scale Edge Info Generator and Conv Edge Fusion. In the Multi-Scale Edge Info Generator, the input image undergoes edge feature extraction via SobelConv convolution. The edge feature map is then subjected to three max-pooling operations and 1×1 convolutions for continuous downsampling by a factor of two, ultimately outputting edge feature information at three scales. The SobelConv convolution extracts edge features from the input image in both the x and y directions, then fuses and superimposes them using addition operations to obtain a complete edge feature map. The three scales of edge features output by the Multi-Scale Edge Info Generator module, along with the enhancement features output by the SMPCGLU module located in different network layers, are input to the Conv Edge Fusion module for cross-channel fusion. The Conv Edge Fusion module sequentially performs 1×1, 3×3, and 1×1 convolutions on the input data to achieve cross-channel fusion of edge features and enhancement features.

[0117] Figure 5This is the structure diagram of the SEAM (Separated and Enhancement Attention Module). The first part of SEAM is a CSMM depthwise separable convolution with residual connections, using three independent channels to achieve different convolution depths. The second part is a two-layer fully connected network used to aggregate information from each channel and enhance the connectivity between channels. The data obtained from the fully connected layer undergoes an exponential transformation, expanding the values ​​from [0, 1] to [1, e], making the result more tolerant to positional errors. Finally, the output of the SEAM module is multiplied by the original features, enabling the model to effectively solve the image occlusion problem.

[0118] S6: Train and validate the GSS_YOLO model using the training and test sets. During training, set the initial learning rate to 0.01 and the number of iterations to 100 rounds, and use the SGD optimizer.

[0119] S7: Use the test set to test the optimal model during the training process to ensure that the model can accurately and efficiently detect QR codes in complex backgrounds, providing reliable technical support for QR code detection.

[0120] Figure 6 This demonstrates the parameter changes of the GSS_YOLO model during dataset training. Figure 6 In this model, `train` and `val` represent the training and validation sets, respectively. `box_loss` (box regression loss) quantifies the positional deviation between the predicted and ground truth boxes. `precision` measures the model's prediction accuracy. `mAP50` represents the average precision when the IoU threshold is ≥0.5. The optimal model during training had a `box_loss` of 0.3016 and an `mAP50` of 0.979 on the test set.

[0121] Figure 7 The performance of the proposed GSS_YOLO model was compared with that of the original YOLOv8, YOLOv8-AR, YOLOv8-PIOU, and YOLOv8-LFI models on the mAP50 metric. Experimental results show that the proposed GSS_YOLO model achieves 97.9% mAP50, significantly outperforming the other comparative models. Therefore, compared to the original YOLOv8 and mainstream improved models, GSS_YOLO achieves a significant improvement in QR code detection accuracy.

Claims

1. A method for accurate QR code detection in complex scenarios based on the GSS_YOLO algorithm, characterized in that, Includes the following steps: S1: Collect QR code images and construct a dataset; S2: Perform data preprocessing on the dataset; S3: Label the preprocessed image, mark the edges of the QR code area, convert the labeling results into a YOLO format dataset, and divide it into training set, validation set and test set; S4: Perform online data augmentation on the dataset; S5: Establish a GSS_YOLO object detection model based on the YOLOv8 framework. The GSS_YOLO object detection model includes Backbone, Neck, and Head modules; the Backbone includes GEIT (Global Edge Information Transfer) and SMPCGLU (Self-moving Point Convolutional Gated Linear Unit) modules. The operating steps of the SMPCGLU module are as follows: S11: Perform SMPConv convolution on the SMPCGLU input data to obtain the output data of the first module; S12: Use the SiLU function to activate the output data of the first module to obtain the output data of the second module; S13: Perform CGLU processing on the output data of the second module to obtain the output data of the third module; S14: Perform Dorp Path regularization on the output data of the third module to obtain the output data of the fourth module; S15: The output data of the fourth module is added to the input data through residual connection to obtain the output data of the SMPCGLU module; The Head includes the SEAM (Separated and Enhancement Attention Module); S6: Train the improved GSS_YOLO model using the training set, and validate the trained model using the validation set; S7: Test the finalized model using the test set.

2. The method for accurate QR code detection in complex scenarios based on the GSS_YOLO algorithm according to claim 1, characterized in that, In step S2, median filtering is used to preprocess the dataset to remove noise.

3. The method for accurate QR code detection in complex scenarios based on the GSS_YOLO algorithm according to claim 1, characterized in that, In step S3, image annotation is performed using labelme, and the ratio of training set, validation set and test set is 7:2:

1.

4. The method for accurate QR code detection in complex scenarios based on the GSS_YOLO algorithm according to claim 1, characterized in that, In step S4, random CoarseDropout and random GridDistortion data augmentation operations are used.

5. The method for accurate QR code detection in complex scenarios based on the GSS_YOLO algorithm according to claim 1, characterized in that, In step S5, the SMPCGLU module is used to replace the C2f module in the Backbone.

6. The method for accurate QR code detection in complex scenarios based on the GSS_YOLO algorithm according to claim 1, characterized in that, In step S5, the GEIT module is integrated into the Backbone. The GEIT module comprises two parts: a multi-scale edge information generator and a convolutional edge fusion module.

7. The method for accurate QR code detection in complex scenarios based on the GSS_YOLO algorithm according to claim 6, characterized in that, The operation steps of the GEIT module are as follows: S21: The input data is subjected to SobelConv convolution in the Multi-Scale Edge Info Generator module to extract the edge features of the image; the SobelConv convolution extracts the edge features of the input data in the x and y directions respectively, and then fuses and superimposes them through addition operation to obtain a complete edge feature map; S22: Max pool the complete edge feature map to obtain a double-downsampled edge feature map, perform a 1×1 convolution on the double-downsampled edge feature map to adjust the number of channels, and obtain the first output data of the Multi-Scale Edge InfoGenerator module; S23: Max pool the double-downsampled edge feature map to obtain a quadruple-downsampled edge feature map, and perform a 1×1 convolution on the quadruple-downsampled edge feature map to adjust the number of channels, thereby obtaining the second output data of the Muttl-Scale EdgeInfo Generator module; S24: Max pool the four-fold downsampled edge feature map to obtain an eight-fold downsampled edge feature map, and perform a 1×1 convolution on the eight-fold downsampled edge feature map to adjust the number of channels, thereby obtaining the third output data of the Muttl-Scale EdgeInfo Generator module. S25: The three output data of the Muttl-Scale Edge Info Generator module, together with the output data of the SMPCGLU module located in different network layers, are input into the Conv Edge Fusion module for cross-channel fusion; S26: The Conv Edge Fusion module performs 1×1 convolution, 3×3 convolution, and 1×1 convolution on the input data in sequence to obtain the output data of the Conv Edge Fusion module.

8. The method for accurate QR code detection in complex scenarios based on the GSS_YOLO algorithm according to claim 1, characterized in that, In step S5, the SEAM module replaces the Head module. The operation steps of the SEAM module are as follows: S31: Perform three independent CSMM depthwise separable convolutions in parallel on the input data of the SEAM module to generate feature maps of three channels; fuse the feature maps of the three channels and add them to the input data through residual connections to obtain the output data of the fifth module; S32: Perform global average pooling on the output data of the fifth module to obtain the output data of the sixth module; S33: Input the output data of the sixth module into a two-layer fully connected network to obtain the output data of the seventh module; S34: Perform an exponential transformation on the output data of the seventh module to obtain the output data of the eighth module; S35: Multiply the output data of the eighth module with the input data channel by channel to obtain the output data of the SEAM module.

9. The method for accurate QR code detection in complex scenarios based on the GSS_YOLO algorithm according to claim 1, characterized in that, In the model training step S6, the gradient is calculated using the backpropagation algorithm, and based on the set hyperparameters, the model's weights and biases are updated using a momentum-driven stochastic gradient descent (SGD) optimizer. Stable convergence is achieved by minimizing the loss function. The hyperparameters include the base learning rate, momentum coefficient, and number of iterations.

10. The method for accurate QR code detection in complex scenarios based on the GSS_YOLO algorithm according to claim 1, characterized in that, In the model validation step S6, the performance of the model is evaluated. If the performance does not meet expectations, the model parameters are adjusted, including the batch size and the number of iterations.