Method for detecting Uyghur in natural scene based on image processing

By replacing the modules of the backbone network and neck network in the YOLOv8 network and combining with the GSConv module, the problem of high computational complexity of the Uyghur detection model on the mobile terminal is solved, and the effect of quickly returning accurate detection results on the mobile terminal is achieved.

CN119992569APending Publication Date: 2025-05-13DALIAN NATIONALITIES UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510063539.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the prior art, when the Uyghur detection model is deployed on the mobile terminal, the detection accuracy and model calculation complexity are contradictory, making it difficult to quickly return the detection results on the mobile terminal.

Method used

By replacing the C2f module of the backbone network in the original YOLOv8 network with the C2f-EMSC module and the C2f module of the neck network with the GS-Bottleneck module, combined with the GSConv module, an improved YOLOv8 network is built to improve feature extraction capabilities and reduce the computational volume.

Benefits of technology

While ensuring detection performance, the calculation amount is significantly reduced, so that the trained detection model can return more accurate detection results faster after it is deployed on the mobile terminal.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992569A_ABST
    Figure CN119992569A_ABST
Patent Text Reader

Abstract

The invention discloses a method for detecting Uygur in a natural scene based on image processing, and relates to the technical field of Uygur detection. A C2f module of a backbone network of an original YOLOv8 network is replaced by a C2f-EMSC module, a Conv module of a neck network is replaced by a GSConv module, and a C2f module of the neck network is replaced by a GS-Bottleneck, so that the improved YOLOv8 network is formed. A backbone network C2f-EMSC module extracts enough Uygur semantic information, so that a C2f module of a neck network is replaced by a GS-Bottleneck module, a Conv module is replaced by a GSConv module, and extraction of the Uygur semantic information in backbone network output features by the neck network is improved; according to the improvement, the calculation amount is reduced under the condition that the detection performance is ensured, so that the detection model can quickly return a more accurate detection result at the mobile terminal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of Uighur detection, and in particular to a method for detecting Uighur in a natural scene based on image processing. Background Art

[0002] Uyghur detection in natural scenes is the preliminary step of Uyghur recognition. Its purpose is to determine whether there is Uyghur text in different scene images (warning signs, street signs, etc.) so as to locate the position of Uyghur text and perform subsequent translation or information collection.

[0003] At present, the Uyghur detection method based on image processing in natural scenes based on deep learning has become the mainstream. Peng Yong et al. proposed "Uyghur detection in natural scenes based on improved single deep neural network". Li Lujingyi proposed a Uyghur detection method based on image processing in natural scenes based on improved YOLOV3 in "Design and Implementation of Uyghur Detection System in Natural Scenes". This method replaces Resblock with Denseblock and replaces the deep separable convolution with 3×3 ordinary convolution in the original network. By reducing the convolution parameters and the amount of calculation, the purpose of speed increase is achieved. Wang Deqing et al. proposed an improved DBNet algorithm in "A Review of Research on Scene Text Recognition Technology". The feature extraction of the network structure mainly adopts ResNeSt50 and feature pyramid network fusion, and then obtains the Uyghur text detection result through DB operation. Yiwen Wang et al. proposed a multi-directional scene Uyghur text detection model based on fine-grained feature representation and spatial feature fusion in "Scene uyghur recognition with embedded coordinate attention", which improved feature extraction and feature fusion and enhanced the ability of the network to represent multi-scale features. The above methods improve the deep learning network's ability to detect Uyghur in natural scenes from different perspectives.

[0004] However, these improved models will result in a large number of additional parameters and calculations, which leads to challenges when deployed on mobile devices. This is because when the Uyghur detection task is deployed on a mobile terminal, it must have high detection accuracy and be lightweight to ensure that the detection results can be returned quickly on the mobile terminal. Summary of the invention

[0005] Based on this, in order to solve the technical problem in the prior art that the detection accuracy and model calculation complexity are contradictory when the Uyghur detection model is deployed on a mobile terminal, the present invention provides a Uyghur detection method in natural scenes based on image processing.

[0006] The present invention provides a method for detecting Uyghur characters in natural scenes based on image processing, comprising:

[0007] The C2f module of the backbone network in the original YOLOv8 network is replaced with the C2f-EMSC module to form an improved backbone network; the C2f module of the neck network in the original YOLOv8 network is replaced with the GS-Bottleneck module, and the Conv module connected to the output end of the C2f module is replaced with the GSConv module to form an improved neck network; an improved YOLOv8 network including the improved backbone network, the improved neck network and the head network in the original YOLOv8 network is constructed;

[0008] Collect Uyghur images to build a dataset, use the dataset to train the improved YOLOv8 network, and obtain a Uyghur detection model;

[0009] The Uyghur image to be detected is input into the Uyghur detection model. The C2f-EMSC module in the backbone network is improved to divide the Uyghur image to be detected into different groups, and convolution operations of different scales are performed in parallel on each group to recognize the text in the Uyghur image to be detected at different scales, and a multi-channel Uyghur semantic feature map is obtained. The multi-channel Uyghur semantic feature map is cross-channel fused by improving the GS-Bottleneck module in the neck network to enhance the semantic association between different channels and obtain a Uyghur semantic fusion feature map. The local feature information of different channels in the Uyghur semantic fusion feature map is exchanged through the GSConv module to obtain a Uyghur semantic enhancement feature map. The Uyghur semantic enhancement feature map is detected by the head network to obtain the Uyghur position in the Uyghur image to be detected.

[0010] Furthermore, the C2f-EMSC module in the improved backbone network is used to divide the Uyghur images to be detected into different groups, and convolution operations of different scales are performed on each group in parallel to perform different scales of recognition on the text in the Uyghur images to be detected, and obtain a multi-channel Uyghur semantic feature map, which specifically includes:

[0011] The feature map input to the C2f-EMSC module is divided into shallow features X_cheap and features to be processed X_group, where the size of the features to be processed X_group is bs×(g×ch)×h×w, where bs represents the batch size, g represents the number of groups, ch represents the number of channels, h represents the height, and w represents the width;

[0012] The size of the feature to be processed X_group is rearranged to bs×ch×h×w×g to obtain several groups; convolution kernels of different sizes are applied to each group to capture Uyghur features of different scales;

[0013] The convolutional groups are rearranged to restore their sizes to obtain the restored features X'_group, and the restored features X'_group and the shallow features X_cheap are concatenated to obtain the concatenated feature map; the concatenated feature map is convolved through a 1x1 convolutional layer to obtain the output of the C2f-EMSC module.

[0014] Furthermore, the method of exchanging local feature information of different channels in the Uyghur semantic fusion feature map through the GSConv module to obtain the Uyghur semantic enhancement feature map specifically includes:

[0015] Perform standard convolution operation on the Uyghur semantic fusion feature map through the standard convolution layer in the GSConv module;

[0016] The output of the standard convolution layer is subjected to a depth-wise separable convolution operation through the depth-wise separable convolution layer in the GSConv module;

[0017] The output of the standard convolutional layer is concatenated with the output of the depthwise separable convolutional layer through the concatenation layer in the GSConv module;

[0018] The output of the concatenated layer is shuffled through the shuffle layer in the GSConv module to rearrange the channels in the concatenated feature map and output the Uyghur semantic enhanced feature map.

[0019] Furthermore, the replacing of the C2f module of the backbone network in the original YOLOv8 network with the C2f-EMSC module is to replace the third C2f module and the fourth C2f module of the backbone network in the original YOLOv8 network with the C2f-EMSC module.

[0020] Furthermore, replacing the C2f module with the GS-Bottleneck module is to replace the C2f module located after each Concat layer in the neck network with the GS-Bottleneck module.

[0021] Furthermore, the original YOLOv8 network is a YOLOv8n network.

[0022] At least one of the above technical solutions adopted by the present invention can achieve the following beneficial effects:

[0023] In the Uyghur detection method in natural scenes based on image processing provided by the present invention, the C2f module of the backbone network in the original YOLOv8 network is replaced with the C2f-EMSC module to improve the feature extraction capability of the YOLOv8 network while reducing the amount of calculation: the C2f-EMSC module applies convolution kernels of different sizes to each group to capture features of different scales, and the convolution operation of each group is performed in parallel, which can more efficiently utilize computing resources. The multi-channel Uyghur semantic feature map obtained by the C2f-EMSC module based on the backbone network contains enough Uyghur semantic information. Therefore, by replacing the C2f module of the neck network in the original YOLOv8 network with the GS-Bottleneck module and the Conv module connected to the output end of the C2f module with the GSConv module, the extraction of Uyghur semantic information in the output features of the backbone network by the neck network is improved. Specifically: the multi-channel Uyghur semantic feature map is cross-channel fused through the GS-Bottleneck module to enhance the semantic association between different channels; the local feature information of different channels in the Uyghur semantic fusion feature map is exchanged through the GSConv module to further enhance the Uyghur semantic feature map. During the enhancement process, since both the GS-Bottleneck module and the GSConv module have lightweight features, they will not cause computing resource burden during the enhancement process; the above improvements reduce the amount of calculation while ensuring the Uyghur detection performance, so that the trained detection model can be deployed after moving and return more accurate detection results faster. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0025] Figure 1 A schematic diagram of a flow chart of a method for detecting Uyghur characters in natural scenes based on image processing provided by the present invention;

[0026] Figure 2 A schematic diagram of SlimEMSC-YOLOv8 provided by the present invention;

[0027] Figure 3 A schematic diagram of the EMSC provided by the present invention;

[0028] Figure 4 A schematic diagram of the EMSC-Bottleneck provided by the present invention;

[0029] Figure 5 A schematic diagram of the C2f-EMSC provided by the present invention;

[0030] Figure 6 A schematic diagram of GSConv provided by the present invention;

[0031] Figure 7 A schematic diagram of one-shot aggregation provided by the present invention;

[0032] Figure 8 A schematic diagram of the GS-bottleneck provided by the present invention;

[0033] Fig. 9 This is a partial image example provided by the present invention, wherein Fig. 9 (a) is the door head. Fig. 9 (b) is a road sign. Fig. 9 (c) in the figure is a road sign;

[0034] Fig.10 The data enhancement result provided by the present invention, wherein Fig.10 (a) is the original image. Fig.10 (b) is the color perturbation result. Fig.10 (c) is the result of horizontal flipping. Fig.10 (d) in the figure is the result of adding noise;

[0035] Fig.11 A schematic diagram of the Uighur detection results provided by the present invention, wherein Fig.11 (a1)~ Fig.11 (c1) in the figure is the original image in three scenes. Fig.11 (a2)~ Fig.11 (c2) is the original YOLOv8 pair Fig.11 (a1)~ Fig.11 The result of Uyghur detection in (c1) is: Fig.11 (a3)~ Fig.11 (c3) is the improved YOLOv8 Fig.11 (a1)~ Fig.11 The result of Uyghur detection in (c1). DETAILED DESCRIPTION

[0036] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in combination with the specific embodiments of the present invention and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the specification, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0037] Uyghur is one of the ancient Chinese minority languages. It records the important national culture of the Uyghur people and is an important part of the splendid culture of the Chinese nation. Text detection in natural scenes is the preliminary link of text recognition in natural scenes. Its purpose is to determine whether there is text in different scene images (warning signs, street signs, etc.). If there is, the location of the text is located. After locating the location of the text, subsequent translation or information collection can be performed. At present, traditional optical character recognition technology is mainly used in document recognition, bill recognition, unmanned driving and other fields. For Uyghur text scenes with high resolution and single background color in documents, traditional OCR technology can be used. However, in natural scenes, due to the complex background and low contrast of Uyghur text in the image, it will be affected by lighting, obstacle occlusion, vertical distribution of fonts, curvature, small size, etc., which makes the Uyghur text in the image have problems such as text position, angle change, low resolution, etc. At the same time, there is a lack of corresponding Uyghur text datasets, which has caused great difficulties for Uyghur text detection research. Although some existing studies have improved the detection effect of Uyghur text, its model is large and the computational efficiency is low, and the detection performance still needs to be further improved.

[0038] Based on this, the present invention provides a method for detecting Uyghur in natural scenes based on image processing. The method is improved on the basis of the YOLOv8n (You Only Look Once Version 8) model, and a lightweight and efficient SlimEMSC-YOLOv8 model is proposed, which is mainly aimed at realizing efficient monitoring of Uyghur in natural environments. The model uses efficient multi-scale extraction convolution to replace the standard convolution in the YOLOv8n model C2f (CSP Bottleneck with 2Convolutions), and uses Slimneck constructed by GSconv to replace the neck part in the YOLOv8n model.

[0039] Example 1

[0040] Figure 1 The flow chart of the method for detecting Uyghur characters in natural scenes based on image processing in this embodiment is shown. Figure 1 The method is described in detail and specifically comprises the following steps:

[0041] S1: Replace the C2f module of the backbone network in the original YOLOv8 network with the C2f-EMSC module to form an improved backbone network; replace the C2f module of the neck network in the original YOLOv8 network with the GS-Bottleneck module, and replace the Conv module connected to the output end of the C2f module with the GSConv module to form an improved neck network; construct an improved YOLOv8 network including the improved backbone network, the improved neck network and the head network in the original YOLOv8 network.

[0042] The YOLOv8 model is proposed on the basis of the YOLOv5 model. The YOLOv8 network structure includes four parts: Input, Backbone, Neck, and Prediction. First, the Backbone module continues to use the CSP idea for feature transfer, and introduces the C2f module to replace the C3 module in the YOLOv5 model. The SPPF module is also retained, which can obtain rich gradient flow information while optimizing the network structure. Secondly, the FPN+PAN structure is continued in the Neck module to achieve multi-scale feature fusion capabilities. Finally, in the Prediction part, the original coupling head is changed to a decoupling head, and the Decoupled head is introduced to separate classification and regression into two separate structures, which improves the model convergence ability. YOLOv8 uses the Anchor Free network, which solves the problem of uneven positive and negative samples by dynamically allocating positive and negative samples. The YOLOv8 model is divided into four models: YOLOv8l, YOLOv8m, YOLOv8s, and YOLOv8n. Among them, YOLOv8n is the least complex model. It has a faster reasoning speed while maintaining high detection accuracy, and is easy to deploy on mobile and embedded devices. Considering that the Uyghur detection task may need to be deployed on mobile terminals (such as cars, mobile phones for sign translation, map information collection, etc.) to meet the requirements of lightweight and high detection accuracy, YOLOv8n was selected as the baseline model for improvement to achieve Uyghur detection in natural scenes. Therefore, improvements are made on the basis of the YOLOv8n model, and a lightweight and efficient SlimEMSC-YOLOv8 model is proposed. The SlimEMSC-YOLOv8 model is mainly aimed at achieving efficient monitoring of Uyghur in natural environments.

[0043] Since Uyghur characters in natural scenes are written in various ways and have complex structures, they are of different sizes and are affected by multiple complex factors such as light changes and dynamic backgrounds, which can easily lead to missed detection and false detection. In order to overcome the problems of low accuracy, large number of model parameters, high computational overhead, and difficult deployment for Uyghur detection tasks, targeted improvements are made based on the YOLOv8n algorithm. First, a multi-scale feature extraction convolution (Efficient Multi-Scale Convolution, EMSConv) is designed. On this basis, the C2f-EMSC module is designed to replace the C2f module in the backbone network Backbone, capturing the multi-scale feature information of Uyghur characters in natural scenes with lower computational overhead. Secondly, in view of the problem of large number of parameters in the original feature fusion network PANet (Path Aggregation Network) of YOLOv8n, an improved method of using Slimneck to replace the neck part of the original model is proposed to reduce the parameters and computational overhead of the network and adapt to the detection needs of different platforms. Figure 2 The overall architecture of the improved algorithm SlimEMSC-YOLOv8 is shown.

[0044] S101: Backbone network improvement.

[0045] Traditional convolution methods only use convolution kernels of a single scale, which makes it difficult to recognize objects or scene features of different scales. However, due to the use of convolution kernels of multiple scales, multi-scale convolution can more effectively obtain features of various scales, thereby extracting more feature information. Traditional target detection algorithms (such as Faster R-CNN) usually use single-scale feature maps for detection, which causes the detector to perform poorly for small objects or distant objects. Therefore, Tsung-Yi Lin et al. proposed the feature pyramid network FPN

[11] , which enables the detector to better handle objects of different scales. The goal is to obtain the function of extracting multi-scale information in convolution similar to the feature pyramid. In order to optimize the convolution structure, a more sophisticated and lower-parameter multi-scale extraction convolution architecture EMSConv is introduced, such as Figure 3 shown.

[0046] The input feature map is first divided into two parts according to the number of channels, namely the shallow feature X_cheap and the feature to be processed X_group, such as Figure 2As shown in the figure, the features to be processed are rearranged from bs×(g×ch)×h×w to bs×ch×h×w×g (where bs represents the batch size, g represents the number of groups, ch represents the number of channels, h represents the height, and w represents the width). The purpose of this is to be able to perform independent convolution operations on each group. Through this split, different sizes of convolution kernels can be applied to each group to capture features of different scales. In this process, the convolution operation of each group is performed in parallel, so this method can make more efficient use of computing resources. The feature map after convolution is rearranged to restore the original shape bs×(g×ch)×h×w to ensure that it can be spliced ​​with the shallow features. Then it is connected to the convolution X_group through a CONCAT layer X_cheap. Finally, the connected feature map is processed by a 1x1 convolution layer to obtain the final output feature map. The main purpose of rearranging the tensor is to enable each convolution kernel to independently process different parts of the input tensor and correctly combine these parts after processing.

[0047] The core idea of ​​multi-scale extraction convolution is to divide the input feature map into multiple groups, and apply different convolution kernels to each group for convolution, which can capture the feature information of different scales. Finally, the obtained feature maps are combined. By fusing features of multiple scales, the impact of information loss can be reduced. In this process, the number of parameters of each convolution kernel is relatively small. Compared with standard convolution, due to the use of smaller convolution kernels, the computational complexity of multi-scale extraction convolution is relatively small, so it can save parameters and reduce the complexity of the model.

[0048] When the size of the convolution kernel of the standard convolution used is 3×3, the ratio of FLOPs of the standard convolution to the multi-scale extraction convolution is as follows:

[0049]

[0050] From the results, we can see that the computational cost of the multi-scale extraction convolution is only one-third of that of the standard convolution. Therefore, this convolution can be used instead of the standard convolution to construct the C2f module.

[0051] First, we design EMSC-Bottleneck based on the Bottleneck module in C2f in YOLOV8, and replace the second convolutional layer in the Bottleneck module in C2f in YOLOV8 with EMSConv, so that the advantages of multi-scale extraction convolution can be continued in the Bottleneck module. The structure of EMSC-Bottleneck is as follows Figure 4 shown.

[0052] Based on EMSC-Bottleneck, we can design the C2f-EMSC module, replacing the original Bottleneck module in YOLOv8 with EMSC-Bottleneck. The workflow of C2f-EMSC is similar to that of C2f. The structural diagram of C2f-EMSC is shown in the figure below. Figure 5 shown.

[0053] Compared with the ordinary C2f module, C2f_EMSC applies multiple Bottleneck_EMSCs. When the number of parameters of each layer of the output network is taken into account, the number of parameters of C2f and C2f_EMSC in the sixth layer is 197632 and 149632 respectively, and the number of parameters of C2f and C2f_EMSC in the eighth layer is 460288 and 364160 respectively, and the number of parameters is reduced by 48000 and 96128 respectively when C2f_EMSC is used. Therefore, by introducing the multi-scale extraction convolution operation, C2f_EMSC enhances the feature extraction capability of the model, inherits the advantages of multi-scale extraction convolution, and reduces the parameters and computational overhead of the model.

[0054] S102: Neck network improvements.

[0055] The computational complexity of standard convolution is relatively large, so when standard convolution is used continuously in a model, it usually makes the model larger. A large model on an edge device may result in reduced accuracy or inability to deploy. At this point, some previous studies may choose depthwise separable convolution (DSC) instead of standard convolution to be deployed in edge devices. Depthwise separable convolution decomposes the convolution operation into two independent steps, namely depthwise convolution and pointwise convolution. Depthwise convolution is used to extract spatial features, and pointwise convolution is used to extract channel features. In addition, depthwise separable convolution uses smaller convolution kernels, which greatly reduces the number of parameters and computation of the model. This makes it more advantageous in environments with limited computing resources, such as mobile devices and embedded systems. However, a large number of depthwise separable convolution layers cannot achieve sufficient accuracy to build a lightweight model, and due to memory bandwidth limitations, depthwise separable convolution is not fast on GPUs, mainly because of high memory access. Moreover, the defects of DSC are also obvious: the channel information of the input image is separated during the calculation process.

[0056] In order to make the output of DSC as close as possible to SC, a new method GSConv is proposed, which is a convolution that mixes the three together by connecting the standard convolution and the depthwise separable convolution in parallel and then connecting them in series with the shuffle operation. The structure of GSConv is as follows Figure 6 shown.

[0057] First, GSConv performs a downsampling operation on the input through a normal convolution, and then uses DWConv to extract features through deep convolution. After this operation, the results of the two Convs are concatenated, and finally a shuffle operation is performed to bring the corresponding channels of the feature maps output by the previous two convolutions together. This is done to completely mix the information from the two convolutions by evenly exchanging the local feature information on different channels.

[0058] When using standard convolution at the 16th and 19th layers of the model, the number of parameters is 36992 and 147712 respectively, while the number of parameters of GSConv at the same positions is 19360 and 75584 respectively. Therefore, compared with standard convolution, GSConv has about half the number of parameters of the former, and maintains similar learning ability with lower computational overhead, but the task of reducing inference time and maintaining accuracy requires a new model.

[0059] GS-bottleneck is built based on GSConv, using one-shot aggregation in VoVNet (such as Figure 7 The GS-bottleneck module is designed as shown in the figure. Figure 8 As shown in the figure. This design has a good balance between the size of the model and the accuracy of the network. When applied to the YOLOv8n model, GS-bottleneck is used to replace the C2f module in the neck part. Compared with the C2f module, GS-bottleneck reduces the computational requirements of the model, reduces the model burden while maintaining the model accuracy, and achieves a good balance between model accuracy and model lightweight.

[0060] In addition to achieving lightweight, the C2f-EMSC module also compensates for GSConv, because the essence of GSConv is the superposition of depthwise separable convolution and standard convolution, which inevitably has the problem of loss of precision due to layer-by-layer stacking of depthwise separable convolution, and due to bandwidth limitations, depthwise separable convolution has the problem of excessive memory access, which makes it not fast on the GPU. Therefore, in the C2f module, EMSConv is used to reduce the memory access of the model and convolution kernels of multiple scales are used to improve the detection accuracy of the model, which better compensates for the loss of precision and excessive memory access in the feature enhancement network, that is, the neck part, due to the use of more GSConv.

[0061] In order to speed up the prediction calculation, images processed by convolutional neural networks (CNNs) almost always pass spatial information into channels step by step in the backbone network. Each time the spatial (width and height) compression and expansion of the channel of the feature map will result in the loss of some semantic information. Dense convolution calculations retain the hidden connections between each channel as much as possible, while sparse convolutions completely cut off these connections. GSConv makes a compromise between the two, but if it is used to replace all standard convolutions in the network, the network layer will become deeper, which will significantly increase the inference time. However, when processing these feature maps in the neck, the spatial information is gradually compressed to the minimum, and the number of channels is gradually increased to the highest. At this time, the feature map already contains rich semantic information, and further spatial compression and channel expansion are no longer necessary. Therefore, if Figure 2 As shown in the figure, in the neck part of the YOLOv8n model, the original Conv is replaced with GSConv, and the original C2f module is replaced with GS-Bottleneck, to build a neck - Slimneck with lower parameters and calculation amount and maintaining accuracy compared with the yolov8n model, so as to reduce the calculation complexity and parameters while improving the detection accuracy.

[0062] S2: Collect Uyghur images to build a dataset, use the dataset to train the improved YOLOv8 network, and obtain the Uyghur detection model.

[0063] S201: Build a Uyghur dataset.

[0064] The Uyghur characters in the pictures are used as the detection target, so a large number of pictures containing Uyghur characters are needed as the data set for experimental research. Since there is no data set suitable for Uyghur target detection in the existing public data set, we collect real natural scene pictures containing Uyghur characters to create a data set. The collected pictures include door heads, plaques, road name signs, road signs, book covers, etc., and the final real data set collected is 2693 pictures, such as Fig. 9 As shown, Fig. 9 (a) is the door head. Fig. 9 (b) is a road sign. Fig. 9 (c) in the figure is a road sign. The texts have different sizes, colors and sizes and are distributed in different positions in the image. In this experiment, the images are unified into a size of 640×640 to facilitate SlimEMSC-YOLOv8 processing. All images in this dataset are labeled with text areas using the Label Image tool. After data enhancement, the data contains 5211 Uyghur datasets in natural scenes. The description of the training sample and test sample set is shown in Table 1. The categories of the dataset are divided into two categories: Uyghur and Chinese.

[0065] Table 1 Description of training samples and test sample sets

[0066] type Resolution quantity Training samples 640×640 2114 natural scene images and 2518 data augmented images Test samples 640×640 579 images of natural scenes

[0067] Enhance the collected images. Data augmentation refers to the technology of generating new data by slightly modifying existing data or merging it with other data. The purpose of this technology is to increase the amount of data. Data augmentation technology has made great contributions to improving the generalization performance of the network and avoiding overfitting of neural networks. In addition, when the training data set samples are scarce, the number of data set samples can be expanded through data augmentation methods to construct a high-quality, more realistic data set that conforms to actual application scenarios. In order to expand the relevant data sets required for model training and enhance the generalization ability and robustness of the model, since the text on road name signs, road signs, book covers, etc. is relatively simple, data enhancement is performed on doorheads and plaques. The data augmentation methods used are as follows: Fig. 9 As shown, color perturbation, horizontal flipping, and noise addition are performed.

[0068] (1) Color disturbance

[0069] Color perturbation is a data augmentation technique used to increase the diversity of training data in machine learning. In the RGB color space, color perturbation can be achieved by adding or removing some colors, or by changing the order of color channels. In the process of Uyghur detection, due to factors such as lighting and weather, images taken at different times of the day will also be different. Therefore, by simulating the lighting changes of images through color perturbation, the diversity of training data can be effectively increased, and the effect is as follows: Fig.10 As shown in (b) in the figure. Assuming that R0 is the original RGB value of the Uyghur image, d represents the brightness transformation coefficient, and R represents the RGB value after the image operation, the adjusted calculation formula is:

[0070] R=R0×(1+d).

[0071] (2) Horizontal flip

[0072] Image flipping can change the display direction of an image. Common image flipping methods include horizontal flipping and vertical flipping. Horizontal flipping refers to performing a symmetrical operation on the image along the horizontal axis, so that the contents on the left and right sides of the image are swapped. This is like the image seen in a mirror. Vertical flipping is a symmetrical operation along the vertical axis, resulting in the exchange of the contents of the upper and lower parts of the image. Because the research object is Uyghur, there will be no vertical flipping, so only horizontal flipping is used, and the effect is as follows: Fig.10 As shown in (c) in .

[0073] Assuming that the size of the input Uyghur image is M×N, the coordinates of a pixel in the original image of the dataset are P0(x, y), and the coordinates after horizontal flipping are P0(x*, y*), then the coordinate transformation of the input Uyghur image after horizontal flipping can be described by formula (3):

[0074]

[0075] (3) Noise enhancement

[0076] In the field of image processing, salt and pepper noise and Gaussian noise are two common types of noise. Their basic principle is to enhance the image by changing the value of each pixel. Specifically, salt and pepper noise is a random noise type that randomly adds some black or white pixels to the image, making the image look more vivid and interesting. Gaussian noise is a continuous noise type that adds some Gaussian distributed noise points to the image, making the image look smoother and more natural. In the process of Uyghur detection, it is often disturbed by noise. By adding noise to the Uyghur image, the real scene can be better simulated, and the effect is as follows: Fig.10 As shown in (d) in .

[0077] S202: Model training environment settings.

[0078] The Pytorch framework is used to build the model. All models are trained on a Nvidia RTX 4090 on a Linux CentOS 7 operating system. The hyperparameters are as follows: the training step size is 300; the optimizer is stochastic gradient descent; the batch size is 30.

[0079] S203: Model training indicator settings.

[0080] We selected the overall model parameters, GFLOPs, recall (R), accuracy (P) and mean average precision (mAP) as the main evaluation indicators of the model. The number of parameters is often an important indicator to measure the size of a model. Models with large parameters are difficult to deploy on some edge platforms. FLOPs is the abbreviation of floating point of operations, which is the number of floating-point operations and can be used to measure the complexity of the algorithm / model. GFLOPs (Giga Floating-point Operations Per Second) means "billion floating-point operations per second". The Giga here means 10^9, that is, one billion. GFLOPs is often used as a reflection of the amount of computing. The lower the GFLOPs, the lower the amount of computing of the model; it is also often used as a GPU performance parameter, but it does not necessarily represent the actual performance of the GPU.

[0081] The calculation of average precision mAP involves the calculation of accuracy and recall. The calculation formulas for accuracy P and recall R are:

[0082]

[0083]

[0084] Here, TP (true positive) refers to the number of correctly predicted positive samples; FP (false positive) refers to the number of incorrectly predicted positive samples; and FN (false negative) represents the number of incorrectly predicted negative samples. The calculation formula for the average precision (AP) for all categories is as follows:

[0085]

[0086] Where m represents the number of positive samples, P(r) is the proportion of positive samples in the first r search results, ΔR(r) is the change in recall rate relative to r in the first r search results, P(R) represents the precision (P) under the feature recall rate (R), and AP represents the average value under the precision-recall rate (PR) curve, which represents the accuracy under different recall rates. To calculate AP, a numerical integration method is used. The AP of all categories is summed and the average is taken to get mAP, which is the model average precision.

[0087]

[0088] Where n is the number of real categories, AP iis the average precision of the i-th category. mAP is divided into mAP0.5 and mAP0.5:0.95.

[0089] S204: Ablation experiment

[0090] In the experiment, each improvement method is introduced one by one, and its impact on the model performance is observed. The experiment is divided into 4 groups, and each group maintains the consistency of the input image and the training hyperparameters to ensure the fairness and accuracy of the experimental results. C2f_EMSC and Slimneck are the improvement methods I proposed in the paper. The "√" in the table indicates that they are used. The results are shown in Table 2:

[0091] Table 2 Ablation experiment results

[0092]

[0093] (*Note: a means using C2f_EMSC; b means using Slimneck to replace backbone)

[0094] As shown in Table 2, the experimental group with serial number 1 is the SlimEMSC-YOLOv8 algorithm, which is 91.7%, with a parameter amount of 2651798 and a GFLOPs of 7.0. The second group removes the backbone part and uses the C2f module modified by multi-scale extraction convolution to replace the original C2f module, using the original C2f module with standard convolution of the YOLOv8 model, and retains the Slimneck of the neck part, which is 91.6%, with a parameter amount of 2861782 and a GFLOPs of 7.9. It can be seen that compared with the SlimEMSC-YOLOv8 algorithm, the detection accuracy has dropped by one percentage point, while the parameter amount has increased by 5% and the GFLOPs has increased by 4%. In the third group, the Slimneck used in the neck part is removed, the original neck is used, and C2f-EMSC is retained. Compared with the SlimEMSC-YOLOv8 algorithm, it has dropped by 0.1%, and the parameter amount and GFLOPs have increased by about 7% and 13%, respectively. The last group is YOLOv8n without any modification. Compared with SlimEMSC-YOLOv8, it is reduced by 0.3%, and the number of parameters is 3006038, which is 13% more than the algorithm, and GFLOPs is also increased by about 16%. This proves that each improved method is effective for Uyghur detection in natural scenes, and the number of parameters and calculations are significantly reduced while the detection accuracy is improved.

[0095] S205: Comparative test

[0096] The algorithm is compared with models with smaller parameters in mainstream object detection algorithms such as YOLOv5n, YOLOv7-tiny, YOLOv8n, DETR, and RT-DETR-l. Table 3 shows the experimental results.

[0097] Table 3 Comparative experimental results

[0098]

[0099] According to the experimental comparison data above, it can be concluded that the algorithm has a higher index than other mainstream target detection algorithms on the Uyghur dataset in natural scenes, reaching 91.7% when ; and 67.8% when ; at the same time, the algorithm has much fewer parameters and calculations than other mainstream target detection algorithms. In summary, the algorithm is more suitable for Uyghur detection deployed in edge platforms.

[0100] S206: Detection effect comparison

[0101] To provide an intuitive demonstration of the model's detection capabilities, we selected the benchmark model YOLOv8n for comparison. To ensure the comprehensiveness of the detection, we selected three images, covering Uyghur and Chinese detection. The detection results are as follows: Fig.11 As shown, Fig.11 (a1)~ Fig.11 (c1) in the figure is the original image in three scenes. Fig.11 (a2)~ Fig.11 (c2) is the original YOLOv8 pair Fig.11 (a1)~ Fig.11 The result of Uyghur detection in (c1) is: Fig.11 (a3)~ Fig.11 (c3) is the improved YOLOv8 Fig.11 (a1)~ Fig.11The result of Uyghur detection in (c1) in the figure. As can be seen from the figure, in the detection process of the first picture, SlimEMSC-YOLOv8 detected all the Chinese and Uyghur characters that were not blocked in the picture, while the YOLOv8n model missed one Chinese and one Uyghur character and also misdetected a digital part; in the detection process of the second picture, SlimEMSC-YOLOv8 had a good effect on italic text detection, and detected all the Chinese and Uyghur characters in the picture, while the YOLOv8n model missed all the Uyghur characters and most of the Chinese characters above the middle plaque; in the third picture, the YOLOv8n model missed an obvious Uyghur character, but SlimEMSC-YOLOv8 detected it, and the confidence of the detected part was higher than that of the YOLOv8n model. From the comparison of detection results, it can be seen that the detection effect of SlimEMSC-YOLOv8 algorithm is better than that of the original YOLOv8n algorithm, and SlimEMSC-YOLOv8 has improved the missing detection problem in the original YOLOv8n to a certain extent.

[0102] S3: Input the Uyghur image to be detected into the Uyghur detection model, divide the Uyghur image to be detected into different groups by improving the C2f-EMSC module in the backbone network, and perform convolution operations of different scales on each group in parallel to recognize the text in the Uyghur image to be detected at different scales, and obtain a multi-channel Uyghur semantic feature map; perform cross-channel fusion of the multi-channel Uyghur semantic feature map by improving the GS-Bottleneck module in the neck network to enhance the semantic association between different channels and obtain a Uyghur semantic fusion feature map; exchange the local feature information of different channels in the Uyghur semantic fusion feature map through the GSConv module to obtain a Uyghur semantic enhancement feature map; detect the Uyghur semantic enhancement feature map through the head network to obtain the Uyghur position in the Uyghur image to be detected.

[0103] Aiming at the Uyghur detection task in natural scenes, a lightweight Uyghur detection method in natural scenes based on image processing, SlimEMSC-YOLOv8, is proposed. Firstly, by replacing the standard convolution in C2f with multi-scale extraction convolution in the backbone part, the feature extraction capability is improved while reducing about 4.8% of parameters and 3% of computational complexity; then, Slimneck is used to replace the neck of the YOLOv8n model to reduce 7% of parameters and 10% of computational complexity while improving the accuracy by 0.2%, so as to balance the accuracy of the model detection text and the model size. The experimental results on the Uyghur dataset in natural scenes show that the improved algorithm reduces the accuracy and computational complexity by 12% and 14% respectively, and improves the accuracy by 0.3%. In future research, the impact of different initial grouping methods of multi-scale extraction convolution on the accuracy and the problem that the accuracy of the model may be affected when too many GSConvs are used should be further considered.

[0104] The main contributions of the present invention are as follows:

[0105] 1) A high-quality, diverse, and standardized Uyghur natural scene image dataset was established to address the lack of datasets in Uyghur detection research in natural scenes. The natural scene images collected in this paper include door heads, plaques, road signs, etc. In order to make the dataset more useful for model training, this paper combines a variety of data enhancement methods to expand the dataset, including color perturbation, horizontal flipping, and adding noise. The final dataset has more than 5,000 images, providing data support for Uyghur detection research in natural scenes.

[0106] 2) In view of the problems of complex structure of existing Uyghur fonts, difficulty in extracting Uyghur features, and easy missed detection of targets, a detection method for Uyghur in natural scenes based on YOLOV8 is studied. An improved group convolution module EMSConv is fused with the C2f module in the YOLOV8 model to replace the C2f module in the backbone part of YOLOV8 to obtain more features from the input image.

[0107] 3) As the YOLOV8 model is large and difficult to deploy on some mobile platforms, Slimneck composed of GSConv is used to improve the neck part of the original model to make the model lightweight while minimizing the loss of model accuracy and making it more suitable for practical application scenarios.

[0108] The above is a method for detecting Uyghur characters in natural scenes based on image processing provided by one or more embodiments of the present invention. Based on the same idea, the present invention also provides a corresponding device for detecting Uyghur characters in natural scenes based on image processing, including:

[0109] A model building module is used to replace the C2f module of the backbone network in the original YOLOv8 network with the C2f-EMSC module to form an improved backbone network; replace the C2f module of the neck network in the original YOLOv8 network with the GS-Bottleneck module, and replace the Conv module connected to the output end of the C2f module with the GSConv module to form an improved neck network; and construct an improved YOLOv8 network including the improved backbone network, the improved neck network, and the head network in the original YOLOv8 network.

[0110] The model training module is used to collect Uyghur images to build a data set, and use the data set to train the improved YOLOv8 network to obtain a Uyghur detection model.

[0111] The detection module is used to input the Uyghur image to be detected into the Uyghur detection model, divide the Uyghur image to be detected into different groups by improving the C2f-EMSC module in the backbone network, and perform convolution operations of different scales on each group in parallel to recognize the text in the Uyghur image to be detected at different scales, and obtain a multi-channel Uyghur semantic feature map; perform cross-channel fusion of the multi-channel Uyghur semantic feature map by improving the GS-Bottleneck module in the neck network to enhance the semantic association between different channels and obtain a Uyghur semantic fusion feature map; exchange local feature information of different channels in the Uyghur semantic fusion feature map through the GSConv module to obtain a Uyghur semantic enhancement feature map; detect the Uyghur semantic enhancement feature map through the head network to obtain the Uyghur position in the Uyghur image to be detected.

[0112] For the specific definition of the Uyghur detection device in natural scenes based on image processing, please refer to the definition of the Uyghur detection method in natural scenes based on image processing above, which will not be repeated here. Each module in the above-mentioned Uyghur detection device in natural scenes based on image processing can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory in the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0113] The present invention also provides a computer-readable storage medium, which stores a computer program, which can be used to execute the above Figure 1 The proposed method for Uyghur detection in natural scenes based on image processing.

[0114] The present invention also provides a computer device structure. At the hardware level, the computer device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include other hardware required for the business. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to achieve the above Figure 1 The proposed method for Uyghur detection in natural scenes based on image processing.

[0115] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided by the present invention can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory or optical memory, etc. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).

[0116] The technical features of the above embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of the present invention.

Claims

1. A method for detecting Uyghur characters in natural scenes based on image processing, characterized in that: include: The C2f module of the backbone network in the original YOLOv8 network is replaced with the C2f-EMSC module to form an improved backbone network; the C2f module of the neck network in the original YOLOv8 network is replaced with the GS-Bottleneck module, and the Conv module connected to the output end of the C2f module is replaced with the GSConv module to form an improved neck network; an improved YOLOv8 network including the improved backbone network, the improved neck network and the head network in the original YOLOv8 network is constructed; Collect Uyghur images to build a dataset, use the dataset to train the improved YOLOv8 network, and obtain a Uyghur detection model; The Uyghur image to be detected is input into the Uyghur detection model. The C2f-EMSC module in the backbone network is improved to divide the Uyghur image to be detected into different groups, and convolution operations of different scales are performed in parallel on each group to recognize the text in the Uyghur image to be detected at different scales, and a multi-channel Uyghur semantic feature map is obtained. The multi-channel Uyghur semantic feature map is cross-channel fused by improving the GS-Bottleneck module in the neck network to enhance the semantic association between different channels and obtain a Uyghur semantic fusion feature map. The local feature information of different channels in the Uyghur semantic fusion feature map is exchanged through the GSConv module to obtain a Uyghur semantic enhancement feature map. The Uyghur semantic enhancement feature map is detected by the head network to obtain the Uyghur position in the Uyghur image to be detected.

2. The method for detecting Uyghur characters in natural scenes based on image processing as claimed in claim 1, characterized in that: The method divides the Uyghur images to be detected into different groups by improving the C2f-EMSC module in the backbone network, and performs convolution operations of different scales on each group in parallel, so as to recognize the text in the Uyghur images to be detected at different scales and obtain a multi-channel Uyghur semantic feature map, which specifically includes: The feature map input to the C2f-EMSC module is divided into shallow features X_cheap and features to be processed X_group, where the size of the features to be processed X_group is bs×(g×ch)×h×w, where bs represents the batch size, g represents the number of groups, ch represents the number of channels, h represents the height, and w represents the width; The size of the feature to be processed X_group is rearranged to bs×ch×h×w×g to obtain several groups; convolution kernels of different sizes are applied to each group to capture Uyghur features of different scales; The convolutional groups are rearranged to restore their sizes to obtain the restored features X'_group, and the restored features X'_group and the shallow features X_cheap are concatenated to obtain the concatenated feature map; the concatenated feature map is convolved through a 1x1 convolutional layer to obtain the output of the C2f-EMSC module.

3. The method for detecting Uyghur characters in natural scenes based on image processing as claimed in claim 1, characterized in that: The method of exchanging local feature information of different channels in the Uyghur semantic fusion feature map through the GSConv module to obtain the Uyghur semantic enhancement feature map specifically includes: Perform standard convolution operation on the Uyghur semantic fusion feature map through the standard convolution layer in the GSConv module; The output of the standard convolution layer is subjected to a depth-wise separable convolution operation through the depth-wise separable convolution layer in the GSConv module; The output of the standard convolutional layer is concatenated with the output of the depthwise separable convolutional layer through the concatenation layer in the GSConv module; The output of the concatenation layer is shuffled through the shuffle layer in the GSConv module to rearrange the channels in the concatenated feature map and output the Uyghur semantic enhanced feature map.

4. The method for detecting Uyghur characters in natural scenes based on image processing as claimed in claim 1, characterized in that: The replacing the C2f module of the backbone network in the original YOLOv8 network with the C2f-EMSC module is to replace the third C2f module and the fourth C2f module of the backbone network in the original YOLOv8 network with the C2f-EMSC module.

5. The method for detecting Uyghur characters in natural scenes based on image processing as claimed in claim 1, characterized in that: The replacing of the C2f module with the GS-Bottleneck module is to replace the C2f module located after each Concat layer in the neck network with the GS-Bottleneck module.

6. The method for detecting Uyghur characters in natural scenes based on image processing as claimed in claim 1, characterized in that: The original YOLOv8 network is a YOLOv8n network.