Character detection model and method based on improved YOLOv7

By improving the YOLOv7 network, using the MobilenetV3 backbone network and a specific module combination, the problem of lack of a fast character detection model in the prior art that is suitable for all scenarios and is implemented efficient character detection operations.

CN120126148APending Publication Date: 2025-06-10XIDIAN UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510284955.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

In the prior art, there is a shortage of a character detection model suitable for all scenarios and is difficult to meet the demands of character recognition tasks in different application scenarios.

Method used

By improving YOLOv7, MobilenetV3 is used as the backbone network, combining PW convolution channel expansion, DW convolution, PW convolution, SE module, and inverse residual structure to reduce the network calculation amount and improve the computing efficiency.

Benefits of technology

On the basis of maintaining accuracy, the calculation amount of character detection is significantly reduced, the computing efficiency is improved, and it is suitable for character detection tasks in various scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126148A_ABST
    Figure CN120126148A_ABST
Patent Text Reader

Abstract

The invention discloses a character detection model based on improved YOLOv7, and belongs to the technical field of optical character recognition, the character detection model uses MobilenetV3 as a backbone network to extract features, a Multi-spliced ConcatBlock (MCB) module in a YOLOv7 original structure network is replaced, the MobilenetV3 comprises a Bcheck module, the Bcheck module comprises a PW convolution channel expansion module, a DW convolution module, a PW convolution module, an SE module and an inverse residual structure, and the inverse residual structure is used for extracting the features. The invention further discloses a character detection method based on the improved YOLOv7. By adopting the model and the method, the YOLOv7 is improved, and on the basis of not excessively losing precision, the calculation amount of a network is reduced, and the operation efficiency of character detection is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of optical character recognition, and in particular, to a character detection model and method based on improved YOLOv7. Background Art

[0002] For traditional optical character recognition (OCR) tasks, two neural networks are required. First, a character detection model is needed to detect and frame the areas in the image where characters exist. Second, the detection boxes to be detected are sent into the character recognition network, and the latter recognizes the characters in the areas to be detected. Most of the character detection models are implemented using object detection networks.

[0003] The YOLO series of networks are currently single-stage object detection networks with good performance and fast speed in the market, which can complete real-time detection. In addition, the accuracy of the classification and detection tasks of this series of models has a large tolerance space. If the YOLO series of networks are used to recognize the appearance characters of integrated circuits, stricter requirements are imposed on speed.

[0004] At present, the potential of neural networks has not been fully developed. For character recognition tasks, different application scenarios require different detection and recognition models, and there is no general character recognition algorithm that can be perfectly applicable to all scenarios. Therefore, there is an urgent need for a character detection model and method that is applicable to all scenarios and has high speed. Summary of the Invention

[0005] The purpose of the present invention is to provide a character detection model and method based on improved YOLOv7. By improving YOLOv7, without excessive loss of accuracy, the computational amount of the network is reduced, and the operation efficiency of character detection is improved.

[0006] To achieve the above purpose, the present invention provides a character detection model based on improved YOLOv7, which uses MobilenetV3 as the backbone network to extract features. The MobilenetV3 includes a Bneck module, and the Bneck module includes a PW convolutional channel expansion, a DW convolution, a PW convolution, an SE module, and an inverted residual structure.

[0007] Preferably, the PW convolutional channel expansion is used to expand the number of channels of the feature map; the DW convolution and the PW convolution are combined to implement the convolution operation through the separable convolution method, reducing the computational amount of character detection; the SE module is used to calculate each channel separately again; the inverted residual structure is used to prevent the loss of original feature information during the convolution operation.

[0008] Preferably, the Bneck module includes 15 layers. Among them, layers 1-3 and 7-10 include the PW convolution channel dilation, the DW convolution, and the PW convolution. Layers 4-6 and 11-15 include the PW convolution channel dilation, the DW convolution, the PW convolution, and the SE module. The SE module includes a pooling layer and two fully connected layers. The inverted residual structure only exists in the layers where the convolution stride is 1 and the number of channels does not change.

[0009] Preferably, the activation function of layers 1-6 of the Bneck module is Relu6, and the activation function of layers 7-15 is H-swish.

[0010] Preferably, the convolution kernel size of layers 1-3 and 7-12 of the Bneck module is 3*3, and the convolution kernel size of layers 4-6 and 13-15 is 5*5.

[0011] The present invention also provides a character detection method based on improved YOLOv7, and the steps include:

[0012] S1. Collect the appearance image of the integrated circuit, and preprocess the image to obtain the image to be detected;

[0013] S2. Use the improved YOLOv7 network to infer the collected image to obtain the character boxes in the areas where characters appear in the image;

[0014] S3. Use the CRNN network to recognize the characters in the obtained character boxes, and recognize the characters in the appearance image of the integrated circuit.

[0015] Preferably, the preprocessing of the image in step S1 includes image denoising, image enhancement, and image reconstruction, and the size after reconstruction is 640*640.

[0016] Preferably, the loss function of the improved YOLOv7 network in step S2 is set to CIOU_loss, and the parameters are trained by backpropagation and gradient descent to obtain the optimal parameters for feature extraction.

[0017] Preferably, step S2 specifically includes:

[0018] Input the image to be detected into the trained improved YOLOv7 network. First, perform convolution calculation on the image using the convolution layer, and then extract features through the Bneck module to obtain three feature maps with different scales;

[0019] Further extract features from the three feature maps with different scales through convolution operations and perform feature fusion;

[0020] Use different detection heads to detect the features after feature fusion to obtain the detection character boxes of feature maps with different scales.

[0021] Therefore, the present invention adopts the above-mentioned character detection model and method based on improved YOLOv7. Firstly, separable convolution is used, which greatly reduces the computational complexity of character detection. Subsequently, the SE module and the inverted residual structure are adopted to make up for the insufficient feature extraction of separable convolution, ensuring that the network will not excessively lose accuracy in the task of character text detection and positioning. Finally, through different activation functions, the nonlinear characteristics of features are increased, thereby increasing the expression ability of the model.

[0022] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Description of the Drawings

[0023] Figure 1 It is the improved YOLOv7 network architecture diagram of the embodiment of the present invention;

[0024] Figure 2 It is the data flow diagram of the Bneck module of the embodiment of the present invention. Detailed Embodiments

[0025] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components of the embodiments of the present invention described and illustrated herein can be arranged and designed in various different configurations. In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "upper", "lower", "inner", "outer", etc. is based on the orientation or positional relationship shown in the accompanying drawings, or the orientation or positional relationship when the product of the present invention is usually placed. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation to the present invention.

[0026] Embodiment

[0027] Referring to Figure 1-2 , the present invention provides a character detection model based on improved YOLOv7, which uses MobilenetV3 to replace the Multi-concatenation concatBlock (MCB) module in the original structure network for feature extraction, and the subsequent part of the network is the same as the original network. MobilenetV3 includes a Bneck module, and the Bneck module includes PW convolutional channel expansion, DW convolution, PW convolution, an SE module, and an inverted residual structure.

[0028] PW convolutional channel expansion is used to expand the number of channels of the feature map, specifically by increasing the number of convolution kernels.

[0029] The DW convolution and the PW convolution are combined to implement the convolution operation through the separable convolution method, reducing the computational complexity of character detection. The computational complexity of the ordinary convolution and the separable convolution is as follows:

[0030] For a feature map with a product of length and width of D and a number of channels of M, first, a convolutional layer with a kernel size of 3*3 is used for convolution operation to obtain a feature map with an output of N channels. The computational complexity required for the ordinary convolution is:

[0031] FLOPs_common = D * 9 * M * N.

[0032] For the same situation, the separable convolution first calculates each channel with a different convolution kernel (DW convolution) to obtain a convolutional layer with the same number of channels as the input. Next, a convolution kernel of size 1*1 is used to change the number of channels of the feature map (PW convolution). The computational complexity required for the separable convolution operation is:

[0033] FLOPs_division = D * M * 9 + D * M * N.

[0034] The ratio of the computational complexity is:

[0035] FLOPs_common / FLOPs_division = 9.

[0036] That is, compared with the ordinary convolution, the separable convolution reduces the parameters to 1 / 9 of the original.

[0037] In MobilenetV3, in order to alleviate the decrease in accuracy caused by the reduction in computational complexity during the parameter extraction process of the separable convolution, the SE attention mechanism is additionally introduced into the convolutional layer. The SE module is used to ensure that under the operation of the separable convolution, in order to avoid the feature dispersion caused by the small computational complexity, each channel is calculated separately again, so as to ensure that the network will not overly lose accuracy in the task of character text detection and localization.

[0038] In the SE module, pooling is performed on each channel of the obtained feature map. After pooling, the output vector is obtained through two fully connected layers. For the first fully connected layer, the number of nodes in the fully connected layer is equal to 1 / 4 of the number of channels of the input feature matrix. The number of channels of the second fully connected layer is the same as the number of channels of the feature matrix. After average pooling + two fully connected layers, the output feature vector can be understood as analyzing a weight relationship for each channel of the feature matrix before the SE module. It assigns a larger weight to the channels it considers more important, and a relatively smaller weight to the channels with less important dimensions.

[0039] The inverted residual structure is used to prevent the loss of original feature information during the convolution operation, and obtain the features that combine the features obtained through convolution calculation with the original features. The inverted residual structure only exists in the layers where the convolution stride is 1 and the number of channels does not change.

[0040] As shown in Table 1, the Bneck module includes 15 layers, which are the 2nd - 16th layers in the YOLOv7 network. Among them, the 2nd - 4th layers and the 8th - 14th layers in the YOLOv7 network include PW convolution channel dilation, DW convolution, and PW convolution. The 5th - 7th layers and the 12th - 16th layers include PW convolution channel dilation, DW convolution, PW convolution, and the SE module. The SE module includes a pooling layer and two fully connected layers.

[0041] It should be noted that although it is named Bneck3*3 in the figure, the convolution kernel sizes of the 5th - 7th layers and the 14th - 16th layers are 5*5. In the table, the activation function type 0 is Relu6, 1 is H-swish, and whether to use the SE module 0 is not used, 1 is used.

[0042] Table 1 Specific parameters of the Bneck module as the 2nd - 16th layers in the YOLOv7 network

[0043]

[0044]

[0045] The present invention also provides a character detection method based on the improved YOLOv7, and the steps include:

[0046] S1. Image acquisition: Use an industrial camera to take pictures of the appearance of the integrated circuit. First, denoise the captured image by using Gaussian filtering, then enhance the image by using histogram equalization, and then resize the image to a size of 640*640 to obtain the image to be detected.

[0047] S2. Use the improved YOLOv7 network to perform inference on the collected image to obtain the character frames in the areas where characters appear in the image, specifically including:

[0048] First, set the loss function of the improved YOLOv7 network to CIOU_loss, and then train the parameters by the way of backpropagation and gradient descent to obtain the optimal parameters for feature extraction.

[0049] Input the image to be detected into the trained improved YOLOv7 network. First, use a convolutional layer with a kernel size of 3*3 and a stride of 2 to perform convolutional calculations on the image, obtaining a feature map with a size of 320*320*16. Next, through the 2-3 layers of Bneck3*3 in the improved YOLOv7 network, after separable convolutional calculations, a feature map with a size of 160*160*24 is obtained. Then, this feature map undergoes convolutional calculations through the 4-7 layers, resulting in a feature Figure 1 , this feature Figure 1 While saving and sending it to the Neck part of the network, continue with convolutional calculations through the 8-11 layers, obtaining a feature map with a size of 40*40*80. Then, through convolutional calculations in the 12-13 layers, a feature with a size of 40*40*112 is obtained Figure 2 , feature Figure 2 While saving and sending it to the Neck part of the network, through convolutional calculations in the 14-16 layers, a feature map with a size of 20*20*160 is obtained and enters the SPPCSPC module. The SPPCSPC module adopts an operation that combines multi-scale convolution and Cat calculation, extracts feature information of different scales, and obtains feature map 5, which is also sent to the Neck part of the network. The Neck part can be simply divided into an FPN network and a PAN network. Among them, the FPN network performs top-down feature fusion of the pyramid, while the PAN network performs bottom-up feature fusion.

[0050] The Neck part first uses the FPN network for feature fusion. Specifically: all three feature maps further extract features through convolutional operations. First, the smallest feature map 5 undergoes an upsampling operation and is concatenated with the feature Figure 2 after convolution, and then through the MCB module, convolutional and concatenation calculations for multi-scale feature fusion are performed to obtain feature map 4. After performing convolution and upsampling operations on feature map 4, it is concatenated with the feature Figure 1 after convolution, obtaining feature map 3.

[0051] Then, the PAN network is used for feature fusion. Specifically: feature map 3 undergoes downsampling and is concatenated with feature map 4 to obtain P4, and then through downsampling and concatenation with feature map 5, P5 is obtained.

[0052] The feature maps 3, P4, and P5 after feature fusion pass through the MCB module and the CBL module, and are respectively sent to the yolo_head_P3, yolo_head_P4, and yolo_head_P5 detection heads for detection, thus realizing object detection at three scales: large, medium, and small. YOLO_head divides the image into N*N grids, predicts several instance boxes in each grid, and fine-tunes the center coordinate position and size of the prediction box.

[0053] S3. Use the CRNN network to perform character recognition on the obtained character boxes to recognize the characters in the integrated circuit appearance image.

[0054] Therefore, the present invention adopts the above-mentioned character detection model and method based on the improved YOLOv7. By improving YOLOv7, on the basis of not overly losing accuracy, the computational amount of the network is reduced, and the operation efficiency of character detection is improved.

[0055] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify or equivalently replace the technical solutions of the present invention, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A character detection model based on improved YOLOv7, characterized by: MobilenetV3 is used as the backbone network to extract features, and the MobilenetV3 includes a Bneck module, and the Bneck module includes a PW convolution channel expansion, a DW convolution, a PW convolution, a SE module and an inverted residual structure.

2. A character detection model based on improved YOLOv7 according to claim 1, characterized in that: The PW convolution channel expansion is used to expand the number of channels of the feature map; the DW convolution is combined with the PW convolution to implement the convolution operation through the separation convolution method, thereby reducing the amount of character detection calculation; the SE module is used to calculate each channel again separately; the inverted residual structure is used to prevent the original feature information from being lost in the convolution operation.

3. A character detection model based on improved YOLOv7 according to claim 1, characterized in that: The Bneck module includes 15 layers, among which layers 1-3 and layers 7-10 include the PW convolution channel expansion, the DW convolution and the PW convolution, layers 4-6 and layers 11-15 include the PW convolution channel expansion, the DW convolution, the PW convolution and the SE module, the SE module includes a pooling layer and two fully connected layers, and the inverted residual structure only exists in the layer where the convolution step size is 1 and the number of channels does not change.

4. A character detection model based on improved YOLOv7 according to claim 3, characterized in that: The activation function of the Bneck module 1-6 layers is Relu6, and the activation function of the 7-15 layers is H-swish.

5. A character detection model based on improved YOLOv7 according to claim 1, characterized in that: The convolution kernel size of the Bneck module 1-3 layers and 7-12 layers is 3*3, and the convolution kernel size of the 4-6 layers and 13-15 layers is 5*5.

6. A character detection method based on improved YOLOv7, applying a character detection model based on improved YOLOv7 according to any one of claims 1 to 5, characterized in that the steps include: S1, collecting the appearance image of the integrated circuit and preprocessing the image to obtain the image to be detected; S2. Use the improved YOLOv7 network to infer the collected image and obtain the character box of the area where the character appears in the image; S3. Use the CRNN network to perform character recognition on the obtained character box to identify the characters in the integrated circuit appearance image.

7. A character detection method based on improved YOLOv7 according to claim 6, characterized in that: In step S1, the image is preprocessed including image denoising, image enhancement and image reconstruction, and the reconstructed size is 640*640.

8. A character detection method based on improved YOLOv7 according to claim 6, characterized in that: The improved YOLOv7 network loss function in step S2 is set to CIOU_loss, and the parameters are trained by back propagation and gradient descent to obtain the optimal parameters for feature extraction.

9. A character detection method based on improved YOLOv7 according to claim 8, characterized in that: Step S2 specifically includes: The image to be detected is input into the trained improved YOLOv7 network. The convolution layer is used to perform convolution calculation on the image, and then the Bneck module is used to extract features to obtain three feature maps of different scales. The three feature maps of different scales are further extracted through convolution operations and feature fusion is performed; Use different detection heads to detect the features after feature fusion and obtain the detection character boxes of feature maps of different scales.