A method for establishing a text detection model, a text detection method and a device
By introducing cascading feature extraction units and attention subunits into the feature extraction module of the text detection model, the problem of low accuracy of text detection models in the prior art is solved, and higher text detection accuracy and model reliability are achieved.
Patent Information
- Application Number
- CN202210371955.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-11
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-04-11
AI Technical Summary
In the prior art, the feature extraction of the text detection model is insufficient, resulting in the low accuracy of the text detection model trained.
A method for establishing a text detection model is proposed. By obtaining text detection training data, and based on these data and the original model, the text detection model is obtained. The original model includes a feature extraction module, which consists of a plurality of cascading feature extraction units, each feature extraction unit including convolutional pooling subunits and attention subunits connected in turn. By setting attention subunits in each feature extraction unit of the feature extraction module for feature extraction, a more informative feature map is obtained.
By obtaining richer feature maps, the accuracy of file detection is improved, thereby improving the reliability of text detection models.
Smart Images

Figure CN114639095B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and particularly relates to a method for establishing a text detection model, a text detection method and a device. Background Art
[0002] Text detection is a branch in the field of image processing. With the development of technology, text detection has been more and more widely applied, such as license plate recognition, invoice recognition and picture text recognition, etc.
[0003] In the prior art, text detection is based on manually labeled text images as the training data of the model, and then feature extraction is performed on the training data, and the model is trained based on the extracted feature data to obtain a text detection model. When extracting features, due to insufficient feature extraction, the accuracy of the trained text detection model is not high. Summary of the Invention
[0004] In view of the problems in the prior art, embodiments of the present invention provide a method for establishing a text detection model, a text detection method and a device, which can at least partially solve the problems existing in the prior art.
[0005] In a first aspect, the present invention proposes a method for establishing a text detection model, including:
[0006] Obtaining text detection training data;
[0007] Based on the text detection training data and an original model, training to obtain a text detection model; wherein, the original model includes a feature extraction module, the feature extraction module includes a plurality of cascaded feature extraction units, and each feature extraction unit includes a convolutional pooling subunit and an attention subunit connected in sequence.
[0008] In a second aspect, the present invention provides a text detection method based on the method for establishing a text detection model according to any one of the above, including:
[0009] Obtaining a text image to be detected;
[0010] Based on the text image to be detected and the text detection model, obtaining a detection result corresponding to the text image to be detected.
[0011] In a third aspect, the present invention proposes a device for establishing a text detection model, including:
[0012] A first obtaining module, configured to obtain text detection training data;
[0013] A training module, configured to train and obtain a text detection model based on the text detection training data and the original model; wherein, the original model includes a feature extraction module, and the feature extraction module includes a plurality of cascaded feature extraction units, and each feature extraction unit includes a convolutional pooling subunit and an attention subunit connected in sequence.
[0014] In a third aspect, the present invention provides a text detection device based on the above-mentioned text detection model establishment device, including:
[0015] A second acquisition module, configured to acquire a text image to be detected;
[0016] A detection module, configured to obtain a detection result corresponding to the text image to be detected based on the text image to be detected and the text detection model.
[0017] In yet another aspect, the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the text detection model establishment method or the text detection method described in any one of the above embodiments.
[0018] In still another aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the text detection model establishment method or the text detection method described in any one of the above embodiments.
[0019] The text detection model establishment method, the text detection method and the device provided by the embodiments of the present invention can acquire text detection training data, train and obtain a text detection model based on the text detection training data and the original model. The original model includes a feature extraction module, and the feature extraction module includes a plurality of cascaded feature extraction units. Each feature extraction unit includes a convolutional pooling subunit and an attention subunit connected in sequence. By setting an attention subunit in each feature extraction unit of the feature extraction module for feature extraction, a feature map with richer information can be obtained, the accuracy of file detection can be improved, and thus the reliability of the text detection model is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings. In the drawings:
[0021] Figure 1 is a schematic flowchart of the text detection model establishment method provided by the first embodiment of the present invention.
[0022] Figure 2 It is a schematic flowchart of a method for establishing a text detection model provided by the second embodiment of the present invention.
[0023] Figure 3 It is a schematic structural diagram of an attention sub-unit provided by the third embodiment of the present invention.
[0024] Figure 4 It is a schematic structural diagram of an attention sub-unit provided by the fourth embodiment of the present invention.
[0025] Figure 5 It is a schematic structural diagram of a feature extraction module provided by the fifth embodiment of the present invention.
[0026] Figure 6 It is a schematic flowchart of a method for establishing a text detection model provided by the sixth embodiment of the present invention.
[0027] Figure 7 It is a schematic structural diagram of a device for establishing a text detection model provided by the seventh embodiment of the present invention.
[0028] Figure 8 It is a schematic structural diagram of a device for establishing a text detection model provided by the eighth embodiment of the present invention.
[0029] Figure 9 It is a schematic structural diagram of a text detection device provided by the ninth embodiment of the present invention.
[0030] Figure 10 It is a schematic physical structure diagram of an electronic device provided by the tenth embodiment of the present invention. Detailed implementation manners
[0031] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer and more understandable, the following further elaborates on the embodiments of the present invention with reference to the accompanying drawings. Here, the illustrative embodiments of the present invention and their descriptions are used to explain the present invention, but not to limit the present invention. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined arbitrarily with each other.
[0032] To facilitate understanding of the technical solutions provided in this application, the following first explains the relevant content of the technical solutions in this application. In reality, many applications related to text recognition have greatly facilitated our lives, such as license plate recognition, invoice recognition, picture-to-text recognition, etc. The method for establishing a text detection model provided by the embodiments of the present invention improves the existing text detection algorithm based on CTPN (Connectionist Text Proposal Network) to establish a new text detection model to improve the accuracy of text detection.
[0033] Taking the server as the execution subject as an example, the following describes the establishment method of the text detection model provided by the embodiments of the present invention and the specific implementation process of the text detection method.
[0034] Figure 1 It is a schematic flowchart of the establishment method of the text detection model provided by the first embodiment of the present invention. As Figure 1 shown, the establishment method of the text detection model provided by the embodiments of the present invention includes:
[0035] S101. Obtain text detection training data;
[0036] Specifically, images including text can be collected, and the text is used as the text label of the image to obtain text detection training data. The server can obtain the text detection training data, and the text detection training data includes multiple text detection training images, and each text detection training image has a corresponding text label.
[0037] For example, in order to improve the diversity of the text detection training data, text detection training images with different font sizes can be collected, text detection training images with different font formats can be collected, and text detection training images with different shooting angles can be collected.
[0038] S102. Based on the text detection training data and the original model, train to obtain a text detection model; wherein, the original model includes a feature extraction module, and the feature extraction module includes a plurality of cascaded feature extraction units, and each feature extraction unit includes a convolutional pooling subunit and an attention subunit connected in sequence.
[0039] Specifically, the server inputs the text detection training data into the original model, and a text detection model can be trained. Among them, the original model includes a feature extraction module and a detection classification module. The feature extraction module is used to extract features from the text detection training images to obtain an enhanced feature map, and the detection classification module is used to detect text in the text area of the enhanced feature map to obtain detected text. The detection classification module can be implemented by a convolutional layer, such as using 1 or 2 convolutional layers. Among them, the specific structure of the detection classification module is set according to actual needs, and the embodiments of the present invention do not make limitations.
[0040] The feature extraction module includes a plurality of cascaded feature extraction units, and each feature extraction unit includes a convolutional pooling subunit and an attention subunit connected in sequence. Each convolutional pooling subunit includes a convolutional layer and a pooling layer. The convolutional layer is used for feature extraction. The convolutional layer can use a convolutional kernel of size 3×3, and the Relu activation function can be used after convolution. The pooling layer is used to reduce the dimension of the extracted feature information to reduce the number of parameters and the amount of computation. The attention subunit performs weighted processing on the input feature image, reduces the feature loss of the image, and enables the image output by the attention subunit to have richer features, which is beneficial to improving the accuracy of subsequent document detection. Among them, the number of feature extraction units is selected according to actual needs, and the embodiments of the present invention do not make any limitations.
[0041] It can be understood that the original model will set initial parameters. During the training process of the model, the parameters of the model will be continuously updated until the training is completed to obtain the text detection model.
[0042] The method for establishing a text detection model provided by the embodiments of the present invention can obtain text detection training data, and based on the text detection training data and the original model, train to obtain a text detection model. The original model includes a feature extraction module. The feature extraction module includes a plurality of cascaded feature extraction units. Each feature extraction unit includes a convolutional pooling subunit and an attention subunit connected in sequence. By setting an attention subunit in each feature extraction unit of the feature extraction module for feature extraction, a feature map with richer information can be obtained, the accuracy of document detection can be improved, and thus the reliability of the text detection model can be improved.
[0043] Figure 2 It is a schematic flowchart of the method for establishing a text detection model provided by the second embodiment of the present invention. As Figure 2 shown, on the basis of the above embodiments, further, the training to obtain a text detection model based on the text detection training data and the original model includes:
[0044] S201. Perform feature extraction on the input image through the convolutional pooling subunit included in each feature extraction unit to obtain a feature extraction map;
[0045] Specifically, each feature extraction unit of the feature extraction module includes a convolutional pooling subunit to perform feature extraction on the input image, obtaining a feature extraction map corresponding to the convolutional pooling subunit. For the first feature extraction unit of the feature extraction module, the input image of the convolutional pooling subunit is the text detection training image included in the text detection training data. For the second feature extraction unit of the feature extraction module, the input image of the convolutional pooling subunit is the enhanced feature map output by the first feature extraction unit. For the third feature extraction unit of the feature extraction module, the input image of the convolutional pooling subunit is the enhanced feature map output by the second feature extraction unit, and so on.
[0046] S202. Perform feature enhancement processing on the feature extraction map through the attention subunit included in each feature extraction unit, obtaining an enhanced feature map.
[0047] Specifically, the feature extraction map output by the convolutional pooling subunit of each feature extraction unit serves as the input of the attention subunit included in each feature extraction unit. Through the attention subunit included in each feature extraction unit, feature enhancement processing will be performed on the input feature extraction map, and the enhanced feature map corresponding to the attention subunit of each feature extraction unit will be input. For the first feature extraction unit of the feature extraction module, the input image of the attention subunit is the feature extraction map output by the convolutional pooling subunit of the first feature extraction unit. For the second feature extraction unit of the feature extraction module, the input image of the attention subunit is the feature extraction map output by the convolutional pooling subunit of the second feature extraction unit. For the third feature extraction unit of the feature extraction module, the input image of the attention subunit is the feature extraction map output by the convolutional pooling subunit of the second feature extraction unit, and so on.
[0048] Performing feature enhancement processing on the image through the attention subunit included in each feature extraction unit can enhance the features of the feature extraction map and enrich the features of the feature extraction map.
[0049] Based on the above embodiments, further, each attention subunit includes a weight extraction channel and a temporal feature extraction channel, where:
[0050] The weight extraction channel includes a first convolutional layer, a first image reconstruction layer, a second convolutional layer, a second image reconstruction layer, a normalization layer, a third convolutional layer, a first normalization layer, a first activation layer, and a fourth convolutional layer; the temporal feature extraction channel includes a temporal extraction layer.
[0051] The input ends of the first convolutional layer, the second convolutional layer, and the timing extraction layer are respectively connected to the output ends of the corresponding convolutional pooling sub-units; the output end of the first convolutional layer is connected to the input end of the first image reconstruction layer, the output end of the second convolutional layer is connected to the input end of the second image reconstruction layer, the output end of the second image reconstruction layer is connected to the input end of the normalization layer, the cross product result of the output result of the first image reconstruction layer and the output result of the normalization layer is used as the input of the third convolutional layer, the output end of the third convolutional layer is connected to the input end of the first normalization layer, the output end of the first normalization layer is connected to the input end of the first activation layer, and the dot product result of the output result of the first activation layer and the output result of the timing extraction layer is used as the input of the fourth convolutional layer.
[0052] Specifically, the weight extraction channel is used to extract a weight parameter from the input feature extraction map, the timing feature extraction channel is used to extract timing features from the input feature extraction map, the output result of the weight extraction channel and the output result of the timing feature extraction channel are multiplied by dot product, and the result of the dot product is used as the input of the fourth convolutional layer, and the output of the fourth convolutional layer is an enhanced feature map.
[0053] The first convolutional layer, the second convolutional layer, the third convolutional layer, and the fourth convolutional layer can use a convolutional kernel with a size of 1×1. The first image reconstruction layer and the second image reconstruction layer are used to reconstruct the shape of the image and can use the Reshape function. The normalization layer can use the Softmax function to normalize the input image, so that the pixel values of the image are converted between 0 and 1 for subsequent processing. The first normalization layer is used to normalize the input image and can be implemented by the method of Layer Normalization. The first activation layer can use the Sigmod function to perform an activation operation on the input image. The timing extraction layer can use Long Short-Term Memory (LSTM for short) or bidirectional LSTM.
[0054] The feature extraction maps output by the convolutional pooling subunit are respectively input into the first convolutional layer, the second convolutional layer, and the temporal extraction layer. After the first convolutional layer extracts features from the feature extraction map, it is input into the first image reconstruction layer for image shape reconstruction. After the second convolutional layer extracts features from the feature extraction map, it is input into the second image reconstruction layer for image shape reconstruction. The output image of the second image reconstruction layer is then normalized through the normalization layer. The output result of the first image reconstruction layer and the output result of the normalization layer are multiplied crosswise to obtain a crosswise multiplication result. The crosswise multiplication result is used as the input of the third convolutional layer. After the third convolutional layer extracts features from the crosswise multiplication result, the output result of the third convolutional layer is obtained. The output result of the third convolutional layer undergoes the normalization process of the first normalization layer and the activation operation of the first activation layer in sequence to obtain the output result of the first activation layer. The output result of the first activation layer includes weight parameters. The temporal extraction layer extracts temporal features from the feature extraction map to obtain the output result of the temporal extraction layer. The output result of the first activation layer and the output result of the temporal extraction layer are multiplied dotwise to obtain a dotwise multiplication result, and the dotwise multiplication result is used as the input of the output convolutional layer.
[0055] For example, Figure 3 is a schematic structural diagram of the attention subunit provided in the third embodiment of the present invention. As Figure 3 shown, from top to bottom on the left side of the figure are the first convolutional layer, the first image reconstruction layer, the third convolutional layer, the first normalization layer, and the first activation layer. From top to bottom in the middle of the figure are the second convolutional layer, the second image reconstruction layer, the normalization layer, and the fourth convolutional layer. On the right side of the figure is the temporal extraction layer. ⊙ represents dotwise multiplication, and
[0056] The feature extraction maps of C×H×W are respectively input into the first convolutional layer, the second convolutional layer, and the temporal extraction layer. The convolutional kernels of the first convolutional layer and the second convolutional layer are of size 1×1, and the LSTM is adopted in the temporal extraction layer. The first convolutional layer performs feature extraction on the input feature extraction map and outputs the first feature map of C×H×W. The first image reconstruction layer reconstructs the input first feature map of C×H×W into a feature map of C / 2×HW. The second convolutional layer performs feature extraction on the input feature extraction map and outputs a feature map of 1×H×W. The second image reconstruction layer reconstructs the input image of 1×H×W into a feature map of HW×1×1. The normalization layer processes the feature map of HW×1×1 through the Softmax function and outputs a feature map of HW×1×1. The feature map of C / 2×HW is multiplied element-wise with the feature map of HW×1×1 output by the normalization layer to obtain a feature map of C / 2×1×1. The feature map of C / 2×1×1 successively undergoes feature extraction by the third convolutional layer, normalization processing by the first normalization layer, and activation operation by the first activation layer to obtain a feature map of C×1×1. The temporal extraction layer performs temporal feature extraction on the feature extraction map of C×H×W and outputs the second feature map of C×H×W. The feature map of C×1×1 output by the first activation layer is multiplied element-wise with the second feature map of C×H×W output by the temporal extraction layer to obtain the third feature map of C×H×W. The fourth convolutional layer performs feature extraction on the third feature map of C×H×W and outputs an enhanced feature map of C×H×W. The convolutional kernels of the third convolutional layer and the fourth convolutional layer are of size 1×1. Among them, C represents the channels of the image, H represents the height of the image, and W represents the width of the image.
[0057] Based on the above embodiments, further, each attention sub-unit includes a weight extraction channel and a temporal feature extraction channel, where:
[0058] The weight extraction channel includes a fifth convolutional layer, a global pooling layer, a sixth convolutional layer, a second normalization layer, and a second activation layer connected in sequence. The temporal feature extraction channel includes a temporal extraction layer;
[0059] The input ends of the fifth convolutional layer and the temporal extraction layer are connected to the output end of the corresponding convolutional pooling sub-unit; the element-wise multiplication result of the output result of the weight extraction channel and the output result of the temporal feature extraction channel is used as the output result of the attention sub-unit.
[0060] Specifically, the weight extraction channel is used to extract a weight parameter from the input feature extraction map, the temporal feature extraction channel is used to extract temporal features from the input feature extraction map, the output result of the weight extraction channel is multiplied element-wise with the output result of the temporal feature extraction channel, and the multiplication result is used as the output result of the attention sub-unit. The output result of the attention sub-unit is an enhanced feature map.
[0061] The fifth convolutional layer and the sixth convolutional layer can use convolutional kernels of size 1×1. The global pooling layer is used to adjust the input feature map of H×W to a feature map of 1×1. The second normalization layer is used to normalize the input image, which can be implemented by the method of Layer Normalization. The second activation layer can use the ReLU function and is used to activate the input image. The temporal extraction layer can use Long Short-Term Memory (LSTM) or bidirectional LSTM.
[0062] The feature extraction maps output by the convolutional pooling subunit are respectively input into the fifth convolutional layer and the temporal extraction layer. After the fifth convolutional layer extracts features from the feature extraction map, it is input into the global pooling layer for image size adjustment. The output image of the global pooling layer successively undergoes feature extraction by the sixth convolutional layer, normalization by the second normalization layer, and activation by the second activation layer to obtain the output result of the weight extraction channel, and the output result of the weight extraction channel includes weight parameters. The temporal extraction layer extracts temporal features from the feature extraction map to obtain the output result of the temporal extraction layer, which is also the output result of the temporal feature extraction channel. The output result of the weight extraction channel is multiplied pointwise with the output result of the temporal feature extraction channel to obtain a pointwise multiplication result, and the pointwise multiplication result is used as the output of the attention subunit.
[0063] For example, Figure 4 is a schematic structural diagram of the attention subunit provided in the fourth embodiment of the present invention. As Figure 4 shown, from top to bottom on the left side of the figure, there are successively the fifth convolutional layer, the global pooling layer, the sixth convolutional layer, the second normalization layer, and the second activation layer. On the right side of the figure is the temporal extraction layer, and ⊙ represents pointwise multiplication.
[0064] The feature extraction maps of C×H×W are respectively input into the fifth convolutional layer and the temporal extraction layer. The convolutional kernels of the fifth convolutional layer and the sixth convolutional layer are of size 1×1, and the temporal extraction layer uses LSTM. The fifth convolutional layer performs feature extraction on the input feature extraction map and outputs the first feature map of C / 2×H×W. The first feature map of C×H×W is processed by the global pooling layer and outputs a feature map of C / 2×1×1. The feature map of C / 2×1×1 undergoes feature extraction by the sixth convolutional layer, normalization processing by the second normalization layer, and activation operation by the second activation layer to obtain a feature map of C×1×1. The temporal extraction layer performs temporal feature extraction on the feature extraction map of C×H×W and outputs the second feature map of C×H×W. The feature map of C×1×1 is multiplied element-wise with the second feature map of C×H×W output by the temporal extraction layer, and the obtained third feature map of C×H×W is used as the output result of the attention sub-unit. Here, C represents the number of channels of the image, H represents the height of the image, and W represents the width of the image.
[0065] Based on the above embodiments, further, the feature extraction module includes 4 - 8 cascaded feature extraction units. For example, it is set that the feature extraction module includes 5 cascaded feature extraction units.
[0066] For example, Figure 5 is a schematic structural diagram of the feature extraction module provided by the fifth embodiment of the present invention. As Figure 5 shown, the feature extraction module includes 5 cascaded feature extraction units. Each feature extraction unit includes a convolutional pooling sub-unit and an attention sub-unit connected in sequence. Each convolutional pooling sub-unit includes a convolutional layer and a pooling layer. The convolutional kernel of the convolutional layer is of size 3×3, the pooling layer uses max pooling with a window size of 2×2, and the ReLU activation function is used between the convolutional layer and the pooling layer.
[0067] Based on the above embodiments, further, training the text detection model based on the text detection training data and the original model includes:
[0068] Using an optimizer to accelerate the training process of the text detection model.
[0069] Specifically, an optimizer is used during the training process of the text detection model. The training process of the neural network is to minimize the loss function. The role of the optimizer is to update and calculate the network parameters that affect model training and model output, so that the loss function reaches the minimum value. By using an optimizer, the convergence of the network can be accelerated, the training time can be reduced, and thus the training process can be accelerated. The optimizer is selected according to actual needs, and the embodiments of the present invention do not make limitations.
[0070] For example, the optimizer can adopt the Adam (Adaptive Moment Estimation) optimizer, that is, an optimizer based on the gradient descent algorithm.
[0071] Based on the above embodiments, further, the text detection training data includes text images with different font sizes, different font formats, and different shooting angles.
[0072] Specifically, in order to improve the diversity of the text detection training data, text images with different font sizes, text images with different font formats, and text images with different shooting angles can be collected as the text detection training data.
[0073] Figure 6 It is a schematic flowchart of the method for establishing a text detection model provided in the fifth embodiment of the present invention. As Figure 6 shown, the text detection method based on the method for establishing a text detection model according to any of the above embodiments provided in the embodiments of the present invention includes:
[0074] S601. Obtain a text image to be detected;
[0075] Specifically, the server can obtain a text image to be detected, and the text image to be detected is an image including text, and it is necessary to detect what text the text image to be detected includes.
[0076] For example, the text image to be detected can be a license plate image, an invoice image, etc.
[0077] S602. Obtain a detection result corresponding to the text image to be detected based on the text image to be detected and the text detection model.
[0078] Specifically, the server inputs the text image to be detected into the text detection model, and after being processed by the text detection model, outputs a detection result corresponding to the text image to be detected, and the detection result is the text corresponding to the text image to be detected. Among them, the text detection model is constructed based on the method for establishing a text detection model according to any of the above embodiments.
[0079] The text detection method provided by the embodiments of the present invention can obtain a text image to be detected, obtain a detection result corresponding to the text image to be detected based on the text image to be detected and the text detection model, and can accurately detect the text in the text image to be detected through the text detection model, improving the accuracy of text detection.
[0080] Next, taking the detection of the invoice number and invoice code of an invoice as an example, the specific implementation processes of the text detection method and the method for establishing a text detection model provided by the embodiments of the present invention are described.
[0081] Collect invoice images. The invoice images can be pictures obtained by photographing or scanning paper invoices, or pictures of electronic invoices. Mark the corresponding areas of the invoice number and invoice code from each invoice image, and then perform image extraction to obtain the invoice number area image and invoice code area image corresponding to each invoice image. The invoice number area image includes the invoice number, and the invoice code area image includes the invoice code. The invoice number corresponding to the invoice number area image is used as the text label corresponding to the invoice number area image, and the invoice number corresponding to the invoice number area image is used as the text label corresponding to the invoice number area image.
[0082] The invoice code area images and invoice number area images corresponding to the collected invoice images are used as text detection training images, and the text detection training images and corresponding text labels are used as text detection training data.
[0083] Construct an original model. The original model includes a feature extraction module. The feature extraction module includes multiple cascaded feature extraction units. Each feature extraction unit includes a convolutional pooling sub-unit and an attention sub-unit connected in sequence. Each attention sub-unit includes a weight extraction channel and a temporal feature extraction channel. The weight extraction channel includes a first convolutional layer, a first image reconstruction layer, a second convolutional layer, a second image reconstruction layer, a normalization layer, a third convolutional layer, a first normalization layer, a first activation layer, and a fourth convolutional layer; the temporal feature extraction channel includes a temporal extraction layer; the input ends of the first convolutional layer, the second convolutional layer, and the temporal extraction layer are respectively connected to the output end of the corresponding convolutional pooling sub-unit; the output end of the first convolutional layer is connected to the input end of the first image reconstruction layer, the output end of the second convolutional layer is connected to the input end of the second image reconstruction layer, the output end of the second image reconstruction layer is connected to the input end of the normalization layer, the cross product result of the output result of the first image reconstruction layer and the output result of the normalization layer is used as the input of the third convolutional layer, the output end of the third convolutional layer is connected to the input end of the first normalization layer, the output end of the first normalization layer is connected to the input end of the first activation layer, and the dot product result of the output result of the first activation layer and the output result of the temporal extraction layer is used as the input of the fourth convolutional layer.
[0084] Divide the text detection training data into a training set and a validation set. Perform model training according to the training set and the original model to obtain a text detection model to be verified, and then verify the text detection model to be verified according to the validation set. After the text detection model to be verified passes the verification, use the text detection model to be verified as the text detection model for invoices. Among them, the Adam optimizer is used during the training process.
[0085] Obtain the invoice picture to be detected, and extract the picture of the invoice number area to be recognized and the picture of the invoice code area from the invoice picture to be detected. Input the picture of the invoice number area to be recognized and the picture of the invoice code area into the text detection model of the invoice respectively, and the invoice code and invoice number corresponding to the invoice picture to be detected can be output.
[0086] During the invoice reimbursement process, it is necessary to register the invoice number and invoice code. Through the text detection model of the invoice, the invoice code and invoice number of the invoice can be detected, so that the invoice registration can be automatically carried out, improving the efficiency of invoice registration.
[0087] Figure 7 It is a schematic structural diagram of the device for establishing a text detection model provided by the seventh embodiment of the present invention. As Figure 7 shown, the device for establishing a text detection model provided by the embodiment of the present invention includes a first acquisition module 701 and a training module 702, wherein:
[0088] The first acquisition module 701 is used to acquire text detection training data; the training module 702 is used to train and obtain a text detection model based on the text detection training data and the original model; wherein, the original model includes a feature extraction module, and the feature extraction module includes a plurality of cascaded feature extraction units, and each feature extraction unit includes a convolutional pooling subunit and an attention subunit connected in sequence.
[0089] Specifically, images including text can be collected, and the text is used as the text label of the image to obtain text detection training data. The first acquisition module 701 can acquire the text detection training data, and the text detection training data includes multiple text detection training images, and each text detection training image has a corresponding text label.
[0090] The training module 702 inputs the text detection training data into the original model, and can train and obtain a text detection model. Wherein, the original model includes a feature extraction module and a detection classification module. The feature extraction module is used to extract features from the text detection training image to obtain an enhanced feature map, and the detection classification module is used to detect the text area in the enhanced feature map to obtain the detected text. The specific structure of the detection classification module is set according to actual needs, and is not limited in the embodiment of the present invention.
[0091] The apparatus for establishing a text detection model provided by an embodiment of the present invention can obtain text detection training data and train a text detection model based on the text detection training data and an original model. The original model includes a feature extraction module, and the feature extraction module includes a plurality of cascaded feature extraction units. Each feature extraction unit includes a convolutional pooling subunit and an attention subunit connected in sequence. By setting an attention subunit in each feature extraction unit of the feature extraction module for feature extraction, a feature map with richer information can be obtained, the accuracy of document detection can be improved, and thus the reliability of the text detection model can be improved.
[0092] Figure 8 FIG. is a schematic structural diagram of the apparatus for establishing a text detection model provided by the eighth embodiment of the present invention. As Figure 8 shown, on the basis of the above embodiments, further, the training module 702 includes an extraction unit 7021 and an enhancement unit 7022, where:
[0093] The extraction unit 7021 performs feature extraction on the input image through the convolutional pooling subunit included in each feature extraction unit to obtain a feature extraction map; the enhancement unit 7022 performs feature enhancement processing on the feature extraction map through the attention subunit included in each feature extraction unit to obtain an enhanced feature map.
[0094] On the basis of the above embodiments, further, each attention subunit includes a weight extraction channel and a temporal feature extraction channel, where:
[0095] The weight extraction channel includes a first convolutional layer, a first image reconstruction layer, a second convolutional layer, a second image reconstruction layer, a normalization layer, a third convolutional layer, a first normalization layer, a first activation layer, and a fourth convolutional layer; the temporal feature extraction channel includes a temporal extraction layer;
[0096] The input ends of the first convolutional layer, the second convolutional layer, and the temporal extraction layer are respectively connected to the output end of the corresponding convolutional pooling subunit; the output end of the first convolutional layer is connected to the input end of the first image reconstruction layer, the output end of the second convolutional layer is connected to the input end of the second image reconstruction layer, the output end of the second image reconstruction layer is connected to the input end of the normalization layer, the cross-product result of the output result of the first image reconstruction layer and the output result of the normalization layer is used as the input of the third convolutional layer, the output end of the third convolutional layer is connected to the input end of the first normalization layer, the output end of the first normalization layer is connected to the input end of the first activation layer, and the dot-product result of the output result of the first activation layer and the output result of the temporal extraction layer is used as the input of the fourth convolutional layer.
[0097] Based on the above embodiments, further, each attention sub-unit includes a weight extraction channel and a temporal feature extraction channel, where:
[0098] The weight extraction channel includes a fifth convolutional layer, a global pooling layer, a sixth convolutional layer, a second normalization layer, and a second activation layer connected in sequence. The temporal feature extraction channel includes a temporal extraction layer;
[0099] The input ends of the fifth convolutional layer and the temporal extraction layer are connected to the output end of the corresponding convolutional pooling sub-unit; the dot product result of the output result of the weight extraction channel and the output result of the temporal feature extraction channel is used as the output result of the attention sub-unit.
[0100] Based on the above embodiments, further, the feature extraction module includes 4-8 cascaded feature extraction units.
[0101] Based on the above embodiments, further, the training module 702 is specifically configured to:
[0102] Accelerate the training process of the text detection model through an optimizer.
[0103] Based on the above embodiments, further, the text detection training data includes text images with different font sizes, different font formats, and different file angles.
[0104] Figure 9 It is a schematic structural diagram of the text detection device provided in the ninth embodiment of the present invention. As Figure 9 shown, the text detection device provided in the embodiment of the present invention includes a second acquisition module 901 and a detection module 902, where:
[0105] The second acquisition module 901 is used to acquire the text image to be detected; the detection module 902 is used to obtain the detection result corresponding to the text image to be detected based on the text image to be detected and the text detection model.
[0106] Specifically, the second acquisition module 901 can acquire the text image to be detected, and the text image to be detected is an image including text, and it is necessary to detect what text the text image to be detected includes.
[0107] The detection module 902 inputs the text image to be detected into the text detection model. After being processed by the text detection model, it outputs the detection result corresponding to the text image to be detected, and the detection result is the text corresponding to the text image to be detected. Among them, the text detection model is constructed based on the text detection model establishment method described in any of the above embodiments.
[0108] The text detection device provided by the embodiment of the present invention can obtain a text image to be detected, and based on the text image to be detected and a text detection model, obtain a detection result corresponding to the text image to be detected. The text in the text image to be detected can be accurately detected through the text detection model, improving the accuracy of text detection.
[0109] The embodiment of the device provided by the embodiment of the present invention can specifically be used to execute the processing flow of the corresponding method embodiment above. Its functions will not be elaborated here and can be referred to the detailed description of the above method embodiment.
[0110] It should be noted that the method for establishing a text detection model, the text detection method and device provided by the embodiment of the present invention can be used in the financial field, and can also be used in any technical field other than the financial field. The embodiment of the present invention does not limit the application fields of the method for establishing a text detection model, the text detection method and device.
[0111] Figure 10 It is a schematic physical structure diagram of an electronic device provided by the tenth embodiment of the present invention. As Figure 10 shown, the electronic device may include: a processor 1001, a communication interface 1002, a memory 1003, and a communication bus 1004. Among them, the processor 1001, the communication interface 1002, and the memory 1003 complete mutual communication through the communication bus 1004. The processor 1001 can call the logical instructions in the memory 1003 to execute the following method: obtain text detection training data; based on the text detection training data and an original model, train to obtain a text detection model; wherein, the original model includes a feature extraction module, and the feature extraction module includes a plurality of cascaded feature extraction units, and each feature extraction unit includes a convolutional pooling sub-unit and an attention sub-unit connected in sequence.
[0112] In addition, when the logical instructions in the above-mentioned memory 1003 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.
[0113] This embodiment discloses a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the methods provided in the above-mentioned method embodiments, for example, including: obtaining text detection training data; training to obtain a text detection model based on the text detection training data and an original model; wherein, the original model includes a feature extraction module, and the feature extraction module includes a plurality of cascaded feature extraction units, and each feature extraction unit includes a convolutional pooling subunit and an attention subunit connected in sequence.
[0114] This embodiment provides a computer-readable storage medium. The computer-readable storage medium stores a computer program. The computer program causes the computer to execute the methods provided in the above-mentioned method embodiments, for example, including: obtaining text detection training data; training to obtain a text detection model based on the text detection training data and an original model; wherein, the original model includes a feature extraction module, and the feature extraction module includes a plurality of cascaded feature extraction units, and each feature extraction unit includes a convolutional pooling subunit and an attention subunit connected in sequence.
[0115] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0116] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks.
[0117] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks.
[0118] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or multiple blocks.
[0119] In the description of this specification, the description with reference to terms such as "one embodiment", "a specific embodiment", "some embodiments", "for example", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0120] The above-described specific embodiments have further elaborated on the objectives, technical solutions, and beneficial effects of the present invention. It should be understood that the above is only the specific embodiments of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A method for establishing a text detection model, characterized in that, Including: Obtain text detection training data; Based on the text detection training data and the original model, train to obtain a text detection model; wherein, the original model includes a feature extraction module, and the feature extraction module includes a plurality of cascaded feature extraction units, and each feature extraction unit includes a convolutional pooling subunit and an attention subunit connected in sequence; The training to obtain a text detection model based on the text detection training data and the original model includes: Perform feature extraction on the input image through the convolutional pooling subunit included in each feature extraction unit to obtain a feature extraction map; Perform feature enhancement processing on the feature extraction map through the attention subunit included in each feature extraction unit to obtain an enhanced feature map; Each attention subunit includes a weight extraction channel and a temporal feature extraction channel, wherein: The weight extraction channel includes a first convolutional layer, a first image reconstruction layer, a second convolutional layer, a second image reconstruction layer, a normalization layer, a third convolutional layer, a first normalization layer, a first activation layer, and a fourth convolutional layer; the temporal feature extraction channel includes a temporal extraction layer; The input ends of the first convolutional layer, the second convolutional layer, and the temporal extraction layer are respectively connected to the output end of the corresponding convolutional pooling subunit; the output end of the first convolutional layer is connected to the input end of the first image reconstruction layer, the output end of the second convolutional layer is connected to the input end of the second image reconstruction layer, the output end of the second image reconstruction layer is connected to the input end of the normalization layer, the cross product result of the output result of the first image reconstruction layer and the output result of the normalization layer is used as the input of the third convolutional layer, the output end of the third convolutional layer is connected to the input end of the first normalization layer, the output end of the first normalization layer is connected to the input end of the first activation layer, and the dot product result of the output result of the first activation layer and the output result of the temporal extraction layer is used as the input of the fourth convolutional layer.
2. The method according to claim 1, characterized in that, The weight extraction channel further includes a fifth convolutional layer, a global pooling layer, a sixth convolutional layer, a second normalization layer, and a second activation layer connected in sequence; The input ends of the fifth convolutional layer and the temporal extraction layer are connected to the output end of the corresponding convolutional pooling subunit; the dot product result of the output result of the weight extraction channel and the output result of the temporal feature extraction channel is used as the output result of the attention subunit.
3. The method according to claim 1, characterized in that, The feature extraction module includes 4-8 cascaded feature extraction units.
4. The method according to claim 1, characterized in that, The training to obtain a text detection model based on the text detection training data and the original model includes: Accelerate the training process of the text detection model through an optimizer.
5. The method according to any one of claims 1 to 4, characterized in that, The text detection training data includes text images with different font sizes, different font formats, and different file angles.
6. A text detection method, characterized in that, Including: Obtain a text image to be detected; Based on the text image to be detected and the text detection model, obtain a detection result corresponding to the text image to be detected; The text detection model is obtained based on the text detection model establishment method according to any one of claims 1 to 5.
7. A device for establishing a text detection model, characterized in that, Including: A first acquisition module, configured to obtain text detection training data; A training module for training a text detection model based on the text detection training data and an original model, wherein the original model includes a feature extraction module, and the feature extraction module includes a plurality of cascaded feature extraction units, and each feature extraction unit includes a convolutional pooling subunit and an attention subunit connected in sequence; The training module includes: An extraction unit for extracting features from an input image through the convolutional pooling subunit included in each feature extraction unit to obtain a feature extraction map; An enhancement unit for performing feature enhancement processing on the feature extraction map through the attention subunit included in each feature extraction unit to obtain an enhanced feature map; Each attention subunit includes a weight extraction channel and a temporal feature extraction channel, wherein: The weight extraction channel includes a first convolutional layer, a first image reconstruction layer, a second convolutional layer, a second image reconstruction layer, a normalization layer, a third convolutional layer, a first normalization layer, a first activation layer, and a fourth convolutional layer; the temporal feature extraction channel includes a temporal extraction layer; The input ends of the first convolutional layer, the second convolutional layer, and the temporal extraction layer are respectively connected to the output end of the corresponding convolutional pooling subunit; the output end of the first convolutional layer is connected to the input end of the first image reconstruction layer, the output end of the second convolutional layer is connected to the input end of the second image reconstruction layer, the output end of the second image reconstruction layer is connected to the input end of the normalization layer, the cross product result of the output result of the first image reconstruction layer and the output result of the normalization layer is used as the input of the third convolutional layer, the output end of the third convolutional layer is connected to the input end of the first normalization layer, the output end of the first normalization layer is connected to the input end of the first activation layer, and the dot product result of the output result of the first activation layer and the output result of the temporal extraction layer is used as the input of the fourth convolutional layer.
8. A text detection device, characterized in that, Includes: A second acquisition module for acquiring a text image to be detected; A detection module for obtaining a detection result corresponding to the text image to be detected based on the text image to be detected and the text detection model; the text detection model is obtained based on the text detection model establishment device of claim 7.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, the method according to any one of claims 1 to 6 is implemented.
10. A computer-readable storage medium, on which a computer program is stored, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 6 are implemented.
11. A computer program product, characterized in that, The computer program product includes a computer program, and when the computer program is executed by the processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Text detection method and system based on feature pyramid and attention fusion
CN113903022A