Text recognition method and device based on seal image and computer equipment

Through the improved DBNet text detection network and multi-module working together, the problem of identifying curved text in seal images is solved, and accurate text recognition of seal images is achieved in scenes with huge bending.

CN120375346APending Publication Date: 2025-07-25SOUTH CHINA NORMAL UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510259319.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

Existing text recognition methods are difficult to accurately recognize in seal images with huge bending degrees, and spatial transformation networks and unsupervised transformation networks are not effective in such scenarios.

Method used

The improved DBNet text detection network is adopted, combined with the object detection module, text classification module, text direction correction module and text recognition module, and accurate identification of seal images is achieved through text box detection, classification and correction.

Benefits of technology

In seal images with huge bending degree, accurate text recognition effect is achieved, improving the accuracy and reliability of recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120375346A_ABST
    Figure CN120375346A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of text recognition, in particular to a seal image-based text recognition method and device, computer equipment and a storage medium, and the method comprises the steps: obtaining a to-be-recognized seal image and a preset seal image text recognition model, inputting the to-be-recognized seal image into the seal image text recognition model, and carrying out the recognition of the to-be-recognized seal image. Performing textbox detection according to the to-be-recognized seal image and the target detection module to obtain a textbox detection image; classifying the plurality of initial textboxes according to the textbox detection image and a text classification module to obtain a textbox classification result; according to the textbox detection image, the textbox classification result and a text direction correction module, performing text direction correction on the plurality of initial textboxes to obtain a textbox correction image; and performing text recognition on the plurality of corrected textboxes according to the textbox corrected images and a text recognition module to obtain a text recognition result of the to-be-recognized seal image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of text recognition, and particularly to a text recognition method, device, computer device and storage medium based on a seal image. Background Art

[0002] Scene Text Recognition (STR) aims to convert the text regions in complex natural scene images into a text form that can be processed by a computer for other downstream tasks. In natural scenes, factors such as text diversity, environmental complexity, and the uncertainty of image acquisition perspectives pose great challenges to STR.

[0003] Current text recognition methods usually adopt a Spatial Transformer Network (STN) or an unsupervised transformation network MORAN. However, due to the existence of a large number of curved text segments in seal images, the above methods are not applicable in scenarios with a large degree of curvature and it is difficult to accurately recognize the text in seal images. Summary of the Invention

[0004] Based on this, the purpose of the present invention is to provide a text recognition method, device, equipment and storage medium based on a seal image. By detecting the text boxes in the seal image to be recognized and correcting the text directions, and then performing text recognition on the corrected text boxes, accurate text recognition of the seal image is achieved in scenarios with a large degree of curvature.

[0005] In a first aspect, an embodiment of the present application provides a text recognition method based on a seal image, including the following steps:

[0006] Obtain a seal image to be recognized and a preset seal image text recognition model, where the seal image text recognition model includes an object detection module, a text classification module, a text direction correction module, and a text recognition module;

[0007] Input the seal image to be recognized into the seal image text recognition model, and perform text box detection according to the seal image to be recognized and the object detection module to obtain a text box detection image, where the text box detection image includes a plurality of initial text boxes;

[0008] Classify a plurality of the initial text boxes according to the text box detection image and the text classification module to obtain a text box classification result, where the text box classification result includes the text types of a plurality of the initial text boxes;

[0009] Based on the text box detection image, the text box classification result, and the text direction correction module, perform text direction correction on a plurality of the initial text boxes to obtain a text box corrected image, where the text box corrected image includes a plurality of corrected text boxes;

[0010] Based on the text box corrected image and the text recognition module, perform text recognition on a plurality of the corrected text boxes to obtain the text recognition result of the seal image to be recognized.

[0011] In a second aspect, an embodiment of the present application provides a text recognition device based on a seal image, including:

[0012] A data acquisition module, configured to acquire a seal image to be recognized and a preset seal image text recognition model, where the seal image text recognition model includes an object detection module, a text classification module, a text direction correction module, and a text recognition module;

[0013] A text box detection module, configured to input the seal image to be recognized into the seal image text recognition model, and perform text box detection according to the seal image to be recognized and the object detection module to obtain a text box detection image, where the text box detection image includes a plurality of initial text boxes;

[0014] A text box classification module, configured to classify a plurality of the initial text boxes according to the text box detection image and the text classification module to obtain a text box classification result, where the text box classification result includes the text types of a plurality of the initial text boxes;

[0015] A text box correction module, configured to perform text direction correction on a plurality of the initial text boxes according to the text box detection image, the text box classification result, and the text direction correction module to obtain a text box corrected image, where the text box corrected image includes a plurality of corrected text boxes;

[0016] A text box text recognition module, configured to perform text recognition on a plurality of the corrected text boxes according to the text box corrected image and the text recognition module to obtain the text recognition result of the seal image to be recognized.

[0017] In a third aspect, an embodiment of the present application provides a computer device, including a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the text recognition method based on a seal image as described in the first aspect are implemented.

[0018] Fourthly, an embodiment of the present application provides a storage medium storing a computer program, which when executed by a processor implements the steps of the text recognition method based on a seal image as described in the first aspect.

[0019] In the embodiments of the present application, a text recognition method, apparatus, computer device, and storage medium based on a seal image are provided. By detecting and correcting the text direction of the text box in the seal image to be recognized, and performing text recognition on the corrected text box, accurate text recognition of the seal image is achieved in a scenario with a large degree of bending.

[0020] For better understanding and implementation, the present invention will be described in detail below with reference to the accompanying drawings. Description of the Drawings

[0021] Figure 1 It is a schematic flowchart of a text recognition method based on a seal image provided by an embodiment of the present application;

[0022] Figure 2 It is a schematic flowchart of S2 in the text recognition method based on a seal image provided by an embodiment of the present application;

[0023] Figure 3 It is a schematic diagram of S3 in the flowchart of the text recognition method based on a seal image provided by an embodiment of the present application;

[0024] Figure 4 It is a schematic diagram of S4 in the flowchart of the text recognition method based on a seal image provided by an embodiment of the present application;

[0025] Figure 5 It is a schematic diagram of S41 in the flowchart of the text recognition method based on a seal image provided by an embodiment of the present application;

[0026] Figure 6 It is a schematic diagram of S5 in the flowchart of the text recognition method based on a seal image provided by an embodiment of the present application;

[0027] Figure 7 It is a schematic diagram of S6 in the flowchart of the text recognition method based on a seal image provided by another embodiment of the present application;

[0028] Figure 8 It is a schematic structural diagram of a text recognition apparatus based on a seal image provided by an embodiment of the present application;

[0029] Figure 9 It is a schematic structural diagram of a computer device provided by an embodiment of the present application. Detailed Embodiments

[0030] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0031] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit the present application. The singular forms "a", "the", and "said" used in this application and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0032] It should be understood that although the terms first, second, third, etc. may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".

[0033] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a text recognition method based on a seal image provided for an embodiment of the present application. The method includes the following steps:

[0034] S1: Obtain the seal image to be recognized and a preset seal image text recognition model.

[0035] The execution subject of the text recognition method based on the seal image is a recognition device for the text recognition method based on the seal image (hereinafter referred to as the recognition device). In an alternative embodiment, the recognition device may be a computer device, a server, or a server cluster formed by combining multiple computer devices.

[0036] In an alternative embodiment, the recognition device may obtain the seal image to be recognized by querying in a preset database. The recognition device obtains a preset seal image text recognition model, where the seal image text recognition model is an improved DBNet text detection network, and the seal image text recognition model includes a target detection module, a text classification module, a text direction correction module, and a text recognition module.

[0037] S2: Input the seal image to be recognized into the seal image text recognition model, and perform text box detection based on the seal image to be recognized and the target detection module to obtain a text box detection image.

[0038] In this embodiment, the recognition device inputs the seal image to be recognized into the seal image text recognition model, and performs text box detection based on the seal image to be recognized and the target detection module to obtain a text box detection image, where the text box detection image includes a plurality of initial text boxes.

[0039] Please refer to Figure 2 , Figure 2 which is a schematic diagram of S2 in the process of the text recognition method based on seal images provided by an embodiment of the present application, including steps S21 to S24, specifically as follows:

[0040] S21: Input the seal image to be recognized into the backbone network, and perform feature extraction according to a plurality of preset scales to obtain feature extraction maps of a plurality of scales.

[0041] The backbone network adopts the lightweight network Mobilenetv3. In this embodiment, the recognition device inputs the seal image to be recognized into the backbone network, and performs feature extraction according to a plurality of preset scales to obtain feature extraction maps of a plurality of scales.

[0042] Specifically, the recognition device scales the seal image to be recognized to a size of 640×640 and inputs it into Mobilenetv3 for feature extraction at scales of 1 / 2, 1 / 4, 1 / 8, 1 / 16, and 1 / 32 respectively. The number of feature channels for each scale is different, corresponding to 128, 256, 512, 1024, and 2048 respectively.

[0043] S22: Input the feature extraction maps of the plurality of scales into the feature fusion network for feature fusion to obtain feature fusion maps of a plurality of scales.

[0044] In this embodiment, the recognition device inputs the feature extraction maps of the plurality of scales into the feature fusion network for feature fusion to obtain feature fusion maps of a plurality of scales.

[0045] Specifically, the feature fusion network adopts a feature pyramid network. The recognition device inputs the feature extraction maps of the plurality of scales into the feature fusion network, and transfers the semantic information of the high-level feature map to the shallow layer through a 2-fold upsampling operation to fuse with the shallow feature map to obtain feature fusion maps of a plurality of scales.

[0046] S23: Input the feature fusion maps of several scales into the feature enhancement network for feature enhancement to obtain feature enhancement maps of several scales.

[0047] The feature enhancement network adopts RES-EMA. RES-EMA can learn and fuse from the channel and spatial perspectives, capture pixel-level pairwise relationships, establish short-term and long-term dependencies, and further improve the model's ability to express features.

[0048] In this embodiment, the recognition device inputs the feature fusion maps of several scales into the feature enhancement network for feature enhancement to obtain feature enhancement maps of several scales.

[0049] S24: Upsample the feature enhancement maps of several scales to a preset scale respectively, and use the element-wise addition method to integrate the upsampled feature enhancement maps of several scales to obtain a feature integration map. Input the feature integration map into the text box detection network for text box detection to obtain several initial text boxes.

[0050] In this embodiment, the recognition device upsamples the feature enhancement maps of several scales to a preset scale respectively, and uses the element-wise addition method to integrate the upsampled feature enhancement maps of several scales to obtain a feature integration map.

[0051] The recognition device inputs the feature integration map into the text box detection network for text box detection to obtain several initial text boxes.

[0052] Specifically, when the recognition device inputs the feature integration map into the text box detection network, the sigmoid activation function is used to obtain a probability map P. The sigmoid function can map the feature values to the range between 0 and 1, which conforms to the value range of probability and can well represent the probability that each pixel point belongs to the text area.

[0053] Perform a binarization operation according to the probability map P and a preset probability threshold. The probability threshold can be set to 0.3. Compare each pixel value in the probability map with the probability threshold. Set the pixels greater than the probability threshold to 1, indicating belonging to the text area; set the pixels less than or equal to the probability threshold to 0, indicating belonging to the non-text area. Use the findContours contour extraction algorithm in OpenCV to find the contours of all text areas in the binarized image, and use the text areas as the initial text boxes, realizing the preliminary separation of text and non-text areas.

[0054] S3: Classify several initial text boxes according to the text box detection image and the text classification module to obtain text box classification results.

[0055] In this embodiment, the recognition device classifies a plurality of the initial text boxes according to the text box detection image and the text classification module to obtain a text box classification result, where the text box classification result includes the text types of the plurality of the initial text boxes, and the text types include curved text, inclined text, and horizontal text.

[0056] Please refer to Figure 3 , Figure 3 which is a schematic diagram of S3 in the process of the text recognition method based on a seal image provided by an embodiment of the present application, including steps S31 to S32, specifically as follows:

[0057] S31: Obtain the text area area and the minimum circumscribed matrix area of the plurality of the initial text boxes, calculate the difference between the text area area and the minimum circumscribed matrix area of the plurality of the initial text boxes, and if the difference is greater than or equal to the product of the text area area and a preset area threshold value, determine the text type of the initial text box as curved text.

[0058] In this embodiment, the recognition device obtains the text area area and the minimum circumscribed matrix area of the plurality of the initial text boxes, calculates the difference between the text area area and the minimum circumscribed matrix area of the plurality of the initial text boxes, and if the difference is greater than or equal to the product of the text area area and a preset area threshold value, determines the text type of the initial text box as curved text.

[0059] S32: If the difference is less than the product of the text area area and a preset area threshold value, obtain the main direction angle of the initial text box. If the main direction angle is less than a preset angle threshold, determine the text type of the initial text box as horizontal text; if the main direction angle is greater than or equal to the angle threshold, determine the text type of the initial text box as inclined text.

[0060] In this embodiment, if the difference is less than the product of the text area area and a preset area threshold value, the recognition device obtains the main direction angle of the initial text box. If the main direction angle is less than a preset angle threshold, determines the text type of the initial text box as horizontal text; if the main direction angle is greater than or equal to the angle threshold, determines the text type of the initial text box as inclined text.

[0061] S4: According to the text box detection image, the text box classification result, and the text direction correction module, perform text direction correction on the plurality of the initial text boxes to obtain a text box corrected image.

[0062] In this embodiment, the recognition device corrects the text directions of a plurality of the initial text boxes according to the text box detection image, the text box classification result, and the text direction correction module, and obtains a text box corrected image, where the text box corrected image includes a plurality of corrected text boxes.

[0063] Please refer to Figure 4 , Figure 4 which is a schematic diagram of S4 in the process of the text recognition method based on the seal image provided by an embodiment of the present application, including steps S41 to S43, specifically as follows:

[0064] S41: Horizontally correct the text directions of the initial text boxes with the text types of curved text and inclined text to obtain a plurality of intermediate text boxes.

[0065] In this embodiment, the recognition device horizontally corrects the text directions of the initial text boxes with the text types of curved text and inclined text to obtain a plurality of intermediate text boxes.

[0066] Please refer to Figure 5 , Figure 5 which is a schematic diagram of S41 in the process of the text recognition method based on the seal image provided by an embodiment of the present application, including steps S411 to S412, specifically as follows:

[0067] S411: If the text type of the initial text box is curved text, fit a circumcircle to the initial text box to obtain the center and radius of the circumcircle corresponding to the initial text box; construct a polar coordinate system with the center as the pole and the radius as the polar axis, map a plurality of detection points in the initial text box to the polar coordinate system, calculate the maximum difference between the polar angles of the plurality of detection points in the polar coordinate system, and map the plurality of detection points in the polar coordinate system back to the rectangular coordinate system according to the maximum difference to obtain the intermediate text box corresponding to the initial text box.

[0068] In this embodiment, if the text type of the initial text box is curved text, the recognition device fits a circumcircle to the initial text box to obtain the center and radius of the circumcircle corresponding to the initial text box; constructs a polar coordinate system with the center as the pole and the radius as the polar axis, maps a plurality of detection points in the initial text box to the polar coordinate system, calculates the maximum difference between the polar angles of the plurality of detection points in the polar coordinate system, and maps the plurality of detection points in the polar coordinate system back to the rectangular coordinate system according to the maximum difference to obtain the intermediate text box corresponding to the initial text box.

[0069] S412: If the text type of the initial text box is slanted text, use the Hough transform method to perform edge detection on the initial text box to obtain line angle distribution data. Based on the line angle distribution data, obtain the angle corresponding to the line with the highest occurrence frequency as the text slant angle. Rotate the initial text box according to the text slant angle to obtain the intermediate text box corresponding to the initial text box.

[0070] The Hough transform is a classic algorithm for finding lines in an image. It can transform the lines in the image space into the parameter space (Hough space) and determine the lines in the image by finding the peaks in the Hough space.

[0071] In this embodiment, if the text type of the initial text box is slanted text, the recognition device uses the Hough transform method to perform edge detection on the initial text box to obtain line angle distribution data. Based on the line angle distribution data, obtain the angle corresponding to the line with the highest occurrence frequency as the text slant angle. Rotate the initial text box according to the text slant angle to obtain the intermediate text box corresponding to the initial text box.

[0072] S42: Use the initial text box with the text type of horizontal text and the intermediate text box as the input text boxes of the preset horizontal direction detection network to obtain the horizontal direction probability data of the input text boxes. Based on the horizontal direction probability data, take the horizontal direction corresponding to the horizontal direction probability vector with the largest dimension as the horizontal direction detection result of the input text boxes. Correct the text directions of the initial text box with the text type of horizontal text and the intermediate text box according to the horizontal direction detection result to obtain several corrected text boxes of the seal image to be recognized.

[0073] The horizontal direction detection network is a text four-direction classification network model trained based on PPLCNet and is used to recognize the text direction of the text.

[0074] Since the text direction of the intermediate text box may be reversed, in this embodiment, the recognition device uses the initial text box with the text type of horizontal text and the intermediate text box as the input text boxes of the preset horizontal direction detection network to obtain the horizontal direction probability data of the input text boxes. Based on the horizontal direction probability data, take the horizontal direction corresponding to the horizontal direction probability vector with the largest dimension as the horizontal direction detection result of the input text boxes, where the horizontal direction detection result includes the normal result and the reversed result.

[0075] Based on the horizontal direction detection result, the recognition device corrects the text directions of the initial text box and the intermediate text box whose horizontal direction detection result is an inverted result and the text type is horizontal text, to obtain several corrected text boxes of the seal image to be recognized, ensuring that the text boxes in the seal image to be recognized are all in the horizontal direction and are upright.

[0076] S5: Based on the text box corrected image and the text recognition module, perform text recognition on several of the corrected text boxes to obtain the text recognition result of the seal image to be recognized.

[0077] In this embodiment, the recognition device performs text recognition on several of the corrected text boxes based on the text box corrected image and the text recognition module, to obtain the text recognition result of the seal image to be recognized.

[0078] The text recognition module includes an encoding unit and a recognition unit; the encoding unit includes several sequentially connected sub-encoding units; the sub-encoding unit includes a first normalization layer, a feature extraction layer, a second normalization layer, and a multi-layer perceptron; please refer to Figure 6 , Figure 6 is a schematic diagram of S5 in the process of the text recognition method based on the seal image provided by an embodiment of the present application, including steps S51 to S56, specifically as follows:

[0079] S51: Take the text box corrected image as the input image of the first sub-encoding unit, and perform normalization processing according to the first normalization layer to obtain a first normalized feature map.

[0080] In this embodiment, the recognition device takes the text box corrected image as the input image of the first sub-encoding unit, and performs normalization processing according to the first normalization layer to obtain a first normalized feature map.

[0081] S52: Perform feature extraction according to the first normalized feature map and the feature extraction layer to obtain a first feature extraction map, and splice the first normalized feature map and the first feature extraction map to obtain a first feature splicing map.

[0082] In this embodiment, the recognition device performs feature extraction according to the first normalized feature map and the feature extraction layer to obtain a first feature extraction map, and splices the first normalized feature map and the first feature extraction map to obtain a first feature splicing map. Specifically, the feature extraction layer of the first sub-encoding unit of the encoding unit is a convolutional layer, and the feature extraction layers of the remaining sub-encoding units are global multi-head self-attention layers.

[0083] S53: Perform normalization processing based on the first feature stitching map and the second normalization layer to obtain a second normalized feature map; perform a non-linear transformation based on the second normalized feature map and a multi-layer perceptron, and stitch the second feature extraction map obtained from the non-linear transformation with the first feature stitching map to obtain a second feature stitching map, which is used as the output image of the first sub-encoding unit; use the output image of the first sub-encoding unit as the input image of the next sub-encoding unit, and repeat the process to obtain the output image of the last sub-encoding unit, which is used as the feature encoding map.

[0084] In this embodiment, the recognition device performs normalization processing based on the first feature stitching map and the second normalization layer to obtain a second normalized feature map; performs a non-linear transformation based on the second normalized feature map and a multi-layer perceptron, and stitches the second feature extraction map obtained from the non-linear transformation with the first feature stitching map to obtain a second feature stitching map, which is used as the output image of the first sub-encoding unit; uses the output image of the first sub-encoding unit as the input image of the next sub-encoding unit, and repeats the process to obtain the output image of the last sub-encoding unit, which is used as the feature encoding map for text recognition to improve the accuracy of text recognition, as follows:

[0085] F' = LN(F) + Mixingblock(LN(F)) + MLP(LN(LN(F) + Mixingblock(LN(F))))

[0086] In the formula, F is the input image, F' is the output image, LN() is the normalization function, Mixingblock() is the feature extraction function, and MLP() is the multi-layer perceptron function.

[0087] S54: Input the feature encoding map into the recognition unit to calculate the character probabilities, and obtain the first predicted character probability sequences of several corrected text boxes.

[0088] In this embodiment, the recognition device inputs the feature encoding map into the recognition unit to calculate the character probabilities, and obtains the first predicted character probability sequences of several corrected text boxes. Specifically, the feature encoding map is mapped to the length of the character dictionary through a fully connected layer, and then converted into character probabilities through the softmax normalization exponential function to obtain the first predicted character probability sequences of several corrected text boxes. Among them, the first predicted character probability sequence includes the first predicted character probability data of several characters, and the first predicted character probability data includes the first predicted character probability vectors of several dimensions.

[0089] S55: Perform CTC decoding on the feature encoding map to obtain the first character sequence of several corrected text boxes of the seal image to be recognized.

[0090] In this embodiment, the recognition device performs CTC decoding on the feature encoding map to obtain the first character sequence of several corrected text boxes of the seal image to be recognized, where the first character sequence includes several first characters.

[0091] S56: According to the first character sequence of several corrected text boxes and the first predicted character probability sequence, perform multi-path conditional probability calculation and accumulation on the corrected text boxes to obtain the second predicted character probability sequence of several corrected text boxes; according to the second predicted character probability data of several characters in the second predicted character probability sequence, obtain the text recognition results of several corrected text boxes.

[0092] To remove the duplicate characters and delimiters contained in the first character sequence obtained by CTC decoding, the recognition device performs multi-path conditional probability calculation and accumulation on the corrected text boxes according to the first character sequence, the first predicted character probability sequence of several corrected text boxes, and a preset first probability calculation algorithm, to obtain the second predicted character probability sequence of several corrected text boxes, where the second predicted character probability sequence includes the second predicted character probability data of several characters, the second predicted character probability data includes the second predicted character probability vectors of several dimensions, and the first probability calculation algorithm is:

[0093]

[0094] In the formula, is the second predicted character probability data, is the first predicted character probability data of the character corresponding to the t-th time step, is the first character corresponding to the t-th time step, T is the number of time steps, is the first character sequence, and σ() is.

[0095] The recognition device obtains the character corresponding to the second predicted character probability vector with the largest dimension as the predicted character according to the second predicted character probability data of several characters in the second predicted character probability sequence, and obtains the text recognition results of several corrected text boxes.

[0096] In an optional embodiment, it further includes step S6: training the text recognition module; please refer to Figure 6 , Figure 6 is the schematic flowchart of S6 in the text recognition method based on the seal image provided by another embodiment of this application, including steps S61 to S64, as follows:

[0097] S61: Obtain a plurality of sample seal images, input the plurality of sample seal images into the seal image text recognition model, and obtain the feature encoding maps of the plurality of sample seal images, the first predicted character probability sequences and the second predicted character probability sequences of a plurality of corrected text boxes.

[0098] In this embodiment, the recognition device obtains a plurality of sample seal images, inputs the plurality of sample seal images into the seal image text recognition model, and obtains the feature encoding maps of the plurality of sample seal images, the first predicted character probability sequences and the second predicted character probability sequences of a plurality of corrected text boxes.

[0099] S62: Obtain the true character probability sequences of a plurality of corrected text boxes of a plurality of sample seal images, calculate a loss value according to the true character probability sequences and the second predicted character probability sequences, and obtain a first loss value.

[0100] Since the first character sequence obtained by CTC decoding may have multiple-path solutions, during the training process, there may be a feature misalignment phenomenon, and there is ambiguity in the label calculation of the CTC path, which makes it easy for the model to be confused when learning the feature representation of the character corresponding to each time step, that is, the position corresponding to the time step, resulting in missing or recognizing redundant characters, thereby reducing the learning effect of feature alignment and feature representation. In this embodiment, the recognition device obtains the true character probability sequences of a plurality of corrected text boxes of a plurality of sample seal images, and calculates a loss value according to the true character probability sequences, the second predicted character probability sequences and a preset first loss calculation algorithm, and obtains a first loss value, so as to find the probability values corresponding to all paths, and maximize this probability value to guide the model to learn, where the first loss calculation algorithm is:

[0101]

[0102] In the formula, L ctc is the first loss value.

[0103] S63: Perform Attention decoding on the feature encoding maps of the plurality of sample seal images to obtain the second character sequences of a plurality of corrected text boxes of the plurality of sample seal images; according to the second character sequences and the first predicted character probability sequences, perform single-path conditional probability calculation on the corrected text boxes, and obtain the third predicted character probability sequences of a plurality of corrected text boxes of the plurality of sample seal images.

[0104] To better guide the model to learn and reduce the negative impact brought by CTC decoding, in this embodiment, the recognition device performs Attention decoding on the feature encoding maps of several sample seal images to obtain the second character sequences of several corrected text boxes of several sample seal images. Attention decoding explicitly models the dependencies between sequences and does not require merging repeated characters or delimiters.

[0105] The recognition device performs single-path conditional probability calculation on the corrected text boxes according to the second character sequence, the first predicted character probability sequence, and a preset second probability calculation algorithm, to obtain the third predicted character probability sequence of several corrected text boxes of several sample seal images. Among them, the third predicted character probability sequence includes the third predicted character probability data of several characters, the second predicted character probability data includes the third predicted character probability vectors of several dimensions, and the second probability calculation algorithm is:

[0106]

[0107] In the formula, is the second predicted character probability data, is the first predicted character probability data of the character corresponding to the t-th time step, is the second character corresponding to the t-th time step.

[0108] S64: Calculate the loss value according to the true character probability sequence and the third predicted character probability sequence to obtain the second loss value. Train the text recognition module according to the first loss value and the second loss value.

[0109] In this embodiment, the recognition device calculates the loss value according to the true character probability sequence, the third predicted character probability sequence, and a preset second loss calculation algorithm to obtain the second loss value. Among them, the second loss calculation algorithm is:

[0110]

[0111] In the formula, L attention is the second loss value.

[0112] The recognition device obtains the total loss value according to the first loss value, the second loss value, and a preset total loss value calculation algorithm, and trains the text recognition module according to the total loss value. Among them, the total loss value calculation algorithm is:

[0113] L = λ1L ctc + λ2L attention

[0114] Where L is the total loss value, λ1 is the first weight parameter, and λ2 is the second weight parameter.

[0115] Please refer to Figure 8 , Figure 8 FIG. 7 is a schematic structural diagram of a text recognition device based on a seal image provided by an embodiment of the present application. The device can implement all or part of the text recognition device based on the seal image through software, hardware, or a combination of both. The device 8 includes:

[0116] A data acquisition module 81, configured to acquire a seal image to be recognized and a preset seal image text recognition model. The seal image text recognition model includes an object detection module, a text classification module, a text direction correction module, and a text recognition module.

[0117] A text box detection module 82, configured to input the seal image to be recognized into the seal image text recognition model, perform text box detection according to the seal image to be recognized and the object detection module, and obtain a text box detection image, where the text box detection image includes a plurality of initial text boxes.

[0118] A text box classification module 83, configured to classify a plurality of the initial text boxes according to the text box detection image and the text classification module, and obtain a text box classification result, where the text box classification result includes text types of the plurality of initial text boxes.

[0119] A text box correction module 84, configured to perform text direction correction on a plurality of the initial text boxes according to the text box detection image, the text box classification result, and the text direction correction module, and obtain a text box correction image, where the text box correction image includes a plurality of corrected text boxes.

[0120] A text box text recognition module 85, configured to perform text recognition on a plurality of the corrected text boxes according to the text box correction image and the text recognition module, and obtain a text recognition result of the seal image to be recognized.

[0121] In the embodiments of the present application, a seal image to be recognized and a preset seal image text recognition model are obtained through a data acquisition module. The seal image text recognition model includes an object detection module, a text classification module, a text direction correction module, and a text recognition module. The seal image to be recognized is input into the seal image text recognition model through a text box detection module. According to the seal image to be recognized and the object detection module, text box detection is performed to obtain a text box detection image, where the text box detection image includes a plurality of initial text boxes. According to the text box detection image and the text classification module, a text box classification module classifies the plurality of initial text boxes to obtain a text box classification result, where the text box classification result includes the text types of the plurality of initial text boxes. According to the text box detection image, the text box classification result, and the text direction correction module, a text box correction module corrects the text directions of the plurality of initial text boxes to obtain a text box correction image, where the text box correction image includes a plurality of corrected text boxes. According to the text box correction image and the text recognition module, a text box text recognition module performs text recognition on the plurality of corrected text boxes to obtain a text recognition result of the seal image to be recognized. By detecting the text boxes in the seal image to be recognized and correcting the text directions, and performing text recognition on the corrected text boxes, accurate text recognition of the seal image is achieved in a scenario with a large degree of bending.

[0122] Please refer to Figure 9 , Figure 9 which is a schematic structural diagram of a computer device provided by an embodiment of the present application. The computer device 9 includes: a processor 91, a memory 92, and a computer program 93 stored in the memory 92 and executable on the processor 91. The computer device can store multiple instructions, and the instructions are suitable for being loaded and executed by the processor 91 to perform the method steps of the above Figures 1 to 7 shown embodiment. The specific execution process can refer to the specific description of the Figures 1 to 7 shown embodiment, and details are not described herein.

[0123] Among them, the processor 91 may include one or more processing cores. The processor 91 uses various interfaces and circuits to connect various parts within the server. By running or executing instructions, programs, code sets, or instruction sets stored in the memory 92, and by invoking the data in the memory 92, it executes various functions of the text recognition device 8 based on the seal image and processes data. Optionally, the processor 91 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 91 may integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem, etc. Among them, the CPU mainly processes the operating system, user interface, application programs, etc.; the GPU is responsible for rendering and drawing the content to be displayed on the touch display screen; the modem is used to process wireless communication. It can be understood that the above-mentioned modem may not be integrated into the processor 91 and may be implemented separately by a single chip.

[0124] Among them, the memory 92 may include a random access memory (RAM) and may also include a read-only memory (ROM). Optionally, the memory 92 includes a non-transitory computer-readable storage medium. The memory 92 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 92 may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing the operating system, instructions for at least one function (such as touch instructions, etc.), instructions for implementing the above-mentioned method embodiments, etc.; the data storage area may store the data involved in the above-mentioned method embodiments. Optionally, the memory 92 may also be at least one storage device located far from the aforementioned processor 91.

[0125] The embodiment of the present application also provides a storage medium, which can store multiple instructions, and the instructions are suitable for being loaded and executed by a processor to perform the above Figures 1 to 7 method steps of the embodiment, and the specific execution process can be referred to Figures 1 to 7 the specific description of the embodiment, and details are not described herein.

[0126] Those skilled in the art can clearly understand that, for the convenience and conciseness of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0127] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0128] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0129] In the embodiments provided by the present invention, it should be understood that the disclosed device / terminal device and method can be implemented in other ways. For example, the device / terminal device embodiments described above are only illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.

[0130] The unit described as a separate component may or may not be physically separated. The component displayed as a unit may or may not be a physical unit, that is, it can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0131] In addition, in each embodiment of the present invention, the functional units may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above integrated unit may be implemented in the form of hardware or in the form of a software functional unit.

[0132] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above method embodiments of the present invention, it may also be completed by a computer program instructing relevant hardware. The computer program may be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above method embodiments may be implemented. Among them, the computer program includes computer program code, and the computer program code may be in the form of source code, object code, an executable file, or some intermediate form, etc.

[0133] The present invention is not limited to the above embodiments. If various modifications or deformations of the present invention do not depart from the spirit and scope of the present invention, and if these modifications and deformations are within the scope of the claims of the present invention and equivalent technical scope, then the present invention also intends to include these modifications and deformations.

Claims

1. A text recognition method based on a seal image, characterized in that, Including the following steps: Obtain a seal image to be recognized and a preset seal image text recognition model, wherein the seal image text recognition model includes an object detection module, a text classification module, a text direction correction module, and a text recognition module; Input the seal image to be recognized into the seal image text recognition model, and perform text box detection according to the seal image to be recognized and the object detection module to obtain a text box detection image, wherein the text box detection image includes a plurality of initial text boxes; Classify the plurality of initial text boxes according to the text box detection image and the text classification module to obtain a text box classification result, wherein the text box classification result includes the text types of the plurality of initial text boxes; Perform text direction correction on the plurality of initial text boxes according to the text box detection image, the text box classification result, and the text direction correction module to obtain a text box correction image, wherein the text box correction image includes a plurality of corrected text boxes; Perform text recognition on the plurality of corrected text boxes according to the text box correction image and the text recognition module to obtain a text recognition result of the seal image to be recognized.

2. The text recognition method based on a seal image according to claim 1, wherein: The object detection module includes a backbone network, a feature fusion network, a feature enhancement network, and a text box detection network; The performing text box detection according to the seal image to be recognized and the object detection module to obtain a text box detection image includes the steps of: Input the seal image to be recognized into the backbone network, and perform feature extraction according to a plurality of preset scales to obtain feature extraction maps of a plurality of scales; Input the feature extraction maps of the plurality of scales into the feature fusion network for feature fusion to obtain feature fusion maps of a plurality of scales; Input the feature fusion maps of a plurality of scales into the feature enhancement network for feature enhancement to obtain feature enhancement maps of a plurality of scales; Upsample the feature enhancement maps of a plurality of scales to a preset scale respectively, and use the element addition method to integrate the upsampled feature enhancement maps of a plurality of scales to obtain a feature integration map, and input the feature integration map into the text box detection network for text box detection to obtain a plurality of the initial text boxes.

3. The text recognition method based on a seal image according to claim 2, wherein: The text types include curved text, inclined text, and horizontal text; The classifying the plurality of initial text boxes according to the text box detection image and the text classification module to obtain a text box classification result includes the steps of: Obtain the text area and the minimum circumscribed matrix area of the plurality of initial text boxes, calculate the difference between the text area and the minimum circumscribed matrix area of the plurality of initial text boxes, and if the difference is greater than or equal to the product of the text area and a preset area threshold, determine the text type of the initial text box as curved text; If the difference is less than the value obtained by multiplying the area of the text region by a preset area threshold, the main direction angle of the initial text box is obtained. If the main direction angle is less than the preset angle threshold, the text type of the initial text box is determined as horizontal text; if the main direction angle is greater than or equal to the angle threshold, the text type of the initial text box is determined as inclined text.

4. The text recognition method based on a seal image according to claim 3, wherein Performing text direction correction on the several initial text boxes according to the text box detection image, the text box classification result, and the text direction correction module to obtain a text box corrected image, including the steps of: Performing text direction horizontal correction on the initial text boxes with the text type of curved text and inclined text to obtain several intermediate text boxes; Using the initial text boxes with the text type of horizontal text and the intermediate text boxes as the input text boxes of a preset horizontal direction detection network to obtain the horizontal direction probability data of the input text boxes. According to the horizontal direction probability data, taking the horizontal direction corresponding to the horizontal direction probability vector with the largest dimension as the horizontal direction detection result of the input text boxes; according to the horizontal direction detection result, performing text direction correction on the initial text boxes with the text type of horizontal text and the intermediate text boxes to obtain several corrected text boxes of the seal image to be recognized.

5. The text recognition method based on a seal image according to claim 4, wherein The step of performing text direction horizontal correction on the initial text boxes with the text type of curved text and inclined text to obtain several intermediate text boxes includes the steps of: If the text type of the initial text box is curved text, performing circumcircle fitting on the initial text box to obtain the center and radius of the circumcircle corresponding to the initial text box; constructing a polar coordinate system with the center as the pole and the radius as the polar axis, mapping several detection points in the initial text box into the polar coordinate system, calculating the maximum difference between the polar angles of several detection points in the polar coordinate system, and mapping several detection points in the polar coordinate system back to the rectangular coordinate system according to the maximum difference to obtain the intermediate text box corresponding to the initial text box; If the text type of the initial text box is inclined text, using the method of Hough transform to perform edge detection on the initial text box to obtain line angle distribution data, obtaining the angle corresponding to the line with the highest occurrence frequency according to the line angle distribution data as the text inclination angle, and rotating the initial text box according to the text inclination angle to obtain the intermediate text box corresponding to the initial text box.

6. The text recognition method based on a seal image according to claim 5, characterized in that, The text recognition module includes an encoding unit and a recognition unit; the encoding unit includes several sequentially connected sub-encoding units; the sub-encoding unit includes a first normalization layer, a feature extraction layer, a second normalization layer, and a multi-layer perceptron; Performing text recognition on the several corrected text boxes according to the text box corrected image and the text recognition module to obtain the text recognition result of the seal image to be recognized, including the steps of: Using the text box corrected image as the input image of the first sub-encoding unit and performing normalization processing according to the first normalization layer to obtain a first normalized feature map; Feature extraction is performed based on the first normalized feature map and the feature extraction layer to obtain a first feature extraction map. The first normalized feature map and the first feature extraction map are concatenated to obtain a first feature concatenation map; Normalization processing is performed based on the first feature concatenation map and the second normalization layer to obtain a second normalized feature map; non-linear transformation is performed based on the second normalized feature map and a multi-layer perceptron. The second feature extraction map obtained by the non-linear transformation is concatenated with the first feature concatenation map to obtain a second feature concatenation map, which is used as the output image of the first sub-encoding unit; the output image of the first sub-encoding unit is used as the input image of the next sub-encoding unit, and the above steps are repeated to obtain the output image of the last sub-encoding unit, which is used as the feature encoding map; The feature encoding map is input into the recognition unit to calculate the character probability, and a first predicted character probability sequence of a plurality of the corrected text boxes is obtained; CTC decoding is performed on the feature encoding map to obtain a first character sequence of a plurality of corrected text boxes of the seal image to be recognized; Based on the first character sequence and the first predicted character probability sequence of a plurality of the corrected text boxes, multi-path conditional probability calculation and accumulation are performed on the corrected text boxes to obtain a second predicted character probability sequence of a plurality of the corrected text boxes; based on the second predicted character probability data of a plurality of characters in the second predicted character probability sequence, text recognition results of a plurality of the corrected text boxes are obtained.

7. The text recognition method based on a seal image according to claim 6, wherein It further includes the step of training the text recognition module; The training of the text recognition module includes the steps of: A plurality of sample seal images are obtained, and the plurality of sample seal images are input into the seal image text recognition model to obtain a feature encoding map of the plurality of sample seal images, a first predicted character probability sequence, and a second predicted character probability sequence of a plurality of corrected text boxes; A true character probability sequence of a plurality of corrected text boxes of a plurality of sample seal images is obtained, and a first loss value is obtained by calculating the loss value based on the true character probability sequence and the second predicted character probability sequence; Attention decoding is performed on the feature encoding map of a plurality of the sample seal images to obtain a second character sequence of a plurality of corrected text boxes of a plurality of the sample seal images; based on the second character sequence and the first predicted character probability sequence, single-path conditional probability calculation is performed on the corrected text boxes to obtain a third predicted character probability sequence of a plurality of corrected text boxes of a plurality of the sample seal images; A second loss value is obtained by calculating the loss value based on the true character probability sequence and the third predicted character probability sequence, and the text recognition module is trained based on the first loss value and the second loss value.

8. A text recognition device based on a seal image, characterized in that, It includes: A data acquisition module, which is used to obtain a seal image to be recognized and a preset seal image text recognition model, where the seal image text recognition model includes an object detection module, a text classification module, a text direction correction module, and a text recognition module; A text box detection module, configured to input the to-be-recognized seal image into the seal image text recognition model, perform text box detection based on the to-be-recognized seal image and a target detection module, and obtain a text box detection image, wherein the text box detection image includes a plurality of initial text boxes; A text box classification module, configured to classify the plurality of initial text boxes according to the text box detection image and a text classification module, and obtain a text box classification result, wherein the text box classification result includes the text types of the plurality of initial text boxes; A text box correction module, configured to perform text direction correction on the plurality of initial text boxes according to the text box detection image, the text box classification result, and a text direction correction module, and obtain a text box correction image, wherein the text box correction image includes a plurality of corrected text boxes; A text box text recognition module, configured to perform text recognition on the plurality of corrected text boxes according to the text box correction image and a text recognition module, and obtain a text recognition result of the to-be-recognized seal image.

9. A computer device, characterized in that, It includes a processor, a memory, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the text recognition method based on a seal image according to any one of claims 1 to 7 are implemented.

10. A storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the text recognition method based on a seal image according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Method, device, processor and readable storage medium for realizing high-precision identification of small inclined character labels for electrical cabinet pressing plate

    CN121600517A

  • Round seal text recognition method based on color filtering and coordinate transformation

    CN122090433A