A method, system and device for detecting reversed characters

By acquiring image sharing features and using inverted character detection methods that detect branches and classification branches, the accuracy of deep learning to identify inverted characters in industrial vision is solved, and the accuracy and efficiency of character recognition are improved.

CN115512358BActive Publication Date: 2025-07-18SHENZHEN LINGYUN VISION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211221905.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-08
Publication Date
2025-07-18
Estimated Expiration
2042-10-08

AI Technical Summary

Technical Problem

Existing deep learning character recognition methods cannot accurately recognize forward and inverted characters in the field of industrial vision, especially in complex scenarios.

Method used

The inverted character detection method is adopted to obtain image sharing features, and a single-channel segmentation diagram is generated using the detection branch to obtain the minimum area of the text line area. The character direction is distinguished by affine transformation and classification branches, and the text box detection results are corrected to identify the inverted characters.

Benefits of technology

It improves the accuracy and efficiency of deep learning character recognition, especially the accuracy of recognizing inverted characters in industrial vision, and enhances the application of deep learning in industrial vision character recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115512358B_ABST
    Figure CN115512358B_ABST
Patent Text Reader

Abstract

This application relates to the technical field of character recognition methods. Specifically, it relates to a method, system, and device for detecting inverted characters, which can, to a certain extent, solve the problem of inaccurate character recognition. The method for detecting inverted characters includes: obtaining an image to be recognized; obtaining shared features in the image and sending the shared features to a detection branch; the detection branch obtains a single-channel segmentation map of the first scale based on the shared features, and obtains a text line region; based on the text line region, the minimum area bounding rectangle of each region in the text line region is obtained, which is characterized as the text box detection result, and the text box detection result includes text box coordinate information; an affine transformation is performed on the shared features according to the coordinate information of the text box to obtain text region features, and the text region features are sent to a classification branch, and the classification branch is used to obtain a character direction classification result based on the text region features; to output the text box detection result or correct the text box detection result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of character recognition methods, and more particularly, to a method, system, and device for detecting inverted characters. Background Art

[0002] Deep learning is widely used in industrial vision, and its performance in complex scenarios is significantly better than that of traditional image processing algorithms. Deep learning is data-driven, mapping images to a high-dimensional feature space and then performing different processing according to different tasks. Among them, deep learning character recognition, as a specialized application of deep learning in the field of vision, can be specifically divided into text detection, text recognition, and text end-to-end recognition according to tasks. The first two can be concatenated as a two-stage solution for deep learning character recognition, that is, first detect the text area in the image and then perform character recognition only on the text area.

[0003] In the implementation of the character recognition process, after the text detection process locates the text area, it only distinguishes the foreground and background of the image and no longer distinguishes specific fonts. Moreover, text recognition methods generally recognize characters in the forward direction from the horizontal direction.

[0004] However, in the business scenarios of the industrial vision field, there are often both forward characters and inverted characters, and the traditional two-stage solution of text detection - text recognition cannot accurately recognize characters. Summary of the Invention

[0005] To solve the problem of inaccurate character recognition, this application provides a method, system, and device for detecting inverted characters.

[0006] The embodiments of this application are implemented as follows:

[0007] The first aspect of the embodiments of this application provides a method for detecting inverted characters, including:

[0008] Obtain an image to be recognized, where the image contains characters;

[0009] Obtain the shared features in the image and send the shared features to the detection branch; the detection branch obtains a single-channel segmentation map of the first scale based on the shared features; where the regions with high confidence in the segmentation map correspond to the text line regions in the image;

[0010] Based on the text line regions, obtain the minimum area bounding rectangle of each region in the text line regions, and the minimum area bounding rectangle is characterized as the text box detection result, and the text box detection result includes text box coordinate information;

[0011] Scale the text box to a second scale, perform an affine transformation on the shared feature according to the coordinate information of the text box to obtain a text region feature, and send the text region feature into a classification branch, where the classification branch is used to obtain a character direction classification result based on the text region feature; wherein, the first scale is greater than the second scale.

[0012] When the character direction classification result is a positive character, directly output the text box detection result; when the character direction classification result is an inverted character, correct the text box detection result and output the corrected text box detection result.

[0013] In some embodiments, in the step of obtaining a single-channel segmentation map of the first scale based on the shared feature in the detection branch, it further includes:

[0014] Generate a single-channel first-scale segmentation map through transposed convolution, and the first-scale segmentation map is a feature map of the same size as the image.

[0015] In some embodiments, in the step of obtaining a single-channel segmentation map of the first scale based on the shared feature in the detection branch, it further includes:

[0016] Generate a single-channel first-scale segmentation map and a threshold map through transposed convolution, and generate a binary map from the first-scale segmentation map according to the threshold map;

[0017] Both the first-scale segmentation map, the threshold map, and the binary map participate in supervised learning.

[0018] In some embodiments, in the step of obtaining the minimum area bounding rectangle of each region in the text line region based on the text line region, it further includes:

[0019] Perform contour detection and contour analysis on the regions with high confidence in the text line region to obtain the minimum area bounding rectangle of each region, and the minimum area bounding rectangle is characterized as the text box detection result, and the text box detection result includes text box coordinate information.

[0020] In some embodiments, in the step of the classification branch obtaining a character direction classification result based on the text region feature, it further includes:

[0021] Perform three downsamplings through a convolutional layer to generate a third-scale feature map, and the third scale is smaller than the second scale;

[0022] Perform global average pooling and fully connected classification processing on the third-scale feature map, and output a two-dimensional one-hot vector;

[0023] Among them, one dimension of the vector represents the forward character, and the other dimension represents the inverted character. The character direction classification result is obtained through the meaning represented by the index corresponding to the maximum value of the vector.

[0024] In some embodiments, the difference between the corrected text box detection result and the text box detection result lies in the different order of reading the vertex coordinates of the text box.

[0025] The second aspect of the embodiments of the present application provides an inverted character detection system, including:

[0026] An image acquisition module, configured to acquire an image to be recognized, where the image contains characters;

[0027] A shared feature acquisition module, which acquires the shared features in the image and transports the shared features to the detection branch;

[0028] A detection module, configured to obtain a single-channel segmentation map of the first scale based on the shared features in the detection branch; wherein, the regions with high confidence in the segmentation map correspond to the text line regions in the image;

[0029] A text box detection result module, configured to obtain the minimum area bounding rectangle of each region in the text line region based on the text line region, and the minimum area bounding rectangle is characterized as the text box detection result, and the text box detection result includes text box coordinate information;

[0030] An affine transformation module, configured to scale the text box to the second scale, perform an affine transformation on the shared features according to the coordinate information of the text box to obtain text region features, and send the text region features to the classification branch; wherein, the first scale is greater than the second scale;

[0031] A classification module, configured to obtain a character direction classification result based on the text region features;

[0032] An output module, configured to directly output the text box detection result when the character direction classification result is a forward character; when the character direction classification result is an inverted character, correct the text box detection result and output the corrected text box detection result.

[0033] In some embodiments, in the step of obtaining a single-channel segmentation map of the first scale based on the shared features in the detection branch, the detection module is further configured to:

[0034] Generate a single-channel first-scale segmentation map through transposed convolution, and the first-scale segmentation map is a feature map of the same size as the image.

[0035] In some embodiments, in the step of obtaining the character direction classification result based on the text region features in the classification branch, the classification module is further configured to:

[0036] Perform three downsamplings through a convolutional layer to generate a third-scale feature map, where the third scale is smaller than the second scale;

[0037] Perform global average pooling and fully connected classification processing on the third-scale feature map, and output a two-dimensional one-hot vector;

[0038] Wherein, one dimension of the vector represents upright characters, and the other dimension represents inverted characters, and the character direction classification result is obtained through the meaning represented by the index corresponding to the maximum value of the vector.

[0039] A second aspect of the embodiments of the present application provides a device for detecting inverted characters, including: a processor, a memory, a communication interface, and a communication bus. The processor, the memory, and the communication interface complete communication with each other through the communication bus;

[0040] The memory is used to store at least one executable instruction, and the executable instruction causes the processor to execute the operations of the inverted character recognition method as described in any one of the first aspects above.

[0041] Advantages of the present application: By sharing the feature layer between the detection branch and the classification branch, the application of deep learning in industrial vision character recognition is improved, and the execution efficiency of deep learning character recognition is accelerated; further, an affine transformation is performed on the shared feature based on the coordinate information of the text box to obtain text region features, which helps to efficiently classify characters in the later stage, and thus helps to accurately recognize characters; further, through the classification of the classification branch, when the character direction classification result is an inverted character, the text box detection result is corrected, and the corrected text box detection result is output, which helps to further improve the accuracy of the character recognition result. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0043] Figure 1 It is a flowchart of a method for detecting inverted characters according to one or more embodiments of the present application;

[0044] Figure 2 It is the text box detection result after being detected by the method for detecting inverted characters according to an embodiment of the present application;

[0045] Figure 3 It is the detection result of the corrected text box after being detected by the reversed character detection method in an embodiment of the present application;

[0046] Figure 4 It is a schematic structural diagram of a reversed character detection method system according to one or more embodiments of the present application;

[0047] Figure 5 It is a schematic structural diagram of a device for the reversed character detection method according to one or more embodiments of the present application. Detailed implementation manners

[0048] To make the objectives, implementation manners, and advantages of the present application clearer, the following will clearly and completely describe the exemplary implementation manners of the present application with reference to the accompanying drawings in the exemplary embodiments of the present application. Obviously, the described exemplary embodiments are only a part rather than all of the embodiments of the present application.

[0049] It should be noted that the brief description of the terms in the present application is only for facilitating the understanding of the subsequent described implementation manners, rather than intending to limit the implementation manners of the present application. Unless otherwise specified, these terms should be understood according to their ordinary and common meanings.

[0050] The terms "first", "second", "third", etc. in the specification, claims, and the above accompanying drawings of the present application are used to distinguish similar or like objects or entities, and do not necessarily mean to limit a specific order or sequence, unless otherwise noted. It should be understood that such used terms can be interchanged under appropriate circumstances.

[0051] The terms "comprising" and "having" and any variations thereof are intended to cover but not exclusively include. For example, a product or device comprising a series of components does not necessarily have to be limited to all the clearly listed components, but may include other components not clearly listed or inherent to these products or devices.

[0052] Figure 1 It is a flowchart of the reversed character detection method in an embodiment.

[0053] As Figure 1 shown, in the first aspect of the present application, a reversed character detection method is disclosed. The reversed character detection method includes the following steps:

[0054] In step 100, an image to be recognized is obtained, and the image contains characters;

[0055] Among them, an image containing characters can be obtained by means such as an operator manually taking a photo, a drone automatically taking a photo, or extracting a photo from a surveillance video, and this image is used as the input image of this method.

[0056] In step 200, shared features in the image are obtained and sent to the detection branch; based on the shared features, the detection branch obtains a single-channel segmentation map at the first scale; among them, the regions with high confidence in the segmentation map correspond to the text line regions in the image.

[0057] Among them, multi-scale features of [1 / 4, 1 / 8, 1 / 16, 1 / 32] are extracted from the image through a mainstream backbone network, and the extracted multi-scale features are decoded layer by layer by adding the current layer features and the upsampled high-level features pixel by pixel and then performing convolution. For example, the 1 / 16-scale feature map and the 1 / 32-scale feature map are added pixel by pixel after upsampling (becoming 1 / 16 after upsampling) to decode the 1 / 16-scale feature map (this feature map can be smoothed by convolution), and in this way, the multi-scale features of [1 / 4, 1 / 8, 1 / 16, 1 / 32] are decoded layer by layer to obtain the decoded multi-scale features of [1 / 4, 1 / 8, 1 / 16, 1 / 32]. The decoded multi-scale features are further fused in the channel dimension to generate a single-scale feature map of [1 / 4], that is, the shared features. Among them, 1 / 4, 1 / 8, 1 / 16, 1 / 32 represent the ratio of the feature map scale to the scale of the image to be recognized.

[0058] It should be noted that each pixel value of this segmentation map represents a confidence level. When the confidence level is greater than the first threshold, it can be considered that the confidence level is high. Generally, the first threshold is set to 0.5. Multiple adjacent pixel values with a confidence level greater than 0.5 cluster together to form a region with a high confidence level.

[0059] After the shared features are sent to the detection branch, the detection branch generates a single-channel segmentation map at the first scale through transposed convolution. This segmentation map is like a mask, and the regions with high confidence levels represent the text line regions. Among them, the first-scale segmentation map is a 1 / 1-scale segmentation map. The reason for generating a single-channel 1 / 1-scale segmentation map is that transposed convolution has the function of upsampling, and upsampling the 1 / 4 scale to 1 / 1 can make the feature map the same size as the original image, which is convenient for subsequent analysis.

[0060] In some embodiments, the detection branch generates a single-channel segmentation map at the first scale and a threshold map through transposed convolution. The first-scale segmentation map generates a binary map according to the threshold map. Based on the first-scale segmentation map, the threshold map, and the binary map all participating in supervised learning for subsequent processing helps to further improve the processing accuracy, can make the predicted text region more refined, suppress background noise, and has more advantages in scenarios with higher accuracy requirements.

[0061] In step 300, based on the text line regions, the minimum area bounding rectangle of each region in the text line regions is obtained. The minimum area bounding rectangle represents the text box detection result, and the text box detection result includes the text box coordinate information.

[0062] In some embodiments, by performing contour detection and contour analysis on the text region, the minimum area bounding rectangle of each region is obtained, which is the text box detection result.

[0063] It should be noted that contour detection is based on topological structure analysis, and the coordinate information of the text box is output through the detection frame segmentation. The method of segmenting the text region and then finding its bounding rectangle can adapt to text lines of various shapes, and at the same time, the rotation angle regression can be omitted during the training process, improving the network convergence situation.

[0064] In step 400, the text box is scaled to the second scale, and an affine transformation is performed on the shared feature according to the coordinate information of the text box to obtain the text region feature, and the text region feature is sent to the classification branch, which is used to obtain the character direction classification result based on the text region feature; wherein, the first scale is greater than the second scale.

[0065] It should be noted that an affine transformation is geometrically defined as an affine transformation or affine mapping between two vector spaces, consisting of a non-singular linear transformation (a transformation using a linear function) followed by a translation transformation. The second scale is 1 / 4 scale. When performing an affine transformation on the shared feature channel by channel, the text region feature can be quickly obtained. After sending the text region feature to the classification branch, the character direction classification result is obtained according to the text region feature to distinguish whether the character is a forward character or an inverted character.

[0066] It should be noted that the height of the text region feature can be fixed at 16, and the width is scaled proportionally. Since the entire network needs to be downsampled by 32 times, the height should be at least 32 pixels in the image, and since the shared feature has been downsampled by 1 / 4, this height should be at least 8.

[0067] In some embodiments, the third-scale feature map is subjected to global average pooling and fully connected classification processing to output a two-dimensional one-hot vector; one dimension of the vector represents the forward character, and the other dimension represents the inverted character, and the character direction classification result is obtained through the meaning represented by the index corresponding to the maximum value of the vector.

[0068] Among them, based on the text area features, the classification branch performs three more downsamplings through the convolutional layer to generate the third-scale feature map. The third scale is smaller than the second scale. This operation is equivalent to generating a [1 / 32]-scale feature map. The [1 / 32]-scale feature map can be regarded as a high-dimensional feature. High-dimensional features are more abstract and can better express the difference between the target and the background. Then, global average pooling is performed based on the third-scale feature map, and global average pooling can be replaced by global max pooling. Then, fully connected classification processing is performed to output a two-dimensional one-hot vector; or the feature map is flattened and then fully connected classification is performed. Flattening means flattening [H, W] into [1, H*W].

[0069] Extract high-dimensional abstract features through a convolutional neural network, and create a detection branch and a classification branch in the shared feature layer. The detection branch outputs the text box detection result. The shared features and the text box detection result are sampled and then input into the classification branch together to obtain the text line direction classification result. Finally, the text box detection result is rotationally corrected according to the text line direction classification result.

[0070] In step 500, when the character direction classification result is a positive character, the text box detection result is directly output; when the character direction classification result is an inverted character, the text box detection result is corrected and the corrected text box detection result is output.

[0071] Among them, as Figure 2 shown, when the character direction classification result is a positive character, the text box detection result is the four vertex coordinates sorted clockwise from the upper left vertex.

[0072] It can be understood that in some other embodiments, when the character direction classification result is an inverted character, the difference between the corrected text box detection result and the text box detection result is that the order of reading the text box vertex coordinates is different. Taking Figure 3 as an example, when the character direction classification result is an inverted character, the corrected text box detection result is the coordinates of the four vertices sorted clockwise from the lower right vertex.

[0073] In a second aspect, the present application also discloses an inverted character detection system, as Figure 4 shown. The inverted character detection system includes the following modules:

[0074] An image acquisition module, configured to acquire an image to be recognized, where the image contains characters;

[0075] A shared feature acquisition module, which acquires the shared features in the image and conveys the shared features to the detection branch;

[0076] A detection module, configured to obtain a single-channel segmentation map at the first scale based on the shared features by the detection branch; among them, the region with high confidence in the segmentation map corresponds to the text line region in the image.

[0077] The text box detection result module is used to obtain the minimum area bounding rectangle of each region in the text line region based on the text line region. The minimum area bounding rectangle is characterized as the text box detection result, and the text box detection result includes the text box coordinate information;

[0078] The affine transformation module is used to scale the text box to the second scale, perform an affine transformation on the shared feature according to the coordinate information of the text box to obtain the text region feature, and send the text region feature to the classification branch; wherein, the first scale is greater than the second scale;

[0079] The classification module is used to obtain the character direction classification result based on the text region feature;

[0080] The output module is used to directly output the text box detection result when the character direction classification result is a positive character; when the character direction classification result is an inverted character, the text box detection result is corrected and the corrected text box detection result is output.

[0081] In some embodiments, in the step of obtaining the single-channel segmentation map of the first scale based on the shared feature in the detection branch, the detection module is further used for:

[0082] Generating a single-channel first-scale segmentation map through transposed convolution, and the first-scale segmentation map is a feature map of the same size as the image.

[0083] In some embodiments, the detection module is further used for: generating a single-channel first-scale segmentation map and a threshold map through transposed convolution, and generating a binary map from the first-scale segmentation map according to the threshold map; the first-scale segmentation map, the threshold map, and the binary map all participate in supervised learning.

[0084] In some embodiments, in the step of obtaining the minimum area bounding rectangle of each region in the text line region based on the text line region, the text box detection result module is further used for performing contour detection and contour analysis on the regions with high confidence in the text line region to obtain the minimum area bounding rectangle of each region, and the minimum area bounding rectangle is characterized as the text box detection result, and the text box detection result includes the text box coordinate information.

[0085] In some embodiments, in the step of obtaining the character direction classification result based on the text region feature in the classification branch, the classification module is further used for:

[0086] Performing three times of downsampling through the convolutional layer to generate a third-scale feature map, the third scale is smaller than the second scale, and the third scale represents the proportion of the image size in the scale feature to the size of the image to be recognized;

[0087] Performing global average pooling and fully connected classification processing on the third-scale feature map, and outputting a two-dimensional one-hot vector;

[0088] Among them, one dimension of the vector represents the forward character, and the other dimension represents the reversed character. The character direction classification result is obtained through the meaning represented by the index corresponding to the maximum value of the vector.

[0089] In some embodiments, the classification module is further configured to correct the text box detection result when the character direction classification result is a reversed character. The difference between the corrected text box detection result and the text box detection result lies in the different order of reading the vertex coordinates of the text box.

[0090] In a third aspect, the present application also discloses a reversed character detection device, as Figure 5 shown. The reversed character detection device includes a processor, a memory, a communication interface, and a communication bus. The processor, the memory, and the communication interface complete mutual communication through the communication bus;

[0091] The memory is used to store at least one executable instruction, and the executable instruction causes the processor to perform the operations of the following reversed character recognition method.

[0092] Perform the operations of the following reversed character recognition method:

[0093] Obtain an image to be recognized, where the image contains characters;

[0094] Obtain the shared features in the image and send the shared features to the detection branch; the detection branch obtains a single-channel segmentation map of the first scale based on the shared features; among them, the region with a high confidence level in the segmentation map corresponds to the text line region in the image;

[0095] Based on the text line region, obtain the minimum area bounding rectangle of each region in the text line region. The minimum area bounding rectangle is characterized as the text box detection result, and the text box detection result includes text box coordinate information;

[0096] Scale the text box to the second scale, perform an affine transformation on the shared features according to the coordinate information of the text box to obtain text region features, and send the text region features to the classification branch. The classification branch is used to obtain the character direction classification result based on the text region features; among them, the first scale is greater than the second scale;

[0097] When the character direction classification result is a forward character, directly output the text box detection result; when the character direction classification result is a reversed character, correct the text box detection result and output the corrected text box detection result.

[0098] In some embodiments, in the step of the detection branch obtaining a single-channel segmentation map of the first scale based on the shared features, a single-channel first-scale segmentation map is generated through transposed convolution, and the first-scale segmentation map is a feature map of the same size as the image.

[0099] In some embodiments, in the step of obtaining a single-channel segmentation map of the first scale based on shared features in the detection branch, a single-channel segmentation map of the first scale and a threshold map are generated through transposed convolution, and a binary map is generated from the single-channel segmentation map of the first scale according to the threshold map; the single-channel segmentation map of the first scale, the threshold map, and the binary map all participate in supervised learning.

[0100] In some embodiments, in the step of obtaining the minimum area bounding rectangle of each region in the text line region based on the text line region, contour detection and contour analysis are performed on the regions with high confidence in the text line region to obtain the minimum area bounding rectangle of each region, and the minimum area bounding rectangle is characterized as the text box detection result, and the text box detection result includes text box coordinate information.

[0101] In some embodiments, in the step of obtaining the character direction classification result based on the text region features in the classification branch, three downsamplings are performed through a convolutional layer to generate a third-scale feature map, where the third scale is smaller than the second scale, and the third scale represents the proportion of the image size in the scale features to the size of the image to be recognized;

[0102] The third-scale feature map is processed through global average pooling and fully connected classification to output a two-dimensional one-hot vector;

[0103] Among them, one dimension of the vector represents the forward character, and the other dimension represents the inverted character, and the character direction classification result is obtained through the meaning represented by the index corresponding to the maximum value of the vector.

[0104] In some embodiments, in the step of obtaining the character direction classification result based on the text region features in the classification branch, when the character direction classification result is an inverted character, the difference between correcting the text box detection result and the text box detection result lies in the different order of reading the vertex coordinates of the text box.

[0105] Fourthly, the present application also discloses a computer-readable storage medium, in which at least one executable instruction is stored. When the executable instruction runs on the inverted character detection device, the inverted character detection device is caused to perform the following operations of the inverted character recognition method:

[0106] Obtain an image to be recognized, where the image contains characters;

[0107] Obtain the shared features in the image and send the shared features to the detection branch; the detection branch obtains a single-channel segmentation map of the first scale based on the shared features; among them, the regions with high confidence in the segmentation map correspond to the text line regions in the image;

[0108] Based on the text line region, obtain the minimum area bounding rectangle of each region in the text line region, and the minimum area bounding rectangle is characterized as the text box detection result, and the text box detection result includes text box coordinate information;

[0109] Scale the text box to the second scale, perform an affine transformation on the shared features according to the coordinate information of the text box to obtain text region features, and send the text region features to the classification branch, where the classification branch is used to obtain a character direction classification result based on the text region features; wherein, the first scale is greater than the second scale.

[0110] When the character direction classification result is a positive character, directly output the text box detection result; when the character direction classification result is an inverted character, correct the text box detection result and output the corrected text box detection result.

[0111] In some embodiments, in the step of obtaining a single-channel segmentation map of the first scale based on the shared features in the detection branch, a single-channel first-scale segmentation map is generated through transposed convolution, and the first-scale segmentation map is a feature map of the same size as the image.

[0112] In some embodiments, in the step of obtaining a single-channel segmentation map of the first scale based on the shared features in the detection branch, a single-channel first-scale segmentation map and a threshold map are generated through transposed convolution, and the first-scale segmentation map generates a binary map according to the threshold map; the first-scale segmentation map, the threshold map, and the binary map all participate in supervised learning.

[0113] In some embodiments, in the step of obtaining the minimum area bounding rectangle of each region in the text line region based on the text line region, perform contour detection and contour analysis on the regions with high confidence in the text line region to obtain the minimum area bounding rectangle of each region, and the minimum area bounding rectangle is characterized as the text box detection result, and the text box detection result includes text box coordinate information.

[0114] In some embodiments, in the step of the classification branch obtaining a character direction classification result based on the text region features, perform three downsamplings through a convolutional layer to generate a third-scale feature map, where the third scale is smaller than the second scale, and the third scale represents the ratio of the image size in the scale features to the size of the image to be recognized;

[0115] Perform global average pooling and fully connected classification processing on the third-scale feature map, and output a two-dimensional one-hot vector;

[0116] Wherein, one dimension of the vector represents a positive character and the other dimension represents an inverted character, and the character direction classification result is obtained through the meaning represented by the index corresponding to the maximum value of the vector.

[0117] In some embodiments, when the character direction classification result is an inverted character, the difference between the corrected text box detection result and the text box detection result is that the order of reading the vertex coordinates of the text box is different.

[0118] The beneficial effects of the embodiments in this part are as follows: by obtaining shared features, obtaining text box detection results based on text line regions, performing an affine transformation on the shared features based on the coordinate information of the text boxes to obtain text region features, and sending the text region features to a classification branch, the classification branch performs text line direction classification on the basis of cropping the shared features; it can effectively identify the direction of characters, and then accurately identify characters.

[0119] Furthermore, through the classification branch, when the character direction classification result is a reversed character, the text box detection result is corrected, and the corrected text box detection result is output, which helps to further improve the accuracy of the character recognition result. The detection branch and the classification branch work together to improve the applicability of deep learning in industrial vision character recognition, empower the text detection-text recognition two-stage solution in the reversed character scenario, and greatly improve the production efficiency of deep learning in industrial vision character recognition.

[0120] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.

[0121] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0122] The above embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several deformations and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. A method for detecting reversed characters, characterized in that, The method includes: Obtain an image to be recognized, where the image contains characters; Obtain the shared features in the image and send the shared features to the detection branch; the detection branch obtains a single-channel segmentation map at a first scale based on the shared features; wherein, the regions with high confidence in the segmentation map correspond to the text line regions in the image; Based on the text line regions, obtain the minimum area bounding rectangle of each region in the text line regions, and the minimum area bounding rectangle is characterized as the text box detection result, and the text box detection result includes text box coordinate information; Scale the text box to a second scale, perform an affine transformation on the shared features according to the coordinate information of the text box to obtain text region features, and send the text region features to the classification branch, and the classification branch is used to obtain a character direction classification result based on the text region features; wherein, the first scale is greater than the second scale; When the character direction classification result is a forward character, directly output the text box detection result; when the character direction classification result is an inverted character, correct the text box detection result and output the corrected text box detection result; In the step of the detection branch obtaining a single-channel segmentation map at a first scale based on the shared features, it further includes: Generate a single-channel first-scale segmentation map and a threshold map through transposed convolution, and the first-scale segmentation map generates a binary map according to the threshold map; The first-scale segmentation map, the threshold map, and the binary map all participate in supervised learning; In the step of the classification branch obtaining a character direction classification result based on the text region features, it further includes: Perform three downsamplings through a convolutional layer to generate a third-scale feature map, and the third scale is smaller than the second scale; Perform global average pooling and fully connected classification processing on the third-scale feature map, and output a two-dimensional one-hot vector; Wherein, one dimension of the vector represents a forward character, and the other dimension represents an inverted character, and the character direction classification result is obtained through the meaning represented by the index corresponding to the maximum value of the vector.

2. The method for detecting inverted characters according to claim 1, characterized in that, In the step of the detection branch obtaining a single-channel segmentation map at a first scale based on the shared features, it further includes: Generate a single-channel first-scale segmentation map through transposed convolution, and the first-scale segmentation map is a feature map of the same size as the image.

3. The inverted character detection method according to claim 1, characterized in that, In the step of obtaining the minimum area bounding rectangle of each region in the text line regions based on the text line regions, it further includes: Perform contour detection and contour analysis on the regions with high confidence in the text line regions to obtain the minimum area bounding rectangle of each region, and the minimum area bounding rectangle is characterized as the text box detection result, and the text box detection result includes text box coordinate information.

4. The method for detecting inverted characters according to claim 1, characterized in that: The difference between the corrected text box detection result and the text box detection result is that the order of reading the vertex coordinates of the text box is different.

5. An inverted character detection system, characterized in that, It includes: An image acquisition module for acquiring an image to be recognized, where the image contains characters; The shared feature acquisition module acquires the shared features in the image and conveys the shared features to the detection branch; The detection module is configured to obtain a single-channel segmentation map of the first scale based on the shared features in the detection branch; wherein, the regions with high confidence in the segmentation map correspond to the text line regions in the image; The text box detection result module is configured to obtain the minimum area bounding rectangle of each region in the text line region based on the text line region, and the minimum area bounding rectangle is characterized as the text box detection result, and the text box detection result includes text box coordinate information; The affine transformation module is configured to scale the text box to the second scale, perform an affine transformation on the shared features according to the coordinate information of the text box to obtain text region features, and send the text region features to the classification branch; wherein, the first scale is greater than the second scale; The classification module is configured to obtain a character direction classification result based on the text region features; The output module is configured to directly output the text box detection result when the character direction classification result is a positive character; when the character direction classification result is an inverted character, correct the text box detection result and output the corrected text box detection result; In the step of obtaining the character direction classification result based on the text region features in the classification branch, the classification module is further configured to: Perform three downsamplings through a convolutional layer to generate a third-scale feature map, and the third scale is smaller than the second scale; Perform global average pooling and fully connected classification processing on the third-scale feature map, and output a two-dimensional one-hot vector; Wherein, one dimension of the vector represents a positive character, and the other dimension represents an inverted character, and the character direction classification result is obtained through the meaning represented by the index corresponding to the maximum value of the vector.

6. The inverted character detection system according to claim 5, wherein In the step of obtaining a single-channel segmentation map of the first scale based on the shared features in the detection branch, the detection module is further configured to: Generate a single-channel first-scale segmentation map through a transposed convolution, and the first-scale segmentation map is a feature map of the same size as the image.

7. An inverted character detection device, characterized in that, Comprising: A processor, a memory, a communication interface and a communication bus, and the processor, the memory and the communication interface complete communication with each other through the communication bus; The memory is used to store at least one executable instruction, and the executable instruction causes the processor to execute the operations of the inverted character detection method according to any one of claims 1-4.

Citation Information

Patent Citations

  • Character recognition method and system based on parameter reconstruction network

    CN114418001A

  • Text recognition method and device

    CN114463761A