Method and apparatus for detecting single characters in an image, electronic device, and storage medium

By detecting text lines in the image and extracting attention maps, the problems of single-word detection error and missed detection are solved, which reduces training costs, improves detection accuracy, and obtains character coordinates more accurate.

CN115294590BActive Publication Date: 2025-06-27SHANGHAI HONGJI INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210958845.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-10
Publication Date
2025-06-27
Estimated Expiration
2042-08-10

AI Technical Summary

Technical Problem

The prior art has problems of false detection and missed detection in single-word detection in images, which is expensive to train and the accuracy of CTC decoding is not ideal, resulting in inaccurate calculation of single-word coordinates.

Method used

The target image is detected through the text line detection algorithm, the feature map of the text line image is extracted and the attention map is calculated, and the attention sub-map of the characters is extracted from the attention map for Gaussian distribution fitting, and the coordinates of the center point of the character are determined.

Benefits of technology

Reduces training costs, improves detection accuracy, and obtains character coordinates more accurate, and does not require additional single-word position annotations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115294590B_ABST
    Figure CN115294590B_ABST
Patent Text Reader

Abstract

The present application provides a method and apparatus for detecting single characters in an image, an electronic device, and a storage medium. The method includes: performing text line detection on a target image through a text line detection algorithm to obtain coordinate information of a text line box; obtaining a text line image from the target image according to the coordinate information of the text line box; extracting a first feature map of the text line image through a deep convolutional network and calculating an attention map corresponding to the first feature map; extracting an attention sub-map corresponding to each character from the attention map, performing Gaussian distribution fitting on the attention sub-map, and using the expected value as the horizontal coordinate of the center point of the character in the attention map; and determining the position coordinates of each character in the target image according to the horizontal coordinate of the center point of each character in the attention map. This solution does not require additional single-character position annotation and has no additional training process, reducing the training cost. Since the attention map is utilized, the obtained character coordinates are more accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to a method and device for detecting a single word in an image, an electronic device, and a computer-readable storage medium. Background Art

[0002] As a common technology in computer vision, text image recognition is widely used in practical projects such as document information extraction, certificate recognition, and qualification review, especially in RPA (Robotic Process Automation) projects. Currently, commonly used text image detection and recognition algorithms can only output results based on the line level, that is, the coordinates of the text line and the corresponding content. However, in some practical scenarios, results based on the single word level are also required, that is, each single word recognition result and its corresponding coordinates on the original image.

[0003] There are three types of existing solutions: single word detection, text line and single word joint detection, and reverse calculation of decoding results based on CTC (Connectionist temporal classification) text recognition. The main problem with the method based on single word detection is that since the features of a single character are not significant enough, there will be more false detections and missed detections during detection, resulting in inaccurate coordinates. The main problem with the method of joint detection of text lines and single words is that it is necessary to mark the coordinates of single words during the training process, which greatly increases the training cost. If the coordinates of single words are obtained only through weak supervision methods instead of manual annotation, the accuracy of single word detection will be compromised. The method of reverse calculation of decoding results based on CTC text recognition has the disadvantage of relying on the accuracy of CTC decoding. When recognizing long text lines, the accuracy of CTC decoding is often not ideal, which makes the calculation of single word coordinates not accurate enough. Summary of the invention

[0004] The embodiment of the present application provides a method for detecting a single word in an image, so as to reduce training costs and improve detection accuracy.

[0005] The present application embodiment provides a method for detecting a single word in an image, comprising:

[0006] Perform text line detection on the target image through a text line detection algorithm to obtain the coordinate information of the text line frame;

[0007] Obtaining a text line image from the target image according to the coordinate information of the text line frame;

[0008] Extracting a first feature map of the text line image through a deep convolutional network, and calculating an attention map corresponding to the first feature map;

[0009] Extract the attention sub - graph corresponding to each character from the attention map, perform Gaussian distribution fitting on the attention sub - graph, and use the expected value as the horizontal coordinate of the center point of the character in the attention map;

[0010] Determine the position coordinates of each character in the target image according to the horizontal coordinates of the center points of each character in the attention map.

[0011] In one embodiment, the obtaining the text line image from the target image according to the coordinate information of the text line box includes:

[0012] Calculate the coordinate information of the positive rectangular box according to the coordinate information of the text line box;

[0013] Determine the perspective transformation matrix according to the coordinate information of the text line box and the coordinate information of the positive rectangular box;

[0014] Calculate the pixel coordinates mapped in the target image for each pixel coordinate in the positive rectangular box according to the perspective transformation matrix;

[0015] Calculate the pixel value of each pixel coordinate in the positive rectangular box by using bilinear interpolation according to the pixel values of the pixel coordinates in the target image, and obtain the text line image.

[0016] In one embodiment, the extracting the first feature map of the text line image through a deep convolutional network and calculating the attention map corresponding to the first feature map includes:

[0017] Extract the first feature map of the text line image through a deep convolutional network, and connect the first feature map head - to - tail along the vertical dimension to obtain a second feature map;

[0018] Multiply the second feature map by the transpose of the second feature map to obtain the attention map.

[0019] In one embodiment, the extracting the attention sub - graph corresponding to each character from the attention map and performing Gaussian distribution fitting on the attention sub - graph, and using the expected value as the horizontal coordinate of the center point of the character in the attention map includes:

[0020] Extract the attention values of each row from the attention map to obtain the attention sub - graph of the characters corresponding to the row;

[0021] For each character, perform Gaussian distribution fitting on the attention sub - graph of the character with the horizontal coordinate of the attention map as the independent variable and the attention value as the dependent variable to obtain the expected value of the independent variable;

[0022] Use the expected value as the horizontal coordinate of the center point of the character in the attention map.

[0023] In one embodiment, determining the position coordinates of each character in the target image according to the horizontal coordinates of the center points of each character in the attention map includes:

[0024] Calculating the segmentation point coordinates between every two adjacent characters in the attention map according to the horizontal coordinates of the center points of each character in the attention map;

[0025] According to the segmentation point coordinates between every two adjacent characters in the attention map, inversely calculating the segmentation coordinates between adjacent characters on the text line image by using bilinear interpolation;

[0026] Determining the position coordinates of each character in the target image according to the segmentation coordinates between adjacent characters on the text line image.

[0027] In one embodiment, inversely calculating the segmentation coordinates between adjacent characters on the text line image by using bilinear interpolation according to the segmentation point coordinates between every two adjacent characters in the attention map includes:

[0028] Calculating the segmentation coordinates between adjacent characters on the text line image corresponding to the segmentation point coordinates between every two adjacent characters in the attention map by using bilinear interpolation according to the width and height of the attention sub-map and the width and height of the text line image.

[0029] In one embodiment, determining the position coordinates of each character in the target image according to the segmentation coordinates between adjacent characters on the text line image includes:

[0030] Mapping to obtain the boundary coordinates of adjacent characters in the target image through the inverse matrix of the perspective transformation matrix according to the segmentation coordinates between adjacent characters on the text line image;

[0031] Determining the position coordinates of each character in the target image according to the boundary coordinates of adjacent characters in the target image.

[0032] The embodiment of the present application further provides a single-character detection device in an image, including:

[0033] A text line detection module, configured to perform text line detection on a target image through a text line detection algorithm to obtain the coordinate information of the text line box;

[0034] A text line image obtaining module, configured to obtain a text line image from the target image according to the coordinate information of the text line box;

[0035] An attention map obtaining module, configured to extract a first feature map of the text line image through a deep convolutional network and calculate the attention map corresponding to the first feature map;

[0036] A central coordinate calculation module, configured to extract an attention sub-graph corresponding to each character from the attention map, perform Gaussian distribution fitting on the attention sub-graph, and use the expected value as the horizontal coordinate of the center point of the character in the attention map;

[0037] A character coordinate determination module, configured to determine the position coordinates of each character in the target image according to the horizontal coordinate of the center point of each character in the attention map.

[0038] An embodiment of the present application further provides an electronic device, including:

[0039] A processor;

[0040] A memory for storing instructions executable by the processor;

[0041] Wherein, the processor is configured to execute the above-mentioned method for detecting single characters in an image.

[0042] An embodiment of the present application further provides a computer-readable storage medium, storing a computer program, which can be executed by a processor to complete the above-mentioned method for detecting single characters in an image.

[0043] The technical solution provided by the above embodiment of the present application extracts a first feature map of a text line image through a deep convolutional network, and calculates an attention map corresponding to the first feature map; extracts an attention sub-graph corresponding to each character from the attention map, performs Gaussian distribution fitting on the attention sub-graph, and uses the expected value as the horizontal coordinate of the center point of the character in the attention map; determines the position coordinates of each character in the target image according to the horizontal coordinate of the center point of each character in the attention map. No additional single-character position annotation is required, and there is no additional training process, reducing the training cost. Since the attention map is used, the obtained character coordinates are more accurate. Description of the Drawings

[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required to be used in the embodiments of the present application.

[0045] Figure 1 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application;

[0046] Figure 2 It is a schematic flow chart of a method for detecting single characters in an image provided by an embodiment of the present application;

[0047] Figure 3 is Figure 2 It is a detailed flow chart corresponding to step S220 in the embodiment;

[0048] Figure 4 It is a schematic diagram showing the change process from the first feature map to the second feature map provided by an embodiment of the present application;

[0049] Figure 5 and Figure 6 and Figure 7 and Figure 8 and Figure 9 are Gaussian distribution fitting curve diagrams corresponding to five characters in the text line image "Comprehensive Hospital" provided by an embodiment of the present application;

[0050] Figure 10 is Figure 2 a detailed flowchart corresponding to step S250 in the corresponding embodiment;

[0051] Figure 11 is a schematic diagram of the dividing line between adjacent characters in the text line image provided by an embodiment of the present application;

[0052] Figure 12 is a schematic diagram of the position coordinates of each character in the target image provided by an embodiment of the present application;

[0053] Figure 13 is a block diagram of a single-character detection device in an image provided by an embodiment of the present application. Detailed implementation manners

[0054] Next, the technical solutions in the embodiments of the present application will be described in conjunction with the accompanying drawings in the embodiments of the present application.

[0055] Similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of the present application, terms such as "first" and "second" are only used for distinguishing descriptions and cannot be construed as indicating or implying relative importance.

[0056] Figure 1 is a schematic structural diagram of an electronic device provided by an embodiment of the present application. The electronic device 100 can be used to execute the single-character detection method in the image provided by an embodiment of the present application. As Figure 1 shown, the electronic device 100 includes: one or more processors 102, and one or more memories 104 that store processor-executable instructions. Among them, the processor 102 is configured to execute the single-character detection method in the following embodiments of the present application.

[0057] The processor 102 can be a gateway, a smart terminal, or a device including a central processing unit (CPU), an image processing unit (GPU), or other forms of processing units with data processing capabilities and / or instruction execution capabilities. It can process the data of other components in the electronic device 100 and can also control other components in the electronic device 100 to perform desired functions.

[0058] The memory 104 can include one or more computer program products, and the computer program products can include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory can include, for example, random access memory (RAM) and / or cache memory, etc. The non-volatile memory can include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions can be stored on the computer-readable storage media, and the processor 102 can run the program instructions to implement the method for detecting single words in an image described below. Various application programs and various data can also be stored in the computer-readable storage media, such as various data used and / or generated by the application programs, etc.

[0059] In one embodiment, Figure 1 It is shown that the electronic device 100 can further include an input device 106, an output device 108, and a data acquisition device 110. These components are interconnected through a bus system 112 and / or other forms of connection mechanisms (not shown). It should be noted that Figure 1 The components and structures of the illustrated electronic device 100 are exemplary and not restrictive. According to needs, the electronic device 100 can also have other components and structures.

[0060] The input device 106 can be a device used by a user to input instructions and can include one or more of a keyboard, a mouse, a microphone, and a touch screen, etc. The output device 108 can output various information (such as images or sounds) to the outside (for example, to the user) and can include one or more of a display, a speaker, etc. The data acquisition device 110 can acquire an image of an object and store the acquired image in the memory 104 for use by other components. Exemplarily, the data acquisition device 110 can be a camera.

[0061] In one embodiment, the various devices in the exemplary electronic device 100 for implementing the method for detecting single words in an image of the embodiments of the present application can be integrally arranged or dispersedly arranged. For example, the processor 102, the memory 104, the input device 106, and the output device 108 can be integrally arranged, while the data acquisition device 110 is separately arranged.

[0062] In one embodiment, the exemplary electronic device 100 for implementing the single-character detection method in the image of the embodiments of the present application can be implemented as an intelligent terminal such as a smart phone, a tablet computer, a server, a desktop computer, a smart watch, a vehicle-mounted device, etc.

[0063] Figure 2 is a schematic flowchart of a single-character detection method in an image provided by the embodiments of the present application. As Figure 2 shown, the method includes: step S210 - step S250.

[0064] Step S210: Perform text line detection on the target image through a text line detection algorithm to obtain the coordinate information of the text line box.

[0065] Among them, the text line detection algorithm can be an existing common text line detection algorithm such as ctpn, seglink, dbnet, etc. The target image refers to an image containing one or more lines of text. The following solution needs to detect the position coordinates of each character from the target image.

[0066] The text line box refers to the smallest quadrilateral that encloses a line of text. The coordinate information of the text line box can be the coordinates of the four vertices of the quadrilateral. Specifically, starting from the upper left corner and rotating clockwise, the coordinates of the four vertices can be (x1, y1), (x2, y2), (x3, y3), (x4, y4) in sequence.

[0067] In one embodiment, before the above step S210, the input original image can be preprocessed to obtain the above target image. The preprocessing method can include direction correction, brightness adjustment, contrast adjustment, etc. Among them, direction correction refers to straightening the tilted or rotated original image.

[0068] Step S220: Obtain a text line image from the target image according to the coordinate information of the text line box.

[0069] When the text line box is a regular rectangular box, the image corresponding to the coordinate information can be directly intercepted from the target image according to the coordinate information of the text line box, that is, the image within the text line box, as the text line image.

[0070] When the text line box is not a regular rectangular box, the inside of the text line box can be mapped into a regular rectangular box through perspective transformation, and then a rectangular text line image can be obtained. Specifically, as Figure 3 shown, the above step S220 can include the following steps S221 - step S224.

[0071] Step S221: Calculate the coordinate information of the regular rectangular box according to the coordinate information of the text line box.

[0072] For example, assume that the coordinate information of the text line box is (x1, y1), (x2, y2), (x3, y3), (x4, y4). Then the width of the positive rectangular box is w = ((x2 - x1) + (x3 - x4)) / 2, and the height is h = ((y3 - y2) + (y4 - y1)) / 2. The four vertex coordinates of the positive rectangular box can be (0, 0), (w - 1, 0), (w - 1, h - 1), (0, h - 1). Among them, the width and height can be represented by the number of pixels.

[0073] Step S222: Determine the perspective transformation matrix according to the coordinate information of the text line box and the coordinate information of the positive rectangular box.

[0074] Specifically, use the corresponding coordinates of the four vertices of the text line box and the positive rectangular box to solve the perspective transformation matrix M (with a size of 3 * 3). The principle is as follows:

[0075]

[0076] Among them, is the perspective transformation matrix M, is one of the vertex coordinates of the text line box, are unknowns. is the corresponding vertex coordinate of the positive rectangular box. Since there are four sets of corresponding points, the perspective transformation matrix M can be obtained by solving the system of equations.

[0077] Step S223: Calculate the pixel coordinates mapped in the target image for each pixel coordinate in the positive rectangular box according to the perspective transformation matrix.

[0078] Among them, each pixel coordinate in the positive rectangular box refers to the coordinate of each pixel point in the positive rectangular box. For example, the coordinate of the lower left vertex can be (0, 0), the coordinate of the second pixel horizontally is (0, 1), the coordinate of the third pixel horizontally is (0, 2) …… and so on. The coordinate of the second pixel vertically is (1, 0), the coordinate of the third pixel vertically is (2, 0) and so on.

[0079] The perspective transformation matrix M is a calculated known quantity. According to the above formula (1), each pixel coordinate in the positive rectangular box is substituted into the above formula (1) as the value of ( ), and the value of the corresponding pixel coordinate (x, y) in the target image can be calculated.

[0080] Step S224: Calculate the pixel value of each pixel coordinate in the positive rectangular box by using bilinear interpolation according to the pixel value of the pixel coordinate in the target image, and obtain the text line image.

[0081] Assume that the pixel value of the pixel coordinate ( ) in the positive rectangular box is , the pixel coordinates mapped by the content in the positive rectangular box in the target image are ( ), then: ). Then:

[0082] ··· (2)

[0083] Among them, respectively represent rounding up and down, represents the pixel value of the pixel coordinate in the target image. Through the above formula (2), the pixel value of each pixel coordinate in the positive rectangular box can be calculated. The pixel values of each pixel coordinate in the positive rectangular box constitute the text line image.

[0084] Step S230: Extract the first feature map of the text line image through a deep convolutional network, and calculate the attention map corresponding to the first feature map.

[0085] In one embodiment, before the above step S230, the size of the text line image can be adjusted to a preset size first, so as to facilitate the processing of the subsequent text recognition model. Specifically, the text line image can be adjusted to the preset size by using bilinear interpolation, such as 32*240. Or, the height of the text line image can be adjusted to a preset value, and the aspect ratio of the text line image is kept unchanged.

[0086] The network structure of the text recognition model adopts the form of CNN + self-attention + CrossEntropy. The text line image is processed by CNN (deep convolutional network) to obtain a feature map with a size of w’*h’*d’. For distinction, it is called the first feature map. w’ represents the width, h’ represents the height, and d’ represents the number of channels.

[0087] In one embodiment, the first feature maps can be connected end to end in the vertical dimension to obtain the second feature map F. Connecting end to end in the vertical dimension means that the feature values of the first column, the second column, the third column... are connected into a new column. Thus, a column can be obtained for each channel's feature map, and finally the second feature map is arranged. As Figure 4 shown, assuming the size of the first feature map is w’*h’*d’ = 30*2*4, taking the first channel as an example, in the order marked 1, 2, 3, 4... in the figure, from left to right, from top to bottom, it is arranged in a column. Similarly, the data of the second channel is arranged in a column, the data of the third channel is arranged in a column, and the data of the fourth channel is arranged in a column. Thus, 4 columns are obtained for the four channels, and the second feature map with the number of rows L = w’*h’ = 30*2 = 60 is obtained.

[0088] Afterwards, according to the calculation method of self-attention, the second feature map F and the transposed F of the second feature map T Multiply them together to get the attention map.

[0089] For example, assuming that the size of the second feature map is L*d'=60*4, the transpose of the second feature map is 4*60, so the second feature map F and the transpose of the second feature map F T By multiplying them together, we can get an attention map of size L*L.

[0090] Step S240: extracting an attention sub-graph corresponding to each character from the attention graph, performing Gaussian distribution fitting on the attention sub-graph, and taking the expected value as the horizontal coordinate of the center point of the character in the attention graph.

[0091] The attention subgraph corresponding to each character is extracted from the attention graph through CrossEntropy (cross entropy module). The attention subgraph is a sequence of attention values ​​extracted from the attention graph.

[0092] For example, if the size of the attention map is L*L, the size of the attention sub-map corresponding to each character is L×1. Specifically, the attention value of each row can be extracted from the attention map to obtain the attention sub-map of the character corresponding to the row. For example, the attention value of the first row is the attention sub-map of the first character. The attention value of the second row is the attention sub-map of the second character, and so on.

[0093] For each character, the horizontal coordinate of the attention map is used as the independent variable x’’ , attention value is the dependent variable f ( x’’ ), Gaussian distribution fitting is performed on the attention subgraph (x'', f(x'')) of the character to obtain the expected value and variance of the independent variable. The expected value can be used as the horizontal coordinate value of the center point of the character in the attention graph. Since the text line image is a rectangular horizontal text line, the y-axis coordinate is not considered. Figure 5 , Figure 6 , Figure 7 , Figure 8 , Figure 9 This is a Gaussian distribution fitting curve diagram corresponding to the five characters in the text line image "General Hospital". Figures 5 to 9 It can be seen that the expected value of the Gaussian curve corresponding to "综合" is between 0-10, the expected value of the Gaussian curve corresponding to "合" is between 0-20, the expected value of the Gaussian curve corresponding to "性" is around 30, the expected value of the Gaussian curve corresponding to "医" is around 40, and the expected value of the Gaussian curve corresponding to "院" is around 50. Therefore, the expected value can be used to represent the horizontal coordinates of each character in the attention map.

[0094] Step S250: Determine the position coordinates of each character in the target image according to the horizontal coordinates of the center points of each character in the attention map.

[0095] The position coordinates can be the four vertex coordinates of a quadrilateral frame enclosing a single character.

[0096] In one embodiment, as Figure 10 shown, the above step S250 specifically includes: step S251 - step S253.

[0097] Step S251: Calculate the segmentation point coordinates between every two adjacent characters in the attention map according to the horizontal coordinates of the center points of each character in the attention map.

[0098] The segmentation point coordinates refer to the horizontal coordinates of the segmentation line between adjacent characters in the attention map. Assume that the horizontal coordinates of the center points of any two adjacent characters are (xa) and (xb), and calculate the segmentation point coordinates as xs = (xa + xb) / 2.

[0099] Step S252: According to the segmentation point coordinates between every two adjacent characters in the attention map, use bilinear interpolation to inversely calculate the segmentation coordinates between adjacent characters on the text line image.

[0100] Among them, the segmentation coordinates refer to the coordinates of the two endpoints of the segmentation line between two adjacent characters on the text line image.

[0101] For example, assume that the segmentation point coordinates of any two adjacent characters in the attention map are (xs), assume that the width and height of the attention sub - map are L*1, and the width and height of the text line image are (Wt, Ht). Use bilinear interpolation to inversely calculate the segmentation coordinates on the text line image. Then the horizontal coordinate of the segmentation line corresponding to the segmentation point coordinate xs on the text line image is: xs’ = round(Wt / L*xs), and as Figure 11 shown, the y - axis coordinate of point A on the lower edge of the segmentation line is 0, and the y - axis coordinate of point B on the upper edge is round((Ht - 1) / 1). round is the rounding operation. That is to say, the coordinates of point A can be expressed as (xs’, 0), and the coordinates of point B can be expressed as (xs’, round((Ht - 1) / 1)). Since the lower - left first pixel starts from (0, 0), the coordinate is represented by Ht - 1.

[0102] Step S253: Determine the position coordinates of each character in the target image according to the segmentation coordinates between adjacent characters on the text line image.

[0103] As Figure 12As shown in the figure, suppose a text recognition result has six characters, such as "Nanjing Yangtze River Bridge". Referring to the above, there is a dividing line between each two adjacent characters in the text line image, and each dividing line corresponds to the coordinates with two endpoints, namely the coordinates of point A and point B described above. Thus, the position coordinates of each character can be obtained, for example, the coordinates of the character "南" are (x0, y0), (xs0, yu0), (xs0, yl0), (x3, y3). That is, the purpose of single word detection in the target image is achieved.

[0104] In one embodiment, it is assumed that the text line image is obtained by Figure 3 The process shown is to obtain a text line image in a regular rectangular frame after perspective transformation from a text line frame. Therefore, the above step S253 can specifically include the following steps: according to the segmentation coordinates between adjacent characters on the text line image, the boundary coordinates of adjacent characters in the target image are mapped by the inverse matrix of the perspective transformation matrix; according to the boundary coordinates of adjacent characters in the target image, the position coordinates of each character in the target image are determined.

[0105] The boundary coordinates refer to the coordinates of the two endpoints of the dividing line between adjacent characters in the target image. For example, assuming that the dividing coordinates between adjacent characters on the text line image, that is, the coordinates of the upper and lower endpoints are (xs', yu') and (xs', yl'), respectively, using the inverse matrix M of the perspective transformation matrix M -1 , calculate the coordinates of the two endpoints of the segmentation line between adjacent characters in the target image: (xs”,yu”,z”)= (xs',yu',1)M -1 , (xs”,yl”,z”)= (xs’,yl’,1)M -1 (xs”, yu”) represents the upper endpoint coordinates of the dividing line between adjacent characters in the target image, and (xs”, yl”) represents the upper endpoint coordinates of the dividing line between adjacent characters in the target image. Figure 12 , the position coordinates of each character in the target image can be obtained.

[0106] The technical solution provided by the above embodiment of the present application extracts the first feature map of the text line image through a deep convolutional network, and calculates the attention map corresponding to the first feature map; extracts the attention sub-map corresponding to each character from the attention map, and performs Gaussian distribution fitting on the attention sub-map, and uses the expected value as the horizontal coordinate of the center point of the character in the attention map; determines the position coordinates of each character in the target image according to the horizontal coordinates of the center point of each character in the attention map. No additional single character position annotation is required, and there is no additional training process, which reduces the training cost. Because the attention map is used, the character coordinates obtained are more accurate.

[0107] In the document processing of DSU (Document Structure Understanding), words belonging to different categories or cells may be detected as the same text line because they are too close to each other. By adopting the solution provided in the embodiments of the present application, the coordinates of each character can be known, and then they can be divided into their respective parts according to the layout structure of the document or table.

[0108] The following are the device embodiments of the present application, which can be used to execute the method embodiments for single-word detection in the above images of the present application. For the details not disclosed in the device embodiments of the present application, please refer to the method embodiments for single-word detection in the images of the present application based on the present application.

[0109] Figure 13 It is a block diagram of a device for single-word detection in an image shown in an embodiment of the present application. As Figure 13 shown, the device includes: a text line detection module 810, a text line image acquisition module 820, an attention map acquisition module 830, a center coordinate calculation module 840, and a character coordinate determination module 850.

[0110] The text line detection module 810 is configured to perform text line detection on a target image through a text line detection algorithm to obtain the coordinate information of the text line box;

[0111] The text line image acquisition module 820 is configured to acquire a text line image from the target image according to the coordinate information of the text line box;

[0112] The attention map acquisition module 830 is configured to extract a first feature map of the text line image through a deep convolutional network and calculate the attention map corresponding to the first feature map;

[0113] The center coordinate calculation module 840 is configured to extract the attention sub-map corresponding to each character from the attention map, perform Gaussian distribution fitting on the attention sub-map, and use the expected value as the horizontal coordinate of the center point of the character in the attention map;

[0114] The character coordinate determination module 850 is configured to determine the position coordinates of each character in the target image according to the horizontal coordinate of the center point of each character in the attention map.

[0115] The implementation processes of the functions and roles of each module in the above device are specifically detailed in the implementation processes of the corresponding steps in the above method for single-word detection in an image, and will not be elaborated here.

[0116] In several embodiments provided in this application, the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and a module, a program segment, or a part of code contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0117] In addition, each functional module in various embodiments of this application can be integrated together to form an independent part, or each module can exist alone, or two or more modules can be integrated to form an independent part.

[0118] If the function is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of this application. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs, etc., which can store program codes.

Claims

1. A method for detecting single characters in an image, characterized in that, Including: Performing text line detection on the target image through a text line detection algorithm to obtain the coordinate information of the text line box; Obtaining a text line image from the target image according to the coordinate information of the text line box; Extracting a first feature map of the text line image through a deep convolutional network and calculating an attention map corresponding to the first feature map; Extracting an attention sub-map corresponding to each character from the attention map, performing Gaussian distribution fitting on the attention sub-map, and taking the expected value as the horizontal coordinate of the center point of the character in the attention map; Determining the position coordinates of each character in the target image according to the horizontal coordinate of the center point of each character in the attention map; Among them, the obtaining a text line image from the target image according to the coordinate information of the text line box includes: Calculating the coordinate information of a positive rectangular box according to the coordinate information of the text line box; Determining a perspective transformation matrix according to the coordinate information of the text line box and the coordinate information of the positive rectangular box; Calculating the pixel coordinates mapped by each pixel coordinate in the positive rectangular box in the target image according to the perspective transformation matrix; Calculating the pixel value of each pixel coordinate in the positive rectangular box by using bilinear interpolation according to the pixel value of the pixel coordinate in the target image to obtain the text line image; Among them, the extracting an attention sub-map corresponding to each character from the attention map, performing Gaussian distribution fitting on the attention sub-map, and taking the expected value as the horizontal coordinate of the center point of the character in the attention map includes: Extracting the attention values of each row from the attention map to obtain the attention sub-map of the character corresponding to the row; For each character, performing Gaussian distribution fitting on the attention sub-map of the character with the horizontal coordinate of the attention map as the independent variable and the attention value as the dependent variable to obtain the expected value of the independent variable; Taking the expected value as the horizontal coordinate of the center point of the character in the attention map.

2. The method according to claim 1, characterized in that, The extracting a first feature map of the text line image through a deep convolutional network and calculating an attention map corresponding to the first feature map includes: Extracting a first feature map of the text line image through a deep convolutional network and connecting the first feature map end to end in the vertical dimension to obtain a second feature map; Multiplying the second feature map by the transpose of the second feature map to obtain the attention map.

3. The method according to claim 1, wherein The determining the position coordinates of each character in the target image according to the horizontal coordinate of the center point of each character in the attention map includes: Calculating the segmentation point coordinates between every two adjacent characters in the attention map according to the horizontal coordinate of the center point of each character in the attention map; Inverse calculating the segmentation coordinates between adjacent characters on the text line image according to the segmentation point coordinates between every two adjacent characters in the attention map by using bilinear interpolation; Determining the position coordinates of each character in the target image according to the segmentation coordinates between adjacent characters on the text line image.

4. The method according to claim 3, wherein According to the segmentation point coordinates between every two adjacent characters in the attention map, inversely calculating the segmentation coordinates between adjacent characters on the text line image by using bilinear interpolation, including: According to the width and height of the attention sub-map and the width and height of the text line image, calculating the segmentation coordinates between adjacent characters on the text line image corresponding to the segmentation point coordinates between every two adjacent characters in the attention map by using bilinear interpolation.

5. The method according to claim 3, characterized in that, The determining the position coordinates of each character in the target image according to the segmentation coordinates between adjacent characters on the text line image includes: According to the segmentation coordinates between adjacent characters on the text line image, mapping to obtain the boundary coordinates of adjacent characters in the target image through the inverse matrix of the perspective transformation matrix; According to the boundary coordinates of adjacent characters in the target image, determining the position coordinates of each character in the target image.

6. An apparatus for detecting single characters in an image, characterized in that, Including: A text line detection module, configured to perform text line detection on the target image through a text line detection algorithm to obtain the coordinate information of the text line box; A text line image obtaining module, configured to obtain a text line image from the target image according to the coordinate information of the text line box; Wherein, the obtaining the text line image from the target image according to the coordinate information of the text line box includes: According to the coordinate information of the text line box, calculating the coordinate information of the positive rectangular box; According to the coordinate information of the text line box and the coordinate information of the positive rectangular box, determining a perspective transformation matrix; According to the perspective transformation matrix, calculating the pixel coordinates mapped by each pixel coordinate in the positive rectangular box in the target image; According to the pixel values of the pixel coordinates in the target image, calculating the pixel values of each pixel coordinate in the positive rectangular box by using bilinear interpolation to obtain the text line image; An attention map obtaining module, configured to extract a first feature map of the text line image through a deep convolutional network and calculate an attention map corresponding to the first feature map; A center coordinate calculation module, configured to extract an attention sub-map corresponding to each character from the attention map, perform Gaussian distribution fitting on the attention sub-map, and use the expected value as the horizontal coordinate of the center point of the character in the attention map; Wherein, the extracting an attention sub-map corresponding to each character from the attention map, performing Gaussian distribution fitting on the attention sub-map, and using the expected value as the horizontal coordinate of the center point of the character in the attention map includes: Extracting the attention values of each row from the attention map to obtain the attention sub-map of the characters corresponding to the row; For each character, performing Gaussian distribution fitting on the attention sub-map of the character with the horizontal coordinate of the attention map as the independent variable and the attention value as the dependent variable to obtain the expected value of the independent variable; Using the expected value as the horizontal coordinate of the center point of the character in the attention map; A character coordinate determination module, configured to determine the position coordinates of each character in the target image according to the horizontal coordinate of the center point of each character in the attention map.

7. An electronic device, characterized in that, The electronic device includes: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to execute the method for detecting single characters in an image according to any one of claims 1-5.

8. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, and the computer program can be executed by the processor to complete the method for detecting single characters in an image according to any one of claims 1-5.

Citation Information

Patent Citations

  • Text recognition method and device, electronic equipment and storage medium

    CN112464798A

  • Character recognition network model training method, character recognition method, apparatuses, terminal, and computer storage medium therefor

    WO2021115159A1