Text Area Determination Method, Apparatus, Storage Medium, and Electronic Device
By determining the rectangular reference area and candidate area in the image, and determining the target text area using geometric relationships and text distribution information, the problems of slow computing speed and insufficient accuracy in the prior art are solved, and efficient text area recognition is achieved.
Patent Information
- Application Number
- CN202210320579.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-29
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2042-03-29
AI Technical Summary
When determining text areas in the prior art, there is a problem of slow calculation speed or insufficient accuracy when determining the text area in an image, and it is impossible to take into account both the accuracy and the calculation speed.
By determining the rectangular reference area in the initial image, the edges of the rectangular reference area are parallel to the reference direction in the initial image, the candidate areas are determined using geometric relationships, and the target text area is determined based on the text distribution information, avoiding image correction and reducing the calculation amount.
While ensuring the accuracy of text area determination, it significantly reduces the calculation amount and improves the calculation speed.
Smart Images

Figure CN114663873B_ABST
Abstract
Description
Background Art
[0002] With the development of image recognition technology, the technology of determining text regions in images to obtain corresponding texts is used more and more widely.
[0003] In the prior art, in the method of determining text regions in images to obtain corresponding texts, there are problems of slow calculation speed or insufficient accuracy, that is, it is impossible to balance accuracy and calculation speed.
[0004] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present disclosure, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0005] The purpose of the present disclosure is to provide a method for determining a text region, a device for determining a text region, a computer-readable medium, and an electronic device, so as to reduce the amount of calculation and improve the calculation speed while ensuring accuracy.
[0006] According to a first aspect of the present disclosure, there is provided a method for determining a text region, including: obtaining an initial image including the text, and determining a rectangular reference region including the text in the initial image, wherein one edge of the rectangular reference region is parallel to a reference direction in the initial image; determining candidate regions according to the rectangular reference region; determining text distribution information of each candidate region in the candidate regions; and determining a target text region in the candidate regions based on the text distribution information.
[0007] According to a second aspect of the present disclosure, there is provided a device for determining a text region, including: a target detection module for obtaining an initial image including the text and determining a rectangular reference region including the text in the initial image, wherein one edge of the rectangular reference region is parallel to a reference direction in the initial image; a first determination module for determining candidate regions according to the rectangular reference region; an information extraction module for determining text distribution information of each candidate region in the candidate regions; and a second determination module for determining a target text region in the candidate regions based on the text distribution information.
[0008] According to a third aspect of the present disclosure, there is provided a computer-readable medium having a computer program stored thereon, and when the computer program is executed by a processor, the above method is implemented.
[0009] According to a fourth aspect of the present disclosure, there is provided an electronic device, including: one or more processors; and a memory for storing one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors implement the above method.
[0010] A method for determining a text region provided by an embodiment of the present disclosure obtains an initial image including text and determines a rectangular reference region including text in the initial image, wherein the edges of the rectangular reference region are parallel to a reference direction in the initial image; determines candidate regions according to the rectangular reference region; determines the text distribution information of each candidate region in the candidate regions; and determines a target text region in the candidate regions based on the text distribution information. Compared with the prior art, on the one hand, a general rectangular reference region is determined in the initial image, and then candidate regions are determined in the rectangular reference region based on geometric relationships, and target detection can be started without image correction, reducing the computational amount. At the same time, the edges of the rectangular reference region are parallel to the reference direction in the initial image, that is, during detection, the angles of the detection frames can be made consistent, reducing the computational amount during the detection process and further reducing the computational amount and improving the computational speed. On the other hand, after determining the candidate regions, the target text region is determined based on the text distribution information of the candidate regions, which can ensure the accuracy of determining the text region. That is, the present disclosure reduces the computational amount and improves the speed of determining the text region while ensuring the accuracy of determining the text region.
[0011] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present disclosure and used together with the specification to explain the principles of the present disclosure. Obviously, the drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts. In the drawings:
[0013] Figure 1 A schematic diagram showing a bank card image without a rotation angle;
[0014] Figure 2 A schematic diagram showing a bank card image with a rotation angle;
[0015] Figure 3 A schematic diagram showing an exemplary system architecture to which the embodiments of the present disclosure can be applied;
[0016] Figure 4 A flowchart schematically showing a method for determining a text region in an exemplary embodiment of the present disclosure;
[0017] Figure 5 A flowchart schematically showing a method for obtaining a rectangular reference region in an exemplary embodiment of the present disclosure;
[0018] Figure 6 Schematic diagram showing the structure of a rectangular reference region in an exemplary embodiment of the present disclosure;
[0019] Figure 7 Schematic diagram showing the geometric relationship between an inscribed rectangle and a rectangular reference region in an exemplary embodiment of the present disclosure;
[0020] Figure 8 Schematic diagram showing another geometric relationship between an inscribed rectangle and a rectangular reference region in an exemplary embodiment of the present disclosure;
[0021] Figure 9 Schematic diagram showing the target image corresponding to the rectangular reference region in an exemplary embodiment of the present disclosure;
[0022] Figure 10 Schematic diagram showing the position of the inscribed rectangle abcd in the initial image in an exemplary embodiment of the present disclosure;
[0023] Figure 11 Schematic diagram showing the target image corresponding to the inscribed rectangle abcd in an exemplary embodiment of the present disclosure;
[0024] Figure 12 Schematic diagram showing the position of the inscribed rectangle efgj in the initial image in an exemplary embodiment of the present disclosure;
[0025] Figure 13 Schematic diagram showing the target image corresponding to the inscribed rectangle efgj in an exemplary embodiment of the present disclosure;
[0026] Figure 14 Schematic diagram showing the position of key points in an exemplary embodiment of the present disclosure;
[0027] Figure 15 Schematic diagram showing the position of the straight line fitted according to the key points in an exemplary embodiment of the present disclosure;
[0028] Figure 16 Schematic diagram showing the position of the minimum bounding rectangle in an exemplary embodiment of the present disclosure;
[0029] Figure 17 Schematic flowchart showing the optimal embodiment of the method for determining the text region of the present disclosure;
[0030] Figure 18 Schematic diagram showing the composition of the text region determination device in an exemplary embodiment of the present disclosure;
[0031] Figure 19 Schematic diagram showing an electronic device to which the embodiments of the present disclosure can be applied. Detailed implementation manners
[0032] Example embodiments will now be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the concept of example embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0033] In addition, the accompanying drawings are only schematic illustrations of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus repetitive descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.
[0034] In the related art, with the popularity of portable electronic devices such as mobile phones and the large-scale construction of communication infrastructures such as 4G and 5G, mobile payment represented by mobile phone payment has become increasingly popular, realizing some features of a "cashless" society to a certain extent.
[0035] The first step of mobile phone payment is for the user to log in to the payment account on the mobile phone and bind their real-name bank card. An important step is for the user to enter their identity information and then enter the bank card number for verification. However, since bank card numbers generally have a large number of digits and their fonts are quite different from regular printed and handwritten characters, the probability of incorrect input by the user is relatively high, resulting in the user having to repeatedly enter and confirm or even have their account temporarily frozen, which greatly affects the experience. Therefore, researching an automatic recognition scheme for bank card numbers based on image recognition has very important application value.
[0036] Currently, the mainstream bank card number region detection schemes are divided into two categories: one is the relatively traditional computer vision methods based on filters, morphological processing, etc., which locate the card number region through edge detection. The detection accuracy of this type of method is poor; the other type of method is based on object detection class CNN (Convolutional Neural Networks), and the model is trained through big data learning to have the ability of autonomous annotation to directly locate the target region. The advantage of this type of method is extremely high performance, but the disadvantage is that the training difficulty is relatively large, and generally only horizontal rectangular frames can be located, so the effect on rotated images is poor. Such as Figure 1 and Figure 2As shown, when the image has a rotation angle, it is necessary to correct the image to ensure the monitoring accuracy, which involves a large amount of calculation. If the image is not corrected, there will be a large number of irrelevant background regions in the horizontal detection box output by the CNN-based object detection scheme, which will affect the performance of the recognition module.
[0037] Figure 3 FIG. shows a schematic diagram of the system architecture. The system architecture 300 may include a terminal 310 and a server 320. Among them, the terminal 310 may be a terminal device such as a smart phone, a tablet computer, a desktop computer, a laptop computer, etc. The server 320 generally refers to a background system that provides relevant services for text area determination in this exemplary embodiment, and may be a single server or a cluster formed by multiple servers. A connection may be formed between the terminal 310 and the server 320 through a wired or wireless communication link for data interaction.
[0038] In one embodiment, the above text area determination method may be executed by the terminal 310. For example, after the user uses the terminal 310 to capture an image or selects an image from the album of the terminal 310, the terminal 310 determines the text area of the image and outputs the target text area.
[0039] In one embodiment, the above text area determination method may be executed by the server 320. For example, after the user uses the terminal 310 to capture an image or selects an image from the album of the terminal 310, the terminal 310 uploads the image to the server 320, and the server 320 determines the text area of the image and returns the target text area to the terminal 310.
[0040] As can be seen from the above, the execution subject of the text area determination method in this exemplary embodiment may be the above terminal 310 or server 320, and the present disclosure does not limit this.
[0041] The following combines Figure 4 to describe the text area determination method in this exemplary embodiment. Figure 4 FIG. shows an exemplary flow of the text area determination method, which may include:
[0042] Step S410, obtain an initial image including the text, and determine a rectangular reference region including the text in the initial image, where one of the edges of the rectangular reference region is parallel to the reference direction in the initial image;
[0043] Step S420, determine candidate regions according to the geometric relationship of the rectangular reference region;
[0044] Step S430, determine the text distribution information of each of the candidate regions in the candidate regions;
[0045] Step S440: Determine a target text region in the candidate region based on the text distribution information.
[0046] Based on the above method, on the one hand, a rectangular reference region is determined in the initial image, and then candidate regions are determined in the rectangular reference region based on geometric relationships. It is not necessary to correct the image to start target detection, reducing the computational load. At the same time, the edges of the rectangular reference region are parallel to the reference direction in the initial image. That is, during detection, the angles of the detection frames can be made consistent, reducing the computational load during the detection process and improving the computational speed. On the other hand, after determining the candidate regions, the target text region is determined based on the text distribution information of the candidate regions, which can ensure the accuracy of determining the text region. That is, the present disclosure reduces the computational load while ensuring the accuracy of determining the text region and improves the speed of determining the text region.
[0047] The following Figure 4 will specifically describe each step.
[0048] Referring to Figure 4 , in step S410, an initial image including the text is obtained, and a rectangular reference region including the text is determined in the initial image, where one of the edges of the rectangular reference region is parallel to the reference direction in the initial image.
[0049] In the present exemplary embodiment, the initial image includes text. The initial image can be an image corresponding to a bank card, an ID card, or other cards, or it can be a document image, and can also be customized according to user needs, and is not specifically limited in the present exemplary embodiment.
[0050] When the above initial image is an image corresponding to a bank card, an ID card, or other cards, the above text may include a bank card number, an ID card number, etc.
[0051] In the present exemplary embodiment, the above initial image can be rectangular, triangular, circular, etc., and the shape of the above initial image is not specifically limited in the present exemplary embodiment. The initial image can be a 3-channel color image or a 1-channel grayscale image, and can also be customized according to user needs, and is not specifically limited in the present exemplary embodiment.
[0052] Among them, the above rectangular reference region is a rectangular region including the above text, and the reference direction can be a direction defined by the user, used to define the deflection angle between the rectangular reference region and the horizontal direction. For example, if the above initial image is rectangular, here, the reference direction can be parallel to one of the edges of the above initial image.
[0053] In the present exemplary embodiment, when obtaining the above-mentioned rectangular reference region, steps S510 to S530 may be included, and steps S510 to S530 will be described in detail below.
[0054] In step S510, object detection is performed on the above-mentioned initial image to obtain a plurality of intermediate text regions.
[0055] In step S520, the accuracy of each of the intermediate text regions and the confidence of each of the intermediate text regions including preset type text are determined.
[0056] Specifically, with reference to Figure 6 As shown, a rectangular coordinate system can be established first, and the above-mentioned initial image is placed in the above-mentioned rectangular coordinate system. At this time, the above-mentioned reference direction can be a direction parallel to any coordinate axis. The above-mentioned initial image can be subjected to object detection with a rectangular detection frame to obtain a plurality of intermediate text regions. At this time, the edges of the above-mentioned rectangular detection frame are parallel to the above-mentioned reference direction, defining the direction of the rectangular detection frame. There is no need to use rectangular detection frames in other directions, which can reduce the number of rectangular detection frames and reduce the computational complexity in the detection process.
[0057] Below, taking the above-mentioned initial image as an image including a bank card and the preset text type as the bank card number as an example for illustration, with reference to Figure 6 As shown, the processor can output n 6-dimensional vectors after subjecting the above-mentioned initial image to a plurality of cascaded convolution, pooling, and other operations. Among them, n is related to the specific network structure, and n can be any positive integer greater than or equal to 1000 and less than or equal to 99999, such as 10000, 20000, etc., or can be customized according to requirements, and is not specifically limited in the present exemplary embodiment. The form of a single prediction vector is (x, y, w, h, c1, c2), where x, y, w, h are positive numbers, representing the abscissa of the center point of the intermediate text region, the ordinate of the center point, the width of the intermediate text region, and the height of the intermediate text region respectively, and then a plurality of intermediate text regions are obtained. Each 6-dimensional vector represents an intermediate text region.
[0058] In the present exemplary embodiment, where c1 and c2 are positive numbers between 0 and 1, representing the confidence of the accuracy of the center point coordinates and the confidence that the region within the frame is the bank card number respectively.
[0059] In the present exemplary embodiment, the product of the above-mentioned c1 and c2 can be used as the confidence of the intermediate text region including the preset type text.
[0060] In step S530, the rectangular reference region is determined from the plurality of intermediate text regions according to the accuracy and the confidence.
[0061] In this exemplary embodiment, after obtaining n intermediate text regions, the one with the largest product of c1 and c2 is taken, that is, the intermediate text region with the highest confidence including the preset type of text among the above intermediate text regions is used as the above rectangular reference region. As Figure 6 shown, the prediction result is denoted as (X, Y, W, H, C1, C2), and the rectangular reference region is rectangle ABCD.
[0062] In an exemplary embodiment of the present disclosure, a preset threshold can be first set. The preset threshold can be 0.5, or 0.4, 0.6, etc., and can also be customized according to user requirements, and is not specifically limited in this exemplary embodiment.
[0063] In this exemplary embodiment, taking the preset threshold as 0.5 as an example, the upper left corner of the initial image is taken as the coordinate origin, the upper left corner of the image is taken as the coordinate origin, the AB direction is the positive X-axis direction, and the AD direction is the positive Y-axis direction for illustration. If C1×C2 is less than the preset threshold, it means that the initial image does not contain the card number area, and subsequent operations can be skipped. If C1×C2 is greater than or equal to the above preset threshold, the coordinates of the center point P of the above rectangular reference region ABCD are determined as (X, Y), and the coordinates of points A, B, C, and D are (X - W / 2, Y - H / 2), (X + W / 2, Y - H / 2), (X + W / 2, Y + H / 2), and (X - W / 2, Y + H / 2) respectively. Where W represents the length of the long side of the rectangular reference region, and H represents the length of the short side of the rectangular reference region, that is, the length in the present disclosure represents the length of the long side, and the width represents the length of the short side. Determining the rectangular reference region in the above manner can reduce the computational complexity of the target detection algorithm while ensuring accuracy.
[0064] After obtaining the above rectangular reference region, step S420 can be executed.
[0065] In step S420, candidate regions are determined according to the rectangular reference region.
[0066] In this exemplary embodiment, the above candidate regions are rectangular regions that may include all texts in the rectangular reference region and have an area less than or equal to the rectangular reference region. Specifically, the candidate regions can be determined according to the geometric relationship of the rectangular reference region. The geometric relationship can include the length, width, aspect ratio, inscribed image, circumscribed image, etc. of the rectangular reference region, and is not specifically limited in this exemplary embodiment.
[0067] In one embodiment, after determining the above rectangular reference region, first determine the inscribed rectangle of the above rectangular reference region. The inscribed rectangle of the above rectangular reference region and the above rectangular reference region itself can be used as the above candidate regions, that is, three candidate regions can be included, specifically including the rectangular reference region itself and two inscribed rectangles of the rectangular reference region.
[0068] In one embodiment, if the W and H of the above rectangular reference region are not equal (that is, the rectangular reference region is not a square), then only two inscribed rectangles are included in the rectangular reference region. It should be noted that the four corners of the inscribed rectangle in this exemplary embodiment are respectively located on the four sides of the rectangular reference region. Refer to Figure 7 and Figure 8 As shown, the inscribed rectangles of the rectangular reference region ABCD can include abcd and efgj. Among them, the sizes of the above two inscribed rectangles are equal and the deflection directions are opposite. At this time, the coordinates of each point of abcd and efgj can be directly obtained by computer fitting to obtain the above two inscribed rectangles, and the inscribed rectangle abcd, the inscribed rectangle efgj, and the rectangular reference region itself are used as the above candidate regions. Taking the above inscribed rectangle as the above candidate region can, on the one hand, accurately locate the position of the text, and on the other hand, the number of inscribed rectangles is fixed, which can directly reduce the number of candidate regions, thereby reducing the calculation amount.
[0069] In an exemplary embodiment of the present disclosure, refer to Figure 7 As shown, the deflection angle ∠α of the inscribed rectangle relative to the rectangular reference region can be first determined, where there is a one-to-one correspondence between the above deflection angle ∠α and the above ∠BDC. For example, when the above ∠BDC = 32°, ∠α = 22°; when the above ∠BDC = 20°, ∠α = 15°. The corresponding relationship between ∠BDC and the above deflection angle in the above different rectangular reference regions can be first determined. After obtaining the above ∠BDC, the above deflection angle can be determined according to the above corresponding relationship, and then the coordinate values of each point of the above inscribed rectangle can be calculated based on the deflection angle.
[0070] In this exemplary embodiment, the corresponding relationship between the above ∠BDC and the deflection angle can be accurate to 1 degree, or can be accurate to 0.5 degree, 0.1 degree, etc., and no specific limitation is made in this exemplary embodiment.
[0071] In an exemplary embodiment, the preset aspect ratio of the above text can also be first determined, and then the above inscribed rectangle is determined based on the above preset aspect ratio and the above deflection angle. Specifically, taking the inscribed rectangle abcd as an example for illustration, let the long side of the inscribed rectangle abcd be w and the short side be h. In the present disclosure, w represents the length of the long side and h represents the length of the short side. Assuming the deflection angle is ∠α and ∠α is m times of ∠BDC, the following geometric logic relationship can be obtained at this time:
[0072] ∠α=m∠BDC
[0073]
[0074] hsinα+wcosα=W
[0075] At the same time, the preset aspect ratio of the above text can be the preset aspect ratio of the above text area, for example, w=kh, where k represents the above preset aspect ratio. At this time, the values of h and w can be solved by the simultaneous equations. Then the coordinates of the above four points abcd are calculated. For example, assuming that the coordinates of point A are (x, y), then the coordinates of point a are (x, y+Hh*cosα), the coordinates of point b are (x+Wh*sinα, y), the coordinates of point c are (x+W, y+h*cosα), and the coordinates of point d are (x+h*sinα, y+H).
[0076] In an exemplary embodiment of the present disclosure, if W and H of the rectangular reference area are equal (that is, the rectangular reference area is a square), then the inscribed rectangle can be determined based on the preset aspect ratio and the deflection angle, and then the candidate area can be determined. When the rectangular reference area is a square, ∠α=∠BDC=45°, then Then based on w=kh, we can calculate the values of h and w, and then calculate the coordinates of the four points abcd. For example, assuming that the coordinates of point A are (x, y), then the coordinates of point a are The coordinates of point b are The coordinates of point c are The coordinates of point d are
[0077] In yet another example implementation, when determining the inscribed rectangle based on a preset aspect ratio and a deflection angle, if the initial image corresponds to a bank card, then it can be approximately determined that ∠α=∠BDC.
[0078] At this point, we can get the following geometric logical relationship:
[0079] sinα=H / (W 2 +H 2 )1 / 2
[0080] cosα=W / (W 2 +H 2 )1 / 2
[0081] hsinα+wcosα=W
[0082] Meanwhile, substituting \(w = kh\) into the above geometric logical relationship, the values of \(h\) and \(w\) can be solved by simultaneously establishing equations. Furthermore, the coordinates of the above four points \(a\), \(b\), \(c\), and \(d\) can be calculated. For example, assuming the coordinates of point \(A\) are \((x,y)\), then the coordinates of point \(a\) are \((x, y + H - h\cos\alpha)\), the coordinates of point \(b\) are \((x + W - h\sin\alpha, y)\), the coordinates of point \(c\) are \((x + W, y + h\cos\alpha)\), and the coordinates of point \(d\) are \((x + h\sin\alpha, y + H)\).
[0083] In the present exemplary embodiment, the above preset aspect ratio is determined according to the initial image and the text in the initial image. For example, if the above initial image is a bank card image and the preset type of text is a bank card number, the above preset aspect ratio \(k\) can be taken as 14. In the present exemplary embodiment, no specific limitation is imposed on the above preset aspect ratio.
[0084] Based on the calculation process of the inscribed rectangle \(efgj\), the calculation process of the inscribed rectangle \(abcd\) can be referred to, and details are not described herein again.
[0085] After obtaining the two inscribed rectangles, the above inscribed rectangles and the rectangle reference region itself are used as the above candidate regions, and then step S430 is executed.
[0086] In step S430, the text distribution information of each of the candidate regions is determined in the candidate regions.
[0087] The text distribution information is information used to represent the above text distribution position, and may include a text distribution coefficient, the distance between each character in the text, etc. No specific limitation is imposed in the present exemplary embodiment.
[0088] In the present exemplary embodiment, referring to Figure 9 and Figure 10 As shown, after obtaining the above multiple candidate regions, the target image corresponding to each candidate region can be obtained, that is, the image of the candidate region is extracted from the above initial image, and the edge of its candidate region is also parallel to the above reference direction. Specifically, referring to Figure 10 As shown, taking the inscribed rectangle \(adcd\) as an example for illustration, it can be assumed that the four points corresponding to the candidate region in the original image are \(P1\), \(P2\), \(P3\), and \(P4\) (where \(P1\) is the point closest to the upper left corner of the image, and \(P1\), \(P2\), \(P3\), \(P4\) are arranged in a clockwise order), their coordinate set is \(PS=\{(x1,y1),(x2,y2),(x3,y3),(x4,y5)\}\), the width and height of the candidate region are \(w1\) and \(h1\) respectively, then the affine transformation matrix from \(PS\) to the point set \(PD =\{(0,0),(w1,0),(w1,h),(0,h1)\}\) is solved, and the corresponding affine transformation is performed on the original image. Figure 9 The target image obtained when the rectangle reference region is used as the candidate region is shown. Referring toFigure 10 and Figure 11 As shown, when the inscribed rectangle abcd is used as the candidate region, the target image is determined. Refer to Figure 12 and Figure 13 , which is the target image determined when the inscribed rectangle efgj is used as the candidate region.
[0089] At this time, the rotation angle of the text in the above target image can be determined. Specifically, refer to Figure 14 , Figure 15 and Figure 16 As shown, taking the candidate image as the rectangle reference region itself as an example, the processor can use the Shi Tomasi algorithm on the above target image to perform corner detection on the candidate region to obtain multiple key points, and then obtain the minimum circumscribed rectangle corresponding to the text according to the above key points. The included angle between the minimum circumscribed rectangle S and the above candidate region is used as the above rotation angle, and the above rotation angle is denoted as V. In the present exemplary embodiment, the maximum number of points for corner detection can be set to 100, the quality level is 0.005, and the minimum distance is 2, which can make the corner points concentrated in the text area. The specific parameters of corner detection can also be customized according to user needs and are not specifically limited in the present exemplary embodiment.
[0090] In the present exemplary embodiment, the above text distribution parameter includes a text distribution coefficient. Refer to Figure 15 As shown, after obtaining multiple key points, the least squares method can be used to perform linear fitting on the above key points, and the slope K of the determined straight line can be obtained. Then, the above text distribution coefficient R can be calculated according to the above slope and rotation angle. Specifically, R = VK.
[0091] After calculating the above text distribution information of the above multiple candidate regions respectively, step S440 can be executed.
[0092] In step S440, a target text region is determined in the candidate region based on the text distribution information.
[0093] To determine the target text region according to the text distribution information, the rotation angle of the text can be determined through the text distribution information, and the candidate region in the target image where the text is closest to horizontal can be used as the target text region, or the area ratio occupied by the text in the candidate region can be determined through the text distribution information, and the one closest to being full can be used as the target text region.
[0094] In an exemplary real-time manner of the present disclosure, the candidate region with the smallest above text distribution coefficient can be used as the above target text region. The text in the candidate region with the smallest above text distribution coefficient is closest to horizontal, that is, the accuracy of the obtained target text region is the highest. Using the above method can improve the accuracy of determining the above target text region.
[0095] In another exemplary embodiment of the present disclosure, the processor may perform the following operations for each candidate region. First, divide the target image corresponding to the candidate region into multiple sub-regions, then determine the text density information of each of the above sub-regions according to the above text distribution coefficient, and then calculate the standard deviation of the text density information of each sub-region in each of the candidate regions.
[0096] After obtaining the standard deviations of multiple above-mentioned candidate regions, the candidate region corresponding to the target image with the smallest standard deviation can be used as the above-mentioned target text region. The smaller the standard deviation of the text density of each region in the candidate region, the more uniform the text distribution. Selecting the candidate region with the most uniform text distribution as the target text region can improve the accuracy of the determined target text region.
[0097] In still another exemplary embodiment of the present disclosure, the processor may also directly use OCR (Optical Character Recognition) to determine the above-mentioned target text region in the above-mentioned candidate region.
[0098] It should be noted that there are various ways to determine the target text region in the candidate region. The above is an exemplary description, and the present disclosure does not specifically limit it.
[0099] Further, referring to Figure 17 As shown, a specific exemplary embodiment is used to illustrate the above method for determining the text region. First, step S1710 may be executed to obtain an initial image including the text and determine a rectangular reference region including the text in the initial image; then step S1720 is executed, and the rectangular reference region and the inscribed rectangle are used as the candidate regions. Specifically, when executing step S1720, step S1721 may be first executed to determine whether the above rectangular reference region is a rectangle or a square. If it is a rectangle, step S1722 is executed to determine the deflection angle according to the geometric relationship, and step S1723 is executed to determine the inscribed rectangle in the rectangular reference region based on the deflection angle. If the above rectangular reference region is a square, step S1724 is executed to obtain the preset length-width ratio of the text and determine the deflection angle according to the geometric relationship; and step S1725 is executed to determine the inscribed rectangle in the rectangular reference region based on the preset length-width ratio and the deflection angle.
[0100] After obtaining the above-mentioned inscribed rectangle, the text distribution coefficient in the text distribution information can be determined, and step S1730 can be executed to obtain the target image corresponding to each of the candidate regions; then step S1740 is executed to perform corner detection on each of the target images to obtain a plurality of key points; then step S1750 is executed to determine the rotation angle of the text in the target image; and step S1760, perform line fitting based on the key points and determine the slope of the line; finally, execute step S1770 to determine the text distribution coefficient based on the rotation angle and the slope.
[0101] After obtaining the above-mentioned text distribution coefficient, step S1780 can be executed to determine the candidate region corresponding to the target image with the smallest absolute value of the text distribution coefficient as the target text region. To complete the determination of the target text region.
[0102] In summary, compared with the prior art, on the one hand, a rough rectangular reference region is determined in the initial image, and then candidate regions are determined in the rectangular reference region based on geometric relationships. It is possible to start target detection without image correction, reducing the computational load. At the same time, the edges of the rectangular reference region are parallel to the reference direction in the initial image, that is, during detection, the angles of the detection frames can be made consistent, reducing the computational load during the detection process and further reducing the computational load, thereby improving the computational speed. On the other hand, the inscribed rectangle of the rectangular reference region and itself are used as candidate regions, and the coordinate values of each vertex of the candidate region are calculated through relevant geometric relationships. Without the need for the computer to perform relatively complex operations while ensuring accuracy, the computational load is further reduced. On the further hand, after determining the candidate regions, the target text region is determined based on the text distribution information of the candidate regions. Specifically, the text distribution coefficient in the text distribution information is obtained through the rotation angle of the text and the key point fitting and line slope, and the text distribution coefficient is used to determine the target text region, which can ensure the accuracy of determining the text region. That is, the present disclosure reduces the computational load while ensuring the accuracy of determining the text region, and improves the speed of determining the text region.
[0103] It should be noted that the above-mentioned drawings are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present disclosure, rather than for limiting purposes. It is easy to understand that the processes shown in the above-mentioned drawings do not indicate or limit the time sequence of these processes. In addition, it is also easy to understand that these processes can be executed synchronously or asynchronously in, for example, multiple modules.
[0104] Further, referring to Figure 18As shown, in the embodiment of this example, a text area determination device 1800 is further provided, including a target detection module 1810, a first determination module 1820, an information extraction module 1830, and an image generation module 1840. Among them:
[0105] The target detection module 1810 can be used to obtain an initial image including text, and determine a rectangular reference area including text in the initial image. Among them, one edge of the rectangular reference area is parallel to the reference direction in the initial image. Specifically, target detection is performed on the initial image using a rectangular detection frame to obtain multiple intermediate text areas; determine the accuracy of each intermediate text area and the confidence level of the preset type of text included in each intermediate text area; determine the rectangular reference area from the multiple intermediate text areas according to the accuracy and the confidence level.
[0106] The first determination module 1820 can be used to determine a candidate area according to the rectangular reference area. Specifically, the inscribed rectangle of the rectangular reference area can be determined first; the rectangular reference area is used as the candidate area, and the inscribed rectangle is used as the candidate area.
[0107] In an example embodiment, when determining the inscribed rectangle of the rectangular reference area, the first determination module 1820 can first determine the deflection angle of the inscribed rectangle relative to the rectangular reference area according to the geometric relationship, and then determine the inscribed rectangle in the rectangular reference area based on the deflection angle.
[0108] In the embodiment of this example, when determining the inscribed rectangle in the rectangular reference area based on the deflection angle, the first determination module 1820 can first obtain the preset aspect ratio of the text, and then determine the inscribed rectangle in the rectangular reference area based on the preset aspect ratio and the deflection angle.
[0109] The information extraction module 1830 can be used to determine the text distribution information of each candidate area in the candidate area. Specifically, the text distribution information includes a text distribution coefficient. When performing the operation of determining the text distribution information of each candidate area in the candidate area, the information extraction module 1830 can first obtain the target image corresponding to each candidate area; then determine the rotation angle of the text in the target image relative to the candidate area; finally, determine the text distribution coefficient according to the rotation angle.
[0110] In the embodiment of this example, when determining the text distribution coefficient according to the rotation angle, the information extraction module 1830 can first perform corner detection on each target image to obtain multiple key points; then, perform line fitting based on the key points and determine the slope of the line; finally, determine the text distribution coefficient based on the rotation angle and the slope.
[0111] The image generation module 1840 can be used to determine the target text area in the candidate area based on the text distribution information.
[0112] In one exemplary embodiment, the image generation module 1840 is configured to determine the candidate region corresponding to the target image with the minimum absolute value of the text distribution coefficient as the target text region.
[0113] In another exemplary embodiment, the image generation module 1840 may be configured to first obtain the target images corresponding to the candidate regions; then divide each target image into multiple sub-regions; secondly, determine the text density information in each sub-region according to the text distribution information; thereafter, calculate the standard deviation of the text density information of each sub-region in each candidate region; and finally determine the candidate region corresponding to the target image with the minimum standard deviation as the target text region.
[0114] The specific details of each module in the above device have been described in detail in the embodiments of the method part. For the details not disclosed, reference may be made to the embodiments of the method part, and thus will not be elaborated here.
[0115] The exemplary embodiments of the present disclosure further provide an electronic device for executing the above text region determination method. The electronic device may be the above terminal 310 or server 320. Generally, the electronic device may include a processor and a memory. The memory is used to store executable instructions of the processor, and the processor is configured to execute the above image text region determination method by executing the executable instructions.
[0116] Next, taking Figure 19 the mobile terminal 1900 in Figure 19 as an example, the structure of the electronic device will be described exemplarily. Those skilled in the art should understand that, except for the components specifically for mobile purposes,
[0117] As Figure 19 shown, the mobile terminal 1900 may specifically include: a processor 1901, a memory 1902, a bus 1903, a mobile communication module 1904, an antenna 1, a wireless communication module 1905, an antenna 2, a display screen 1906, a camera module 1907, an audio module 1908, a power module 1909, and a sensor module 1910.
[0118] The processor 1901 may include one or more processing units. For example, the processor 210 may include an AP (Application Processor), a modem processor, a GPU (Graphics Processing Unit), an ISP (Image Signal Processor), a controller, an encoder, a decoder, a DSP (Digital Signal Processor), a baseband processor, and / or an NPU (Neural-Network Processing Unit), etc. The method for determining the text area in this exemplary embodiment may be executed by the AP, the GPU, or the DSP. When the method involves neural network-related processing, it may be executed by the NPU.
[0119] The processor 1901 may form a connection with the memory 1902 or other components through the bus 1903.
[0120] The memory 1902 may be used to store computer-executable program code, and the executable program code includes instructions. The processor 1901 executes various functional applications and data processing of the mobile terminal 1900 by running the instructions stored in the memory 1902. The memory 1902 may also store application data, such as storing files such as images and videos.
[0121] The communication function of the mobile terminal 1900 may be implemented through a mobile communication module 1904, an antenna 1, a wireless communication module 1905, an antenna 2, a modem processor, and a baseband processor, etc. The antenna 1 and the antenna 2 are used to transmit and receive electromagnetic wave signals. The mobile communication module 204 may provide 2G, 3G, 4G, 5G, etc. mobile communication solutions applied to the mobile terminal 1900. The wireless communication module 1905 may provide wireless communication solutions such as wireless local area network, Bluetooth, and near-field communication applied to the mobile terminal 1900.
[0122] The sensor module 1910 may include a depth sensor 19101, a pressure sensor 19102, a gyroscope sensor 19103, a barometric pressure sensor 19104, etc., to implement corresponding sensing and detection functions.
[0123] Those skilled in the art can understand that various aspects of the present disclosure can be implemented as a system, a method, or a program product. Therefore, various aspects of the present disclosure can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuit", "module", or "system" here.
[0124] Exemplary embodiments of the present disclosure also provide a computer-readable storage medium, on which a program product capable of implementing the above methods in this specification is stored. In some possible embodiments, various aspects of the present disclosure may also be implemented in the form of a program product, which includes program code. When the program product runs on a terminal device, the program code is used to cause the terminal device to execute the steps according to various exemplary embodiments of the present disclosure described in the "Exemplary Method" section above in this specification.
[0125] It should be noted that the computer-readable medium shown in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0126] In the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0127] In addition, program code for performing the operations of the present disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computing device, partly on the user's device, as a stand-alone software package, partly on the user's computing device and partly on a remote computing device, or entirely on the remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user's computing device through any kind of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., by using an Internet service provider to connect through the Internet).
[0128] Other embodiments of the present disclosure will be readily apparent to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of the present disclosure are pointed out by the claims.
[0129] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the figures, and various modifications and changes may be made without departing from its scope. The scope of the present disclosure is limited only by the appended claims.
Claims
1. A method for determining a text area, characterized in that including: obtaining an initial image including text, and determining a rectangular reference region including the text in the initial image, wherein one edge of the rectangular reference region is parallel to a reference direction in the initial image; determining a candidate region according to the rectangular reference region; determining text distribution information of each candidate region in the candidate region, wherein the text distribution information is information for representing text distribution positions; determining a target text region based on the text distribution information; wherein, the determining the candidate region according to the rectangular reference region includes: determining an inscribed rectangle of the rectangular reference region; taking the rectangular reference region as the candidate region, and taking the inscribed rectangle as the candidate region.
2. The method according to claim 1, characterized in that, The determining the rectangular reference region including the text in the initial image includes: performing object detection on the initial image to obtain a plurality of intermediate text regions; determining the accuracy of each intermediate text region and the confidence of each intermediate text region including preset type text; determining the rectangular reference region from the plurality of intermediate text regions according to the accuracy and the confidence.
3. The method according to claim 1, wherein the determining the inscribed rectangle of the rectangular reference region includes: determining a deflection angle of the inscribed rectangle relative to the rectangular reference region according to geometric relationships, the geometric relationships including the aspect ratio of the rectangular reference region; determining the inscribed rectangle in the rectangular reference region based on the deflection angle.
4. The method according to claim 3, wherein the determining the inscribed rectangle in the rectangular reference region based on the deflection angle includes: obtaining a preset aspect ratio of the text; determining the inscribed rectangle in the rectangular reference region based on the preset aspect ratio and the deflection angle.
5. The method according to claim 1, characterized in that The text distribution information includes a text distribution coefficient, and the determining the text distribution information of each candidate region in the candidate region includes: obtaining a target image corresponding to each candidate region; determining a rotation angle of the text in the target image relative to the candidate region; determining the text distribution coefficient according to the rotation angle.
6. The method according to claim 5, characterized in that, The determining the text distribution coefficient according to the rotation angle includes: performing corner detection on each target image to obtain a plurality of key points; performing straight line fitting based on the key points, and determining the slope of the straight line; determining the text distribution coefficient based on the rotation angle and the slope.
7. The method according to claim 5, characterized in that The determining the target text region based on the text distribution information includes: determining the candidate region corresponding to the target image with the smallest absolute value of the text distribution coefficient as the target text region.
8. The method according to claim 1, wherein The determining the target text region based on the text distribution information includes: obtaining a target image corresponding to each candidate region; dividing each target image into a plurality of sub-regions; determining text density information in each sub-region according to the text distribution information; calculating the standard deviation of the text density information of each candidate region; determining the candidate region corresponding to the target image with the smallest standard deviation as the target text region.
9. A text area determination device, characterized in that including: A target detection module, configured to obtain an initial image including text and determine a rectangular reference region including the text in the initial image, wherein one edge of the rectangular reference region is parallel to a reference direction in the initial image; A first determination module, configured to determine candidate regions according to the rectangular reference region; An information extraction module, configured to determine text distribution information of each of the candidate regions in the candidate regions, where the text distribution information is information used to represent text distribution positions; A second determination module, configured to determine a target text region in the candidate regions based on the text distribution information; Wherein, determining candidate regions according to the rectangular reference region includes: determining an inscribed rectangle of the rectangular reference region; using the rectangular reference region as the candidate region and using the inscribed rectangle as the candidate region.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the text region determination method according to any one of claims 1 to 8.
11. An electronic device, characterized in that, Including: One or more processors; And A memory, configured to store one or more programs, and when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the text region determination method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Text positioning method and device for frame fine tuning, computer device and storage medium
CN109977949A
Method and device for positioning text area in image
JP2013016168A