Text Detection Method, Apparatus, Electronic Device and Medium
By segmenting the target image into multiple sub-maps and merging its probability maps, the problem of difficult to detect abnormal aspect ratio images in the prior art is solved, and a high-accurate text detection effect is achieved.
Patent Information
- Application Number
- CN202210308940.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-25
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2042-03-25
AI Technical Summary
The existing text detection method based on deep learning is difficult to apply to images with abnormal aspect ratios such as long screenshots, resulting in poorer text area detection effect.
By segmenting the target image, multiple target sub-maps are obtained, and these sub-maps are input into the text detection model to obtain their respective probability maps. These probability maps are then merged to generate a probability map of the target image, and finally determine the text area based on the probability map.
Accurate text detection of abnormal aspect ratio images such as long screenshots is realized, improving the accuracy and user experience of text detection results.
Smart Images

Figure CN114782463B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular, to a text detection method, apparatus, electronic device, and medium. Background Art
[0002] With the continuous development of office software, people's requirements for office software processing are also getting higher and higher. They hope to meet different application scenarios while meeting normal office requirements.
[0003] Currently, in practical applications, users may only obtain images containing text, such as images containing text searched from the Internet, and cannot obtain editable tables. At this time, it is necessary to perform text detection on the image to determine the text area. After determining the text area, the text within the text area can be recognized subsequently, and finally data with the original editable content can be generated.
[0004] In the prior art, one method for determining the text area is a text detection method based on deep learning. Due to the distribution of the training set, most text detection methods based on deep learning can only support input images with square resolutions or input images with an A4 ratio, and cannot be directly applied to images with abnormal aspect ratios such as long screenshots.
[0005] To make the text detection method based on deep learning applicable to images with abnormal aspect ratios such as long screenshots, one method is to first scale the image with an abnormal aspect ratio, and then use the text detection method based on deep learning to process the scaled image. However, the image scaling operation is likely to cause the text in the image to be deformed or the resolution to be too low, resulting in a poor effect of text area detection. Another method is to first crop the image with an abnormal aspect ratio into multiple square images, then perform text detection on each square image through the text detection method based on deep learning, and then synthesize the text detection results of the entire image through brute-force stitching. This processing method will result in an unsatisfactory detection result at the seam due to cropping the text area, and the user experience is poor. Summary of the Invention
[0006] Based on the problems existing in the prior art, the present invention proposes a text detection method, apparatus, electronic device, and medium, which can meet the advantages of accurately determining the position information of the text box, improving the accuracy of the text detection result, and enhancing the user experience.
[0007] In a first aspect, the present invention provides a text detection method, including:
[0008] Segmenting a target image to be detected to obtain a plurality of target sub-images; wherein, there is an overlapping area between adjacent target sub-images among the plurality of target sub-images;
[0009] Input the multiple target sub - images into a text detection model to obtain a probability map corresponding to the multiple target sub - images;
[0010] Merge the probability maps corresponding to the multiple target sub - images to obtain a probability map of the target image;
[0011] Determine the text region in the target image according to the probability map of the target image;
[0012] Wherein, the probability map is used to describe the probability value that the content of the pixel in the image corresponding to the probability map is text.
[0013] Further, according to the text detection method provided by the present invention, the step of segmenting the target image to be detected to obtain multiple target sub - images includes:
[0014] Determine the size of the segmented image according to the input resolution of the pre - trained text detection model;
[0015] Segment the target image to be detected according to the size of the segmented image to obtain multiple target sub - images.
[0016] Further, according to the text detection method provided by the present invention, the step of segmenting the target image to be detected according to the size of the segmented image to obtain multiple target sub - images includes:
[0017] Determine the size of the segmentation window according to the size of the segmented image;
[0018] Move the segmentation window on the target image based on a preset overlap ratio to segment the target image and obtain multiple target sub - images.
[0019] Further, according to the text detection method provided by the present invention, the step of moving the segmentation window on the target image based on a preset overlap ratio to segment the target image and obtain multiple target sub - images includes:
[0020] Determine the initial position of the segmentation window on the target image;
[0021] Sequentially move the segmentation window on the target image in a first direction based on a preset overlap ratio to segment the target image;
[0022] Judge whether the segmentation window reaches the segmentation termination position of the target image;
[0023] In the case that the segmentation window has not reached the segmentation termination position of the target image, move the segmentation window in the second direction of the target image based on a preset overlapping ratio, and then re - execute the step of sequentially moving the segmentation window in the first direction of the target image based on the preset overlapping ratio to segment the target image;
[0024] In the case that the segmentation window reaches the segmentation termination position of the target image, obtain multiple target sub - images of the target image based on the segmentation result of the target image.
[0025] Further, according to the text detection method provided by the present invention, the step of sequentially moving the segmentation window in the first direction of the target image based on a preset overlapping ratio to segment the target image includes:
[0026] Determine the target number of moving times of the segmentation window according to the length of the target image in the first direction, the length of the segmentation window in the first direction of the target image, and the preset overlapping ratio;
[0027] Sequentially move the segmentation window in the first direction of the target image based on the preset overlapping ratio, and segment the target image through the segmentation window until the number of moving times of the segmentation window reaches the target number of moving times;
[0028] In the case that there are target sub - images in the segmentation result of the target image whose page size does not meet the specified size, adjust the target sub - images whose page size does not meet the specified size to the specified size.
[0029] Further, according to the text detection method provided by the present invention, before moving the segmentation window on the target image based on a preset overlapping ratio to segment the target image, the method further includes:
[0030] Adjust the resolution of the target image according to the size of the segmented image.
[0031] Further, according to the text detection method provided by the present invention, the adjusting the resolution of the target image according to the size of the segmented image includes:
[0032] Calculate the difference value between the length of the first side of the segmented image and the length of the first side of the target image; wherein, the first side of the segmented image is any side of the segmented image, and the first side of the target image is the shorter side of the target image;
[0033] In the case that the difference value is less than a preset threshold, adjust the resolution of the target image proportionally so that the length of the first side of the segmented image is the same as the length of the adjusted first side of the target image.
[0034] Further, according to the text detection method provided by the present invention, the merging of the probability maps corresponding to the multiple target sub-images to obtain the probability map of the target image includes:
[0035] Determine the overlapping regions and non-overlapping regions in the multiple target sub-images;
[0036] According to the probability values of each pixel in the overlapping region in the probability maps corresponding to the multiple target sub-images, calculate the probability value of each pixel in the overlapping region to obtain the probability map corresponding to the overlapping region;
[0037] According to the probability maps corresponding to the multiple target sub-images, obtain the probability map corresponding to the non-overlapping region;
[0038] Merge the probability map corresponding to the non-overlapping region in the multiple target sub-images and the probability map corresponding to the overlapping region in the multiple target sub-images to obtain the probability map of the target image.
[0039] Further, according to the text detection method provided by the present invention, the calculating the probability value of each pixel in the overlapping region according to the probability values of each pixel in the overlapping region in the probability maps corresponding to the multiple target sub-images includes:
[0040] Determine the magnitudes and the number of the multiple probability values corresponding to a first pixel in the probability maps corresponding to the multiple target sub-images; wherein, the first pixel is any pixel in the overlapping region;
[0041] According to the relationship between the multiple probability values corresponding to the first pixel and a preset probability threshold, determine the probability value of the first pixel by means of voting;
[0042] Or,
[0043] Calculate the average value of the multiple probability values, and use the average value as the probability value of the first pixel.
[0044] Further, according to the text detection method provided by the present invention, the determining the text region in the target image according to the probability map of the target image includes:
[0045] According to the probability map of the target image, determine the probability value of each pixel in the target image;
[0046] According to the probability values of each pixel in the target image, determine the coordinate values of the text region in the target image;
[0047] According to the coordinate values of the text region in the target image, determine the text region of the target image.
[0048] In a second aspect, the present invention further provides a text detection device, including:
[0049] A segmentation model for segmenting a target image to be detected to obtain a plurality of target sub-images; wherein, there is an overlapping area between adjacent target sub-images among the plurality of target sub-images;
[0050] An input module for inputting the plurality of target sub-images into a text detection model to obtain probability maps corresponding to the plurality of target sub-images;
[0051] A merging module for merging the probability maps corresponding to the plurality of target sub-images to obtain a probability map of the target image;
[0052] A determination module for determining a text area in the target image according to the probability map of the target image;
[0053] Wherein, the probability map is used to describe the probability value that the content of the pixel in the image corresponding to the probability map is text.
[0054] In a third aspect, the present invention further provides an electronic device, including: a processor, a memory and a bus, wherein,
[0055] The processor and the memory communicate with each other through the bus;
[0056] The memory stores program instructions executable by the processor, and the processor can execute the steps of the text detection method described in any one of the above by calling the program instructions.
[0057] In a fourth aspect, the present invention further provides a non-transitory computer-readable storage medium, and the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions enable the computer to execute the steps of the text detection method described in any one of the above.
[0058] The present invention provides a text detection method, device, electronic device and medium. The method includes: segmenting a target image to be detected to obtain a plurality of target sub-images; wherein, there is an overlapping area between adjacent target sub-images among the plurality of target sub-images; inputting the plurality of target sub-images into a text detection model to obtain probability maps corresponding to the plurality of target sub-images; merging the probability maps corresponding to the plurality of target sub-images to obtain a probability map of the target image; and determining a text area in the target image according to the probability map of the target image. The text detection method provided by the present invention can accurately determine the position information of the text in the target image, improve the accuracy of the file detection result, and enhance the user experience. Description of the Drawings
[0059] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.
[0060] Figure 1 is a schematic flowchart of the text detection method provided by the present invention;
[0061] Figure 2 is an exemplary diagram of a text detection method provided by the present invention;
[0062] Figure 3 is an exemplary diagram of a text detection result provided by the present invention;
[0063] Figure 4 is an exemplary diagram of a text detection method provided by the present invention;
[0064] Figure 5 is a schematic structural diagram of the text detection device provided by the present invention;
[0065] Figure 6 is a schematic structural diagram of the electronic device provided by the present invention. Detailed implementation manners
[0066] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0067] Figure 1 is a schematic flowchart of the text detection method provided by the embodiments of the present invention. As Figure 1 shown, the text detection method provided by the present invention includes the following steps:
[0068] Step 101: Segment the target image to be detected to obtain multiple target sub-images; wherein, there are overlapping regions between adjacent target sub-images among the multiple target sub-images.
[0069] In this embodiment, when the resolution of the target image meets the preset condition, it is necessary to perform segmentation processing on the target image to be detected to obtain multiple target sub-images. Among them, the target sub-image is an image segmented from the target image, and there is an overlapping area between adjacent target sub-images. The preset condition may be that the resolution of the target image is 3 times the resolution of the target sub-image. It should be noted that the size of the overlapping area can be 30% of the overlapping content, that is, the cropping start end of each target sub-image is 30% of the end of the previous target sub-image, or it can be 20% of the overlapping content. The size of the overlapping area can be set between 0% and 100% according to the actual needs of the user, and is not limited here.
[0070] Step 102: Input the multiple target sub-images into a text detection model to obtain probability maps corresponding to the multiple target sub-images.
[0071] In this embodiment, it is necessary to input the multiple target sub-images obtained in step 101 into a pre-trained text detection model to obtain probability maps corresponding to the multiple target sub-images. Among them, the probability map is a map used to describe the probability value that the content of the pixel in the corresponding image is text. For example, it is determined from a certain probability map that the probability value of pixel 1 is a. By comparing the relationship between the probability value a and a preset threshold, it is determined whether the content corresponding to the pixel 1 is text. For example, when the probability value a is greater than or equal to the preset threshold, it is determined that the content corresponding to the pixel 1 is text. It should be noted that the judgment method can be a method of comparing with a preset threshold, or other judgment methods. For specific details, please refer to the following embodiments and will not be introduced in detail here.
[0072] It should be noted that the multiple target sub-images can be input into the pre-trained text detection model separately in a certain order to obtain probability maps corresponding to each target sub-image respectively.
[0073] It should be noted that in this embodiment, the text detection model is trained based on a sample image and the probability map corresponding to the sample image. Input the sample image into the text detection model to be trained to obtain the corresponding probability map after text detection processing. Compare the obtained probability map with the probability map corresponding to the sample image to determine whether they are consistent. When the training success rate meets the preset training stop condition, stop the training to obtain the text detection model.
[0074] It should be noted that the text detection model is a segmentation-based model. The output of the text detection model is a probability map, and each pixel point on the probability map represents the probability that the content of the corresponding position of the target sub-image belongs to text. The text detection model is a relatively general model, such as the DB-NET model, the East model, the Pse-net model, etc. The text detection model in this embodiment is obtained by training the model.
[0075] Step 103: Merge the probability maps corresponding to the multiple target sub-images to obtain the probability map of the target image.
[0076] In this embodiment, it is necessary to merge the probability maps corresponding to the multiple target sub-images obtained after being processed by the text detection model to obtain the complete probability map corresponding to the target image. It should be noted that there are also overlapping regions in the probability maps corresponding to the multiple target sub-images obtained in this embodiment, where the overlapping regions are determined according to the coordinate information of each probability map. For example, each pixel in the target sub-image is associated with the coordinate information of the pixel in the original target image; the region where the pixels with the same coordinate information in the multiple target sub-images are located is determined as the overlapping region of the multiple target sub-images. For the pixels in the overlapping region, the probability value of the pixel can be calculated based on the probability values of the corresponding pixels in different target sub-images with an overlapping relationship, so as to obtain the probability map of the overlapping region, and then the probability map of the overlapping region and the probability map of the non-overlapping region are stitched together, and the stitched probability map is determined as the probability map of the target image. The determination methods of the probability map of the overlapping region and the probability map of the non-overlapping region can be seen in the following embodiments, and will not be introduced in detail here.
[0077] Step 104: Determine the text region in the target image according to the probability map of the target image;
[0078] Among them, the probability map is used to describe the probability value that the content of the pixel in the image corresponding to the probability map is text; the text detection model is trained based on the sample image and the probability map corresponding to the sample image. It should be noted that the determination method of the text region is: when the probability value of a certain pixel exceeds the preset threshold, it is determined that the content represented by the pixel is text content, and then the region where the pixels corresponding to the text content are connected can be determined as the text region. Among them, the connection method can use the dilation algorithm in image processing technology.
[0079] In this embodiment, it is necessary to determine the text region in the target image according to the probability map of the target image determined in step 103. Specifically, the position information of the text in the target image is determined according to the size of the probability value of whether the content corresponding to each pixel in the probability map is text, and then the text region of the target image is obtained according to the determined position information of multiple texts, such as determining the position information of the text box in the target image, and obtaining the text region of the target image according to the determined position information of multiple text boxes. It should be noted that the text region can refer to the region containing multiple text boxes, rather than a single text box, and can be specifically set according to the actual needs of the user, and will not be specifically limited here.
[0080] According to the text detection method provided by the present invention, by segmenting the target image to be detected, a plurality of target sub-images are obtained. There is an overlapping area between adjacent target sub-images among the plurality of target sub-images. Then, the plurality of target sub-images are input into a text detection model to obtain probability maps corresponding to the plurality of target sub-images. The probability maps corresponding to the plurality of target sub-images are merged to obtain a probability map of the target image. The text area in the target image is determined according to the probability map of the target image. The text detection method provided by the present invention can accurately determine the position information of the text box in the target image, improve the accuracy of the text detection result, and enhance the user experience.
[0081] Based on any of the above embodiments, in this embodiment, the segmenting the target image to be detected to obtain a plurality of target sub-images includes:
[0082] Determine the size of the segmented image according to the input resolution of the pre-trained text detection model;
[0083] According to the size of the segmented image, segment the target image to be detected to obtain a plurality of target sub-images.
[0084] In this embodiment, it is necessary to determine the resolution size of the segmented image according to the size of the input resolution of the pre-trained text detection model, and then segment the target image to be detected according to the determined size of the segmented image to obtain a plurality of target sub-images. Among them, the size of the input resolution of the text detection model can be 800*800, or it can also be 600*600. In this embodiment, the size of the input resolution of the text detection model is determined to be the size of the segmented image. It should be noted that the resolution of the target image is also determined according to the number of pixel points in the horizontal and vertical directions of the target image, that is, the expression of the resolution of the target image is "number of horizontal pixel points × number of vertical pixel points".
[0085] In this embodiment, the size of the input resolution of the text detection model is 800*800, which is a square. It is necessary to determine the size of the input resolution of the text detection model to be the size of the segmented image. For example, if the input resolution of the text detection model is 800*800, then the size of the determined segmented image is also 800*800. It should be noted that in this embodiment, the target image to be detected can be cropped and segmented according to a preset direction to obtain a plurality of target sub-images. Among them, the resolution size of each target sub-image is equal to the resolution size of the segmented image. The specific processing flow is shown in the following embodiments and will not be introduced in detail here.
[0086] For example, according to the input resolution of the text detection model (such as 800*800), the size of the segmented image is determined to be 800*800. Suppose the resolution of the target image obtained is 800*2263. According to the determined size of the segmented image, and then with a 30% overlap ratio, it is determined that the target image can be cropped into (2263 - 800) / (800 - (800*30%)) + 1 = 3.6125 target sub-images. Among them, for the case where the number of target sub-images after segmentation has a decimal, the size of the last target sub-image cannot meet the size of the segmented image, that is, 800*800. A complete target sub-image can be obtained by supplementing the target value, that is, a target sub-image with a size of 800*800, so that finally 4 target sub-images are obtained. For example, a complete target sub-image is determined by supplementing target values such as 0 or 1 or 255.
[0087] According to the text detection method provided by the present invention, the size of the segmented image is determined according to the input resolution of the pre-trained text detection model. Then, according to the size of the segmented image, the target image to be detected is segmented with overlap to obtain multiple target sub-images. The complete target image is obtained through the obtained multiple target sub-images, avoiding the situation of inaccurate recognition caused by changing the resolution of the target image, and also avoiding the situation that the text area at the segmentation line cannot be recognized due to simple segmentation of the target image. Thus, it is possible to achieve precise processing of the target image to be detected, ensure the accuracy of the cropping and segmentation processing of the target image, and improve the efficiency of text detection.
[0088] Based on any of the above embodiments, in this embodiment, the step of segmenting the target image to be detected according to the size of the segmented image to obtain multiple target sub-images includes:
[0089] Determine the size of the segmentation window according to the size of the segmented image;
[0090] Move the segmentation window on the target image based on a pre-set overlap ratio to segment the target image, obtaining multiple target sub-images.
[0091] In this embodiment, it is necessary to determine the size of the segmentation window according to the determined size of the segmented image, and then move the segmentation window on the target image according to the pre-set overlap ratio to segment the target image, obtaining multiple target sub-images. Among them, the segmentation window refers to the sliding window used to segment the image. It should be noted that the size of the segmentation window is the same as the size of the segmented image. The shape of the segmentation window can be a square, a rectangle, or any other arbitrary shape, which is not specifically limited here.
[0092] In this embodiment, it is necessary to preset the overlapping ratio of the target image, and move the segmentation window according to the set overlapping ratio to achieve the segmentation of the target image and obtain multiple target sub-images. It should be noted that the overlapping ratio refers to the ratio of the overlapping content between the target sub-image and the previous target sub-image to the entire content of the target sub-image. That is to say, for each movement of the segmentation window, the overlapping ratio between the area where the segmentation window is located before the movement and the area where the segmentation window is located after the movement is the preset overlapping ratio. For example, if the target sub-image corresponding to the segmentation window before the movement is target sub-image A, and the target sub-image after the movement of the segmentation window is target sub-image B, when the preset overlapping ratio of the target image is 30%, the overlapping content between target sub-image A and target sub-image B accounts for 30% of the entire content of target sub-image A. Among them, the size of the overlapping ratio is set to any ratio between 0% and 100%, preferably 1% to 30%.
[0093] It should be noted that the size of the segmentation window is determined according to the size of the segmented image determined by the input resolution of the text detection model. If the size of the segmented image is N, then the size of the determined segmentation window is also N. For example, if the input resolution of the text detection model is 200*100, then the resolution of the determined segmented image is 200*100, and the size of the determined segmentation window is also 200*100.
[0094] In addition, during the movement of the segmentation window on the target image, the area where the segmentation window is mapped to the target image is determined as the segmentation area, and the pixels corresponding to the segmentation area on the target image are determined as the pixels of the target sub-image.
[0095] In this embodiment, there is no limitation on the moving segmentation method of the segmentation window. The "Z"-shaped moving segmentation method or the multi-vertical line moving segmentation method can be adopted.
[0096] In addition, in this embodiment, there is no limitation on whether to perform segmentation on the horizontal or vertical direction of the target image first. For example, the segmentation window can be first segmented along the horizontal direction of the target image to obtain at least one target sub-image, and then the segmentation window is moved along the vertical direction of the target image, and the segmentation window is continuously segmented along the horizontal direction of the target image to obtain at least one other target sub-image; the above process is repeated until all pixels of the target image correspond to target sub-images. For another example, the segmentation window can be first segmented along the vertical direction of the target image to obtain at least one target sub-image, and then the segmentation window is moved along the horizontal direction of the target image, and the segmentation window is continuously segmented downward along the vertical direction of the target image to obtain at least one other target sub-image; the above process is repeated until all pixels of the target image correspond to target sub-images. The specific segmentation process is shown in the following embodiments and will not be introduced in detail here, and the segmentation method can be set according to the actual needs of the user and is not limited here.
[0097] According to the document detection method provided by the present invention, the size of the segmentation window is determined according to the size of the segmented image, and the segmentation window is moved on the target image based on a preset overlapping ratio to segment the target image, obtaining a plurality of target sub-images, which can ensure the accuracy of the target image segmentation process, while ensuring the integrity of the text data and improving the accuracy of subsequent text recognition.
[0098] Based on any of the above embodiments, in this embodiment, the moving the segmentation window on the target image based on a preset overlapping ratio to segment the target image to obtain a plurality of target sub-images includes:
[0099] Determining an initial position of the segmentation window on the target image;
[0100] Sequentially moving the segmentation window on the target image in a first direction based on a preset overlapping ratio to segment the target image;
[0101] Judging whether the segmentation window reaches a segmentation termination position of the target image;
[0102] In the case that the segmentation window does not reach the segmentation termination position of the target image, moving the segmentation window on the target image in a second direction based on a preset overlapping ratio, and then re-executing the step of sequentially moving the segmentation window on the target image in the first direction based on a preset overlapping ratio to segment the target image;
[0103] In the case that the segmentation window reaches the segmentation termination position of the target image, obtaining a plurality of target sub-images of the target image based on the segmentation result of the target image.
[0104] In this embodiment, it is necessary to determine the initial position of the segmentation window on the target image, and then move the segmentation window sequentially in the first direction of the target image based on a preset overlap ratio to segment the target image. When the segmentation window reaches the segmentation termination position of the target image, the segmentation stops, and multiple target sub-images of the target image are obtained; when the segmentation window does not reach the segmentation termination position of the target image, it is necessary to move the segmentation window in the second direction of the target image also based on the preset overlap ratio to achieve the segmentation of the target image. It should be noted that the overlap ratio in the second direction and the overlap ratio in the first direction may be different or the same. The specific size of the overlap ratio can be set according to the actual needs of the user and is not specifically limited here.
[0105] It should be noted that the first direction can be the vertical direction of the target image or the horizontal direction of the target image. When the first direction is the vertical direction of the target image, the second direction is the horizontal direction of the target image; when the first direction is the horizontal direction of the target image, the second direction is the vertical direction of the target image. It can be specifically set according to the actual needs of the user and is not limited here.
[0106] It should be noted that the initial position of the target image refers to the position of a certain starting edge of the target image, and the segmentation termination position is the position where the segmentation window is located in the target image after moving the segmentation window to traverse the entire target image; for example, the initial position can be the top-leftmost edge of the target image or the bottom-leftmost edge of the target image. If the initial position of the target image is the top-leftmost edge of the target image, the segmentation termination position may be the bottom-rightmost edge or the top-rightmost edge of the target image, that is, the segmentation termination position refers to the position where the segmentation window terminates in the target image. When the segmentation window of the target image is segmented in a "Z" shape, when the initial position is the topmost and leftmost edges of the target image, the segmentation termination position may be the bottommost and rightmost edges of the target image or the topmost and rightmost edges of the target image, subject to moving the segmentation window to traverse the entire target image; when the segmentation window of the target image is segmented in a multi-vertical line form, when the initial position is the topmost and leftmost edges of the target image, the segmentation termination position may be the bottommost and rightmost edges of the target image. Among them, when the last page of content needs to supplement the value 0 or 1, the segmentation termination position is the position where the content terminates. It should be noted that the segmentation termination position is related to the initial position, the moving method of the segmentation window, and the size of the target object. The determination of the segmentation termination position can be specifically determined according to the actual needs of the user and is not limited here. For example, when it is determined that the union of the pixels of the segmented target sub-images includes all the pixels of the entire target image, it can be determined that the segmentation window reaches the segmentation termination position.
[0107] For example, assuming that the first direction is the vertical direction of the target image and the second direction is the horizontal direction of the target image, align the leftmost side of the segmentation window with the leftmost side of the target image and the uppermost side of the segmentation window with the uppermost side of the target image. Then, perform overlapping segmentation in the vertical direction of the target image. After the segmentation window reaches the bottommost part of the target image, move the segmentation window to the right along the horizontal direction of the target image, and then move the uppermost side of the segmentation window upward to align it with the uppermost side of the target image, so that there is a certain overlapping area between the obtained target sub-images and adjacent target sub-images. Then, perform overlapping segmentation again in the vertical direction of the target image until all pixels in the target image are segmented by the segmentation window, and then stop the segmentation.
[0108] Among them, in this embodiment, the segmentation window reaching the rightmost side and the bottommost part of the target image means that the final segmentation window includes the pixels on the rightmost side and the bottommost part of the target image. It should be noted that the segmentation termination position can also be the rightmost side of the target image and the position where the content of the target image ends, not necessarily reaching the bottommost part. In this embodiment, the segmentation ends by moving along a straight line to the rightmost side and the bottommost part of the target image, which belongs to the multi-vertical line form of moving segmentation. In other embodiments, it can also be segmented in a "Z" shape, which can be specifically set according to the actual needs of the user and will not be specifically limited here.
[0109] According to the text detection method provided by the present invention, determine the initial position of the segmentation window on the target image, and then move the segmentation window in the first direction of the target image based on a preset overlapping ratio in sequence to segment the target image; when the segmentation window does not reach the segmentation termination position of the target image, move the segmentation window in the second direction of the target image based on the preset overlapping ratio, and then re-execute the step of moving the segmentation window in the first direction of the target image based on the preset overlapping ratio in sequence to segment the target image; when the segmentation window reaches the segmentation termination position of the target image, based on the segmentation result of the target image, obtain multiple target sub-images of the target image, and improve the efficiency of text detection and enhance the user experience through the cropping and segmentation processing of the target image.
[0110] Based on any of the above embodiments, in this embodiment, the moving the segmentation window in the first direction of the target image based on a preset overlapping ratio in sequence to segment the target image includes:
[0111] Determine the target number of moving times of the segmentation window according to the length of the target image in the first direction, the length of the segmentation window in the first direction of the target image, and the preset overlapping ratio;
[0112] Move the segmentation window sequentially in the first direction of the target image based on a preset overlapping ratio, and segment the target image through the segmentation window until the number of times the segmentation window is moved reaches the target number of movement values;
[0113] In the case that there is a target sub - image in the segmentation result of the target image whose page size does not meet the specified size, adjust the target sub - image whose page size does not meet the specified size to the specified size.
[0114] In this embodiment, it is necessary to pre - determine the target number of movement values of the segmentation window according to the length of the target image in the first direction, the length of the segmentation window in the first direction of the target image, and the preset overlapping ratio. When moving the segmentation window sequentially in the first direction of the target image based on the preset overlapping ratio, use the segmentation window to segment the target image in turn until the number of times the segmentation window is moved reaches the target number of movement values, and then end the segmentation. Among them, if there is a case where the page size in the segmentation result does not meet the specified size, the segmentation result is processed by filling with target values until it meets the specified size. It should be noted that if the target image is a binary image, the target value can be 1, 0, or 255; if the target image is a color image, any color can be filled as long as there is no text content.
[0115] In addition, in this embodiment, the method of supplementing the target value is adopted. In other embodiments, the overlapping ratio between the target sub-graph whose page size does not meet the specified size and the adjacent target sub-graph can also be adjusted. When the length of the target sub-graph in the first direction is less than the length of the segmentation window in the first direction, the overlapping ratio of the previous target sub-graph adjacent to the target sub-graph in the first direction can be adjusted. The adjustment method is to determine that the length of the target sub-graph in the first direction is the first length; the length of the segmentation window in the first direction is the second length; determine the difference between the second length and the first length; move the segmentation window corresponding to the target sub-graph in the first direction by the distance corresponding to the difference, so that the overlapping ratio between the re-segmented target sub-graph and the previous target sub-graph is adjusted, and further the page size of the re-segmented target sub-graph meets the specified size, where the previous target sub-graph is the target sub-graph adjacent to the target sub-graph in the first direction. In this embodiment, for the target sub-graph whose page size does not meet the specified size, the overlapping ratio between the target sub-graph and the previous target sub-graph can be re-adjusted, without the need to supplement most of the target values, thus avoiding affecting the recognition rate due to supplementing the target values. For example, if the length of the target image in the first direction is 100 unit lengths (such as centimeters or pixels), the length of the segmentation window in the first direction is 50 unit lengths, and the original overlapping ratio is 20%, then the target sub-graph A, target sub-graph B, and target sub-graph C can be obtained. Among them, if the length of the target sub-graph C in the first direction is only 20, then the segmentation window corresponding to the target sub-graph C can be moved 30 unit lengths in the first direction, so that the overlapping ratio between the re-segmented target sub-graph C and the target sub-graph B is adjusted from 20% to 80%, that is, the overlapping part is 40 unit lengths, and further the length of the re-segmented target sub-graph C in the first direction is equal to the length of the segmentation window in the first direction, that is, the page size of the target sub-graph C meets the specified size.
[0116] Of course, the page size is the length in the first direction * the length in the second direction. The page size not meeting the specified size means that the length in the first direction or the length in the second direction does not meet the specified length. When the length of the target sub-graph in the second direction does not meet the specified length of the segmentation window in the second direction, similarly, the method of supplementing the target value can be adopted to complete it or adjust the overlapping ratio between the target sub-graph whose page size does not meet the specified size and the adjacent sub-graph.
[0117] The specific size of the overlapping ratio can be set according to the actual needs of the user, and no specific limitation is made here. Moreover, the overlapping ratio in the first direction can be different from the overlapping ratio in the second direction.
[0118] It should be noted that the page size of each segmentation result in the segmentation result of the target image needs to meet the specified size. The page size refers to the size of each side length of the segmentation result, that is, the segmentation result of the target image contains multiple segmentation results with the same page size; among them, the specified size refers to the size of the segmented image determined in advance. For example, if the resolution of the determined segmented image is 600*600, and the resolution of the obtained segmentation result is 400*600, then the resolution is adjusted to 600*600 along the horizontal direction of the segmentation result by the method of supplementing the target value.
[0119] It should be noted that the target movement number value can be obtained according to the length of the target image in the first direction, the length of the segmentation window in the first direction of the target image, and the preset overlap ratio. The size of the target movement number value is the same as the number of obtained target sub-images. The specific calculation method can be seen in the above embodiments and will not be introduced in detail here.
[0120] It should be noted that when the data of the target sub-image does not meet one piece, it is supplemented by the method of supplementing the target value to obtain a complete target sub-image. Specifically, Padding is used. Because most of the number of target sub-images will have a decimal part, but the resolution of the input text detection model is fixed. The remaining decimal part of the target sub-image needs to be filled to the size of a complete target sub-image, and the filled part can be the target value.
[0121] According to the text detection method provided by the present invention, first, according to the length of the target image in the first direction, the length of the segmentation window in the first direction of the target image, and the preset overlap ratio, the target movement number value of the segmentation window is determined. The segmentation window is sequentially moved on the first direction of the target image based on the preset overlap ratio, and the target image is segmented by the segmentation window until the movement number of the segmentation window reaches the target movement number value. Among them, when there is a segmentation result with a page size that does not meet the specified size in the segmentation result of the target image, the segmentation result with a page size that does not meet the specified size is supplemented by the method of supplementing the target value, or the overlap ratio between the target sub-image with a page size that does not meet the specified size and the adjacent sub-image is adjusted, which can ensure the accuracy of the target image segmentation process, improve the accuracy of text detection, and enhance the user experience.
[0122] Based on any of the above embodiments, in this embodiment, before moving the segmentation window on the target image based on the preset overlap ratio to segment the target image, the method further includes:
[0123] Adjust the resolution of the target image according to the size of the segmented image.
[0124] In this embodiment, it is necessary to adjust the resolution of the target image according to the determined size of the segmented image. For example, when the resolution of the target image meets certain preset requirements, a certain scaling process is performed on the target image so that the length of the short side of the target image is equal to the length of the segmented image. It should be noted that when it is determined that both the short side length and the long side length of the target image to be detected are greater than the short side length of the determined segmented image, the target image is scaled proportionally according to the short side length of the target image to be detected, and the short side length of the target image is scaled to be the same as the short side length of the determined segmented image. Here, scaling refers to reducing or enlarging the target image.
[0125] In this embodiment, assuming that the size of the segmented image is N, the length of the long side of the target image is H, and the length of the short side is W, it is necessary to scale the short side length W of the target image to N, and scale the long side length H of the target image to H*N / W. Then, the segmentation window is moved and cropped along the long side direction of the target image. According to the overlapping ratio of 30%, the side length of the overlapping area between every two adjacent target sub-images can be calculated as N*30%. In the calculation process, the first target sub-image does not need to calculate the overlapping area, so it is necessary to subtract the size of the first target sub-image H*N / W - N. The remaining target sub-images need to calculate the overlapping area, that is, the remaining target sub-images all need to segment an area of N - N*30%. That is, the determined calculation method is: in addition to the first target sub-image, (H*N / W - N) / (N - N*30%) target sub-images are required. Therefore, the number of target sub-images into which the target image is finally segmented is (H*N / W - N) / (N - N*30%) + 1.
[0126] For example, according to the input resolution of the text detection model (such as 800*800), it is determined that the size of the segmented image is 800*800. Assuming that the resolution of the target image obtained is 1587*4490, when the resolution of the short side of the target image has a difference of less than a preset multiple (such as 3 times) from the resolution of the segmented image, the target image (such as 1587*4490) is scaled proportionally by the short side to 800*2263. Then, according to the overlapping ratio of 30%, it is determined that the target image can be cropped into (2263 - 800) / (800 - (800*30%)) + 1 = 3.6125 target sub-images. When there is a decimal in the number of obtained target sub-images, a complete target sub-image is obtained by supplementing the target value, that is, a total of 4 target sub-images are obtained.
[0127] According to the text detection method provided by the present invention, the resolution of the target image is adjusted according to the size of the segmented image. By adjusting the resolution, the number of target sub-images to be recognized is reduced. It is possible to reduce the computing power and recognition time required for recognition while ensuring the detection accuracy of the document detection result, thereby improving the processing efficiency of text detection.
[0128] Based on any of the above embodiments, in this embodiment, adjusting the resolution of the target image according to the size of the segmented image includes:
[0129] Calculating the difference value between the length of the first side of the segmented image and the length of the first side of the target image; wherein, the first side of the segmented image is any side of the segmented image, and the first side of the target image is the shorter side of the target image;
[0130] In the case where the difference value is less than a preset threshold, the resolution of the target image is adjusted proportionally so that the length of the first side of the segmented image is the same as the length of the first side of the adjusted target image.
[0131] In this embodiment, it is necessary to calculate the difference value between the length of the first side of the segmented image and the length of the first side of the target image. When the difference value is less than the preset threshold, the resolution of the target image is adjusted proportionally so that the length of the first side of the adjusted target image is the same as the length of the first side of the segmented image. Among them, the difference value can be the difference between the length of the first side of the segmented image and the length of the first side of the target image, or the multiple between the length of the first side of the segmented image and the length of the first side of the target image. The preset threshold can refer to the difference between the resolution of the shorter side of the target image and the size of the segmented image, or the multiple between the resolution of the shorter side of the target image and the size of the segmented image; the first side of the target image refers to the shorter side of the target image. For example, if the resolution of the target image is 1500*2500, the first side is the side corresponding to the length of 1500. It should be noted that proportional adjustment means adjusting the long side and the short side of the target image according to the same adjustment ratio.
[0132] It should be noted that since the segmented image can also be a square in other embodiments, that is, the first side of the segmented image is any side of the segmented image, and the first side of the segmented image is not limited.
[0133] It should be noted that adjusting the resolution of the target image in this embodiment means performing an equal-proportion scaling process on the target image according to the relationship between the length of the short side of the target image and the size of the segmented image. When the difference between the length of the short side of the target image and the resolution of the segmented image is less than three times, an equal-proportion scaling process is performed on the target image. If it exceeds three times, it is not recommended to perform an equal-proportion scaling process on the target image to avoid serious distortion of the scaled target image due to too large a difference, which reduces the detection accuracy of the text area.
[0134] According to the text detection method provided by the present invention, it is necessary to calculate the difference value between the length of the first side of the segmented image and the length of the first side of the target image. When the difference value is less than the preset threshold, the resolution of the target image is adjusted proportionally so that the length of the first side of the segmented image is the same as the length of the first side of the adjusted target image, which can ensure the detection accuracy of the file detection result and improve the processing efficiency of text detection.
[0135] Based on any of the above embodiments, in this embodiment, the merging of the probability maps corresponding to the multiple target sub-images to obtain the probability map of the target image includes:
[0136] Determine the overlapping regions and non-overlapping regions in the multiple target sub-images;
[0137] According to the probability values of each pixel in the overlapping region in the probability maps corresponding to the multiple target sub-images, calculate the probability values of each pixel in the overlapping region to obtain the probability map corresponding to the overlapping region;
[0138] According to the probability maps corresponding to the multiple target sub-images, obtain the probability map corresponding to the non-overlapping region;
[0139] Merge the probability maps corresponding to the non-overlapping regions in the multiple target sub-images and the probability map corresponding to the overlapping region in the multiple target sub-images to obtain the probability map of the target image.
[0140] In this embodiment, it is necessary to determine the overlapping regions and non-overlapping regions of multiple target sub-images, and then calculate the probability values of each pixel in the overlapping regions based on the probability values of each pixel in the overlapping regions in the probability maps corresponding to the multiple target sub-images, so as to obtain the probability map corresponding to the overlapping regions. Then, according to the complete probability maps corresponding to the multiple target sub-images, the probability map corresponding to the non-overlapping regions is obtained. The probability maps corresponding to the non-overlapping regions and the overlapping regions of the multiple target sub-images are combined to obtain the probability map of the target image. It should be noted that in this embodiment, the overlapping region refers to the region where the text contents corresponding to adjacent target sub-images repeat each other. For example, 30% of the content at the bottom of target sub-image A is the same as 30% of the content at the front end of target sub-image B. The same part is determined as the overlapping region. In this embodiment, the overlapping region and the non-overlapping region are determined according to the magnitude of the probability value. The specific judgment method can be set according to the actual needs of the user and is not specifically limited herein.
[0141] It should be noted that the probability map corresponding to the overlapping regions of the multiple target sub-images is calculated and confirmed based on the probability values in the probability maps corresponding to each pixel value in the overlapping regions.
[0142] According to the text detection method provided by the present invention, based on the probability values of each pixel in the overlapping regions in the probability maps corresponding to the multiple target sub-images, the probability values of each pixel in the overlapping regions are calculated to obtain the probability map corresponding to the overlapping regions. Then, according to the probability maps corresponding to the multiple target sub-images, the probability map corresponding to the non-overlapping regions is obtained. The probability maps corresponding to the non-overlapping regions and the overlapping regions in the multiple target sub-images are combined to obtain the probability map of the target image, which can ensure the accuracy of the text detection result, improve the efficiency of text detection, and enhance the user experience.
[0143] Based on any of the above embodiments, in this embodiment, the calculating the probability values of each pixel in the overlapping regions based on the probability values of each pixel in the overlapping regions in the probability maps corresponding to the multiple target sub-images includes:
[0144] Determine the magnitudes and the number of the multiple probability values corresponding to a first pixel in the probability maps corresponding to the multiple target sub-images; wherein, the first pixel is any pixel in the overlapping region;
[0145] Determine the probability value of the first pixel by voting according to the relationship between the multiple probability values corresponding to the first pixel and a preset probability threshold;
[0146] Or,
[0147] Calculate the average value of the multiple probability values and use the average value as the probability value of the first pixel.
[0148] In this embodiment, it is necessary to determine the magnitudes and the number of multiple probability values corresponding to a first pixel in the probability maps corresponding to multiple target sub - graphs, and then, according to the relationship between the multiple probability values corresponding to the first pixel and a preset probability threshold, determine the probability value of the first pixel by voting; or, calculate the average value of the multiple probability values corresponding to the first pixel, and use this average value as the probability value of the first pixel. It should be noted that the first pixel is any pixel within the overlapping area.
[0149] It should be noted that, according to the relationship between the multiple probability values corresponding to the first pixel and the preset probability threshold, the probability value of the first pixel is determined by voting. Among them, the magnitude of the preset probability threshold can be any value and can be set according to the actual needs of the user. In addition, the specific voting determination method is shown in the following embodiments and will not be specifically limited here.
[0150] It should be noted that in this embodiment, the average value of the multiple probability values corresponding to the first pixel can also be calculated, and this average value is used as the probability value of the first pixel. For a detailed introduction, see the following embodiments.
[0151] According to the text detection method provided by the present invention, based on the magnitudes of the multiple probability values corresponding to the determined first pixel in the probability maps corresponding to multiple target sub - graphs and the relationship with the preset threshold, or by calculating the average value of the multiple probability values, the probability value of the first pixel is determined, which improves the accuracy of text detection and enhances the user experience.
[0152] Based on any of the above - mentioned embodiments, in this embodiment, calculating the probability values of each pixel within the overlapping area according to the probability values of each pixel within the overlapping area in the probability maps corresponding to the multiple target sub - graphs includes:
[0153] Determine the magnitudes of the multiple probability values corresponding to a first pixel in the probability maps corresponding to the multiple target sub - graphs; where the first pixel is any pixel within the overlapping area;
[0154] In the case where the magnitudes of the multiple probability values corresponding to the first pixel are all greater than the preset probability threshold, use a first probability value as the probability value of the first pixel; where the first probability value is the maximum value among the multiple probability values corresponding to the first pixel;
[0155] In the case where the magnitudes of the multiple probability values corresponding to the first pixel are all less than the preset probability threshold, use a second probability value as the probability value of the first pixel; where the second probability value is the minimum value among the multiple probability values corresponding to the first pixel;
[0156] Further, in one embodiment, when some of the multiple probability values corresponding to the first pixel are greater than a preset probability threshold and some are less than the preset probability threshold, the third probability value is used as the probability value of the first pixel, or no probability value is set for the first pixel; wherein, the third probability value is any one of the multiple probability values corresponding to the first pixel.
[0157] In another embodiment, when some of the multiple probability values corresponding to the first pixel are greater than a preset probability threshold and some are less than the preset probability threshold, the voting method or the averaging method can also be used to set the probability value for the first pixel.
[0158] In this embodiment, it is necessary to determine the magnitudes of the multiple probability values corresponding to the first pixel in the probability maps corresponding to multiple target subgraphs, and then determine the probability value of the first pixel according to the relationship between the magnitudes of the multiple probability values corresponding to the first pixel and the preset threshold, where the first pixel is any pixel within the overlapping region.
[0159] When all of the multiple probability values corresponding to the first pixel are greater than the preset probability threshold, the maximum value among the multiple probability values corresponding to the first pixel is determined as the probability value of the first pixel. Suppose the preset probability threshold is 0.5. When it is determined that the multiple probability values corresponding to the first pixel are 0.6, 0.7, 0.65, and 0.66 respectively, all of which are greater than the preset probability threshold 0.5, then the maximum probability value 0.7 is determined as the probability value of the first pixel. Similarly, when all of the multiple probability values corresponding to the first pixel are less than the preset probability threshold, the minimum value among the multiple probability values corresponding to the first pixel is determined as the probability value of the first pixel.
[0160] When some of the multiple probability values corresponding to the first pixel are greater than the preset probability threshold and some are less than the preset probability threshold, any one of the multiple probability values corresponding to the first pixel is determined as the probability value of the first pixel, or no probability value is set for this first pixel.
[0161] It should be noted that there can be multiple specific processing methods for determining the probability values of each pixel in the overlapping region of multiple target subgraphs. Suppose there are N target subgraphs with an overlapping region, and x is set as the pixel at the same position in the N target subgraphs. x1 represents the pixel x in target sub Figure 1 graph, and x2 represents the pixel x in target sub Figure 2The pixels x, …, xN therein represent the pixels x in the target sub-graph N, P(xn) represents the probability value that the pixel x in the target sub-graph n is text, and the preset probability threshold is M. Among the N target sub-graphs, if the probability values of the N target sub-graphs are all greater than the preset probability threshold M, then the maximum probability value among the N probability values is taken as the probability value of the pixel x; among the N target sub-graphs, if the N probability values are all less than the preset probability threshold M, then the minimum probability value among the N probability values is taken as the probability value of the pixel x; in the case where there are both probability values greater than the preset probability threshold M and probability values less than the preset probability threshold M, any one of the probability values can be selected to be determined as the probability value of the pixel x, or the pixel x can be directly ignored without determining its probability value.
[0162] According to the text detection method provided by the present invention, based on the magnitudes of the multiple probability values corresponding to the determined first pixel in the probability graphs corresponding to the multiple target sub-graphs and the relationship with the preset threshold, the probability value of the first pixel is determined, which improves the accuracy of text detection and enhances the user experience.
[0163] Based on any one of the above embodiments, in this embodiment, calculating the probability values of the respective pixels in the overlapping region based on the probability values of the respective pixels in the overlapping region in the probability graphs corresponding to the multiple target sub-graphs includes:
[0164] Determining the magnitudes of the multiple probability values corresponding to the first pixel in the probability graphs corresponding to the multiple target sub-graphs and the number of the multiple probability values; wherein, the first pixel is any pixel in the overlapping region;
[0165] Determining a first number of the probability values among the multiple probability values whose magnitudes are greater than the preset threshold and a second number of the probability values among the multiple probability values whose magnitudes are less than or equal to the preset threshold;
[0166] In the case where the first number is greater than the second number, taking the maximum value among the multiple probability values as the probability value of the first pixel;
[0167] In the case where the first number is less than or equal to the second number, taking the minimum value among the multiple probability values as the probability value of the first pixel.
[0168] In this embodiment, it is necessary to determine the magnitudes of the multiple probability values corresponding to the first pixel in the probability graphs corresponding to the multiple target sub-graphs and the number of the multiple probability values, and then it is necessary to determine a first number of the probability values among the multiple probability values whose magnitudes are greater than the preset threshold and a second number of the probability values among the multiple probability values whose magnitudes are less than or equal to the preset threshold. Among them, the magnitude of the probability value of the preset threshold can be set according to the actual needs of the user, and no specific limitation is made here.
[0169] When the determined first quantity is greater than the second quantity, the maximum value among the multiple probability values is used as the probability value of the first pixel. When the first quantity is less than or equal to the second quantity, the minimum value among the multiple probability values is used as the probability value of the first pixel.
[0170] For example, among the N obtained target subgraphs, the probability value of the preset threshold is M. The first quantity of the probability values greater than the preset threshold probability value M among the multiple probability values is N1, and the second quantity of the probability values less than or equal to the preset threshold probability value M among the multiple probability values is N2, and N1 + N2 = N;
[0171] When the first quantity N1 is greater than the second quantity N2, the maximum value of the multiple probability values corresponding to the first quantity N1 is taken as the probability value of the first pixel. When the first quantity N1 is less than the second quantity N2, the minimum value of the multiple probability values corresponding to the second quantity N2 is taken as the probability value of the first pixel.
[0172] According to the text detection method provided by the present invention, based on the magnitudes of the multiple probability values corresponding to the determined first pixel in the probability maps corresponding to the multiple target subgraphs and the first quantity and the second quantity determined according to the relationship between the multiple probability values and the probability value of the preset threshold, and determining the probability value of the first pixel according to the relationship between the first quantity and the second quantity, the accuracy of the text detection result is improved, and the efficiency of text processing is improved.
[0173] Based on any of the above embodiments, in this embodiment, calculating the probability value of the pixel in the overlapping region according to the probability value of the pixel in the overlapping region in the probability maps corresponding to the multiple target subgraphs includes:
[0174] Determining the magnitudes of the multiple probability values corresponding to the first pixel in the probability maps corresponding to the multiple target subgraphs; wherein, the first pixel is any pixel in the overlapping region;
[0175] Calculating the average value of the multiple probability values, and taking the average value as the probability value of the first pixel.
[0176] In this embodiment, it is also necessary to determine the magnitudes of the multiple probability values corresponding to the first pixel in the probability maps corresponding to the multiple target subgraphs and calculate the average value of the multiple probability values. In this embodiment, the calculated average value of the probability values is determined as the probability value of the first pixel. Suppose the preset probability threshold is 0.5. When the multiple probability values corresponding to the first pixel are determined to be 0.6, 0.7, 0.5, and 0.4 respectively, and the calculated average value of the multiple probability values is 0.55, then the average value 0.55 is determined as the probability value of the first pixel.
[0177] According to the text detection method provided by the present invention, the average value of multiple calculated probability values is determined as the probability value of the first pixel, which improves the accuracy of the text detection result and the efficiency of text processing.
[0178] Based on any of the above embodiments, in this embodiment, determining the text region in the target image according to the probability map of the target image includes:
[0179] Determining the probability value of each pixel in the target image according to the probability map of the target image;
[0180] Determining the coordinate value of the text region in the target image according to the probability value of each pixel in the target image;
[0181] Determining the text region of the target image according to the coordinate value of the text region in the target image.
[0182] In this embodiment, according to the probability map of the target image, the probability value of each pixel in the target image is determined, and then according to the probability value of each pixel, the coordinate value of the text region in the target image is determined, and then the text region of the target image is determined. Among them, the text region refers to a region containing multiple text boxes and a region with multiple text contents.
[0183] According to the text detection method provided by the present invention, first, according to the probability map of the target image processed by the text detection model, the probability value of each pixel in the target image is determined, and then according to the probability value of each pixel in the target image, the coordinate value of the text region in the target image is determined, and then according to the coordinate value of the text region in the target image, the text region of the target image is determined, ensuring the accuracy of text detection and the non-loss of text data, and improving the speed of text detection processing.
[0184] Based on any of the above embodiments, in this embodiment, according to the input resolution of the text detection model (such as 800*800), if the length of the short side of the target image is less than 3 times the difference from the input resolution, the target image (such as 1587*4490) needs to be scaled proportionally to 800*2263 according to the short side, and then the target image is cropped and segmented according to a 30% overlap ratio. The target image can be cropped into (2263 - 800) / (800 - (800*30%)) + 1 = 3.6125 target sub-images. For the target sub-images whose page size does not meet the specified size, the target sub-images less than 1 are supplemented to form integer target sub-images according to the supplementary target value, as Figure 2 shown, where Figure 2 the bottom black part in is the result after padding with 0. It should be noted that there will be no position information of text content in the part processed by padding with the target value, and thus the text box position information of this part of the content cannot be obtained.
[0185] In this embodiment, all target sub - graphs need to be stitched and packed and sent into the text detection model to obtain 4 corresponding probability maps. Then, the obtained multiple probability maps are stitched, and pixel - point calculations are performed on the overlapping regions to obtain the complete probability map of the target image. Finally, the corresponding text - box coordinate information is obtained according to the probability map to determine the text region of the target image, as specifically shown in Figure 3 the figure. Among them, the white area is the text region of the obtained target image. It should be noted that when cropping and segmenting the target image in this embodiment, a certain overlapping region is set. If no overlapping region is set, important text information may be missed during equal - ratio cropping and segmentation, as shown in Figure 4 the area pointed by the arrow in the figure. The two target sub - graphs will cut open the content pointed by the arrow, resulting in inaccurate text detection.
[0186] It should be noted that the size of the segmented image needs to be determined according to the input resolution of the text detection model, and there is an equal relationship between the two. In addition, in this embodiment, scaling processing is performed according to the ratio of the short side of the target image to the segmented image, provided that the difference between the target image and the input resolution does not exceed the preset threshold; if it exceeds the preset threshold, no scaling processing is performed on the target image, but multiple target sub - graphs are cropped and segmented from the target image both horizontally and vertically at an overlapping ratio of 30%. Among them, in this embodiment, the size of the preset threshold is 3 times. When the difference between the short - side resolution of the target image and the input resolution is less than 3 times, corresponding scaling processing is performed.
[0187] It should be noted that the determination method of the probability value of the pixels in the overlapping region of N target sub - graphs is as follows: there are N target sub - graphs overlapping. Let x be the pixel at the same position in the N target sub - graphs. x1 represents the pixel x in target sub - graph Figure 1 , x2 represents the pixel x in target sub - graph Figure 2 , ……, xN represents the pixel x in target sub - graph N. P(xn) represents the probability that the pixel x in target sub - graph N is text. Let M be the probability value of the preset threshold, and the determination method of the probability value of the pixels in the overlapping region of N target sub - graphs is described as follows.
[0188] (1) Greedy method
[0189] Among the N target subgraphs, if the probability values of the N target subgraphs are all greater than the preset probability threshold M, then the maximum probability value among the N probability values is taken as the probability value of pixel x; among the N target subgraphs, if the N probability values are all less than the preset probability threshold M, then the minimum probability value among the N probability values is taken as the probability value of this pixel x; in the case where there are both probability values greater than the preset probability threshold M and probability values less than the preset probability threshold M, any one of the probability values is selected and determined as the probability value of pixel x, or it is directly ignored and its probability value is not determined.
[0190] (2) Voting method
[0191] For example, among the N target subgraphs obtained, the preset threshold probability value is M, the first quantity of probability values greater than the preset threshold probability value M among the multiple probability values is N1, the second quantity of probability values less than or equal to the preset threshold probability value M among the multiple probability values is N2, and N1 + N2 = N;
[0192] When the first quantity N1 is greater than the second quantity N2, the maximum value among the multiple probability values corresponding to the first quantity N1 is taken as the probability value of the first pixel. When the first quantity N1 is less than the second quantity N2, the minimum value among the multiple probability values corresponding to the second quantity N2 is taken as the probability value of the first pixel.
[0193] ③ Mean method
[0194] Determine the magnitudes of the multiple probability values corresponding to a pixel in the probability maps corresponding to the multiple target subgraphs, calculate the average value of the multiple probability values, and determine the average value of the calculated probability values as the probability value of this pixel.
[0195] Figure 5 A text detection device provided by the present invention, as Figure 5 shown, the text detection device provided by the present invention includes:
[0196] A segmentation module 501, configured to segment a target image to be detected to obtain multiple target subgraphs; wherein, there is an overlapping area between adjacent target subgraphs among the multiple target subgraphs;
[0197] An input module 502, configured to input the multiple target subgraphs into a text detection model to obtain probability maps corresponding to the multiple target subgraphs;
[0198] A merging module 503, configured to merge the probability maps corresponding to the multiple target subgraphs to obtain a probability map of the target image;
[0199] A determination module 504, configured to determine a text region in the target image according to the probability map of the target image;
[0200] Among them, the probability map is used to describe the probability value that the content of the pixels in the image corresponding to the probability map is text.
[0201] According to the text detection device provided by the present invention, by segmenting the target image to be detected, a plurality of target sub-images are obtained; among them, there are overlapping regions between adjacent target sub-images in the plurality of target sub-images; the plurality of target sub-images are input into the text detection model to obtain probability maps corresponding to the plurality of target sub-images; the probability maps corresponding to the plurality of target sub-images are merged to obtain the probability map of the target image; the text region in the target image is determined according to the probability map of the target image. The text detection device provided by the present invention can accurately determine the position information of the text box, improve the accuracy of the file detection result, and enhance the user experience.
[0202] Furthermore, the segmentation module 501 is further configured to:
[0203] Determine the size of the segmented image according to the input resolution of the pre-trained text detection model;
[0204] Segment the target image to be detected according to the size of the segmented image to obtain a plurality of target sub-images.
[0205] According to the text detection model provided by the present invention, determine the size of the segmented image according to the input resolution of the pre-trained text detection model, and then segment the target image to be detected according to the size of the segmented image to obtain a plurality of target sub-images. It can achieve precise processing of the target image to be detected, ensure the accuracy of the cropping and segmentation processing of the target image, and improve the efficiency of text detection.
[0206] Furthermore, the segmentation module 501 is further configured to:
[0207] Determine the size of the segmentation window according to the size of the segmented image;
[0208] Move the segmentation window on the target image based on a preset overlapping ratio to segment the target image to obtain a plurality of target sub-images.
[0209] According to the document detection device provided by the present invention, determine the size of the segmentation window according to the size of the segmented image, move the segmentation window on the target image based on a preset overlapping ratio to segment the target image to obtain a plurality of target sub-images, which can ensure the accuracy of the target image segmentation processing, and at the same time ensure the integrity of the text data, and improve the accuracy of subsequent text recognition.
[0210] Furthermore, the segmentation module 501 is further configured to:
[0211] Determine the initial position of the segmentation window on the target image;
[0212] Move the segmentation window sequentially in a first direction of the target image based on a preset overlapping ratio to segment the target image;
[0213] Determine whether the segmentation window reaches a segmentation termination position of the target image;
[0214] In a case where the segmentation window does not reach the segmentation termination position of the target image, move the segmentation window in a second direction of the target image based on the preset overlapping ratio, and then re - execute the step of moving the segmentation window sequentially in the first direction of the target image based on the preset overlapping ratio to segment the target image;
[0215] In a case where the segmentation window reaches the segmentation termination position of the target image, obtain a plurality of target sub - images of the target image based on the segmentation result of the target image.
[0216] According to the text detection device provided by the present invention, determine an initial position of a segmentation window on a target image, and then move the segmentation window sequentially in a first direction of the target image based on a preset overlapping ratio to segment the target image; when the segmentation window does not reach the segmentation termination position of the target image, move the segmentation window in a second direction of the target image based on the preset overlapping ratio, and then re - execute the step of moving the segmentation window sequentially in the first direction of the target image based on the preset overlapping ratio to segment the target image; when the segmentation window reaches the segmentation termination position of the target image, obtain a plurality of target sub - images of the target image based on the segmentation result of the target image. The efficiency of text detection is improved by the cropping and segmentation process of the target image, and the user experience is enhanced.
[0217] Further, the segmentation module 501 is further configured to:
[0218] Determine a target moving number value of the segmentation window according to a length of the target image in the first direction, a length of the segmentation window in the first direction of the target image, and the preset overlapping ratio;
[0219] Move the segmentation window sequentially in the first direction of the target image based on the preset overlapping ratio, and segment the target image through the segmentation window until the moving number of the segmentation window reaches the target moving number value;
[0220] In a case where there is a target sub - image in the segmentation result of the target image whose page size does not meet a specified size, adjust the target sub - image whose page size does not meet the specified size to the specified size.
[0221] According to the text detection device provided by the present invention, first, based on the length of the target image in the first direction, the length of the segmentation window in the first direction of the target image, and a preset overlapping ratio, a target number of moving times of the segmentation window is determined. The segmentation window is sequentially moved in the first direction of the target image based on the preset overlapping ratio to segment the target image through the segmentation window until the number of moving times of the segmentation window reaches the target number of moving times. Among them, when there is a target sub-image in the segmentation result of the target image whose page size does not meet the specified size, the target sub-image whose page size does not meet the specified size is adjusted to the specified size, which can ensure the accuracy of the target image segmentation process, provide the accuracy of text detection, and improve the user experience.
[0222] Further, the text detection device is further configured to:
[0223] Adjust the resolution of the target image according to the size of the segmented image.
[0224] According to the text detection device provided by the present invention, adjusting the resolution of the target image according to the size of the segmented image can ensure the detection accuracy of the file detection result and improve the processing efficiency of text detection.
[0225] Further, the text detection device is further configured to:
[0226] Calculate the difference value between the length of the first side of the segmented image and the length of the first side of the target image; wherein, the first side of the segmented image is any side of the segmented image, and the first side of the target image is the shorter side of the target image;
[0227] In the case where the difference value is less than a preset threshold, adjust the resolution of the target image proportionally so that the length of the first side of the segmented image is the same as the length of the first side of the adjusted target image.
[0228] According to the text detection device provided by the present invention, it is necessary to calculate the difference between the length of the first side of the segmented image and the length of the first side of the target image. When the difference is less than the preset threshold, adjust the resolution of the target image proportionally so that the length of the first side of the segmented image is the same as the length of the first side of the adjusted target image, which can ensure the detection accuracy of the file detection result and improve the processing efficiency of text detection.
[0229] Further, the merging module 503 is further configured to:
[0230] Determine the overlapping area and non-overlapping area among the multiple target sub-images;
[0231] Calculate the probability values of the respective pixels within the overlapping region based on the probability values of the respective pixels within the overlapping region in the probability maps corresponding to the multiple target sub - graphs, to obtain the probability map corresponding to the overlapping region;
[0232] Based on the probability maps corresponding to the multiple target sub - graphs, obtain the probability map corresponding to the non - overlapping region;
[0233] Merge the probability map corresponding to the non - overlapping region in the multiple target sub - graphs and the probability map corresponding to the overlapping region in the multiple target sub - graphs, to obtain the probability map of the target image.
[0234] According to the text detection device provided by the present invention, calculate the probability values of the respective pixels within the overlapping region based on the probability values of the respective pixels within the overlapping region in the probability maps corresponding to the multiple target sub - graphs, to obtain the probability map corresponding to the overlapping region, then obtain the probability map corresponding to the non - overlapping region based on the probability maps corresponding to the multiple target sub - graphs, and merge the probability map corresponding to the non - overlapping region and the probability map corresponding to the overlapping region in the multiple target sub - graphs, to obtain the probability map of the target image, which can ensure the accuracy of the text detection result, improve the efficiency of text detection, and enhance the user experience.
[0235] Furthermore, the merging module 503 is further configured to:
[0236] Determine the magnitudes and the number of the multiple probability values corresponding to a first pixel in the probability maps corresponding to the multiple target sub - graphs; wherein, the first pixel is any pixel within the overlapping region;
[0237] Determine the probability value of the first pixel by voting according to the relationship between the multiple probability values corresponding to the first pixel and a preset probability threshold;
[0238] Or,
[0239] Calculate the average value of the multiple probability values, and use the average value as the probability value of the first pixel.
[0240] According to the text detection device provided by the present invention, determine the probability value of the first pixel by calculating the magnitudes of the multiple probability values corresponding to the determined first pixel in the probability maps corresponding to the multiple target sub - graphs and the relationship with the preset threshold, or by calculating the average value of the multiple probability values, which improves the accuracy of text detection and enhances the user experience.
[0241] Furthermore, the determining module 504 is further configured to:
[0242] Determine the probability values of the respective pixels in the target image according to the probability map of the target image;
[0243] Determine the coordinate values of the text region in the target image according to the probability values of each pixel in the target image;
[0244] Determine the text region of the target image according to the coordinate values of the text region in the target image.
[0245] According to the text detection device provided by the present invention, first determine the probability values of each pixel in the target image according to the probability map of the target image processed by the text detection model, then determine the coordinate values of the text region in the target image according to the probability values of each pixel in the target image, and further determine the text region of the target image according to the coordinate values of the text region in the target image, which ensures the accuracy of text detection and the non-loss of text data, and improves the speed of text detection processing.
[0246] Since the principle of the device in the embodiment of the present invention is the same as that of the method in the above embodiment, the more detailed explanation content will not be elaborated here.
[0247] Figure 6 It is a schematic diagram of the entity structure of the electronic device provided in the embodiment of the present invention. As Figure 6 shown, the present invention provides an electronic device, including: a processor 601, a memory 602, and a bus 603;
[0248] Among them, the processor 601 and the memory 602 communicate with each other through the bus 603;
[0249] The processor 601 is configured to call program instructions in the memory 602 to execute the methods provided in the above method embodiments, for example, including: segmenting a target image to be detected to obtain a plurality of target sub-images; wherein, there is an overlapping region between adjacent target sub-images among the plurality of target sub-images; inputting the plurality of target sub-images into a text detection model to obtain probability maps corresponding to the plurality of target sub-images; merging the probability maps corresponding to the plurality of target sub-images to obtain a probability map of the target image; determining the text region in the target image according to the probability map of the target image; wherein, the probability map is used to describe the probability value that the content of the pixel in the image corresponding to the probability map is text.
[0250] In an embodiment of the present invention, a non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions cause the computer to execute the methods provided in the above method embodiments. For example, it includes: segmenting a target image to be detected to obtain a plurality of target sub-images; wherein, there is an overlapping area between adjacent target sub-images among the plurality of target sub-images; inputting the plurality of target sub-images into a text detection model to obtain probability maps corresponding to the plurality of target sub-images; merging the probability maps corresponding to the plurality of target sub-images to obtain a probability map of the target image; determining a text area in the target image according to the probability map of the target image; wherein, the probability map is used to describe the probability value that the content of the pixel in the image corresponding to the probability map is text.
[0251] The present invention also provides a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the methods provided in the above methods. The method includes: segmenting a target image to be detected to obtain a plurality of target sub-images; wherein, there is an overlapping area between adjacent target sub-images among the plurality of target sub-images; inputting the plurality of target sub-images into a text detection model to obtain probability maps corresponding to the plurality of target sub-images; merging the probability maps corresponding to the plurality of target sub-images to obtain a probability map of the target image; determining a text area in the target image according to the probability map of the target image; wherein, the probability map is used to describe the probability value that the content of the pixel in the image corresponding to the probability map is text.
[0252] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: various media such as ROM, RAM, magnetic disk, or optical disk that can store program codes.
[0253] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A text detection method, characterized in that, Including: Segmenting a target image to be detected to obtain a plurality of target sub-images; wherein, there is an overlapping area between adjacent target sub-images among the plurality of target sub-images; Inputting the plurality of target sub-images into a text detection model to obtain probability maps corresponding to the plurality of target sub-images; Merging the probability maps corresponding to the plurality of target sub-images to obtain a probability map of the target image; Determining a text area in the target image according to the probability map of the target image; Wherein, the probability map is used to describe the probability value that the content of the pixel in the image corresponding to the probability map is text; The merging the probability maps corresponding to the plurality of target sub-images to obtain a probability map of the target image includes: Determining an overlapping area and a non-overlapping area in the plurality of target sub-images; Calculating the probability value of each pixel in the overlapping area according to the probability values of each pixel in the overlapping area in the probability maps corresponding to the plurality of target sub-images to obtain a probability map corresponding to the overlapping area; Obtaining a probability map corresponding to the non-overlapping area according to the probability maps corresponding to the plurality of target sub-images; Merging the probability map corresponding to the non-overlapping area in the plurality of target sub-images and the probability map corresponding to the overlapping area in the plurality of target sub-images to obtain a probability map of the target image.
2. The text detection method according to claim 1, characterized in that, The segmenting the target image to be detected to obtain a plurality of target sub-images includes: Determining the size of the segmented image according to the input resolution of a pre-trained text detection model; Segmenting the target image to be detected according to the size of the segmented image to obtain a plurality of target sub-images.
3. The text detection method according to claim 2, characterized in that, The segmenting the target image to be detected according to the size of the segmented image to obtain a plurality of target sub-images includes: Determining the size of a segmentation window according to the size of the segmented image; Moving the segmentation window on the target image based on a preset overlapping ratio to segment the target image to obtain a plurality of target sub-images.
4. The text detection method according to claim 3, characterized in that, The moving the segmentation window on the target image based on a preset overlapping ratio to segment the target image to obtain a plurality of target sub-images includes: Determining an initial position of the segmentation window on the target image; Sequentially moving the segmentation window on the target image in a first direction based on a preset overlapping ratio to segment the target image; Judging whether the segmentation window reaches a segmentation termination position of the target image; In the case that the segmentation window does not reach the segmentation termination position of the target image, moving the segmentation window in a second direction on the target image based on a preset overlapping ratio, and then re-executing the step of sequentially moving the segmentation window on the target image in the first direction based on a preset overlapping ratio to segment the target image; In the case that the segmentation window reaches the segmentation termination position of the target image, obtaining a plurality of target sub-images of the target image based on the segmentation result of the target image.
5. The text detection method according to claim 4, characterized in that, The sequentially moving the segmentation window on the target image in a first direction based on a preset overlapping ratio to segment the target image includes: Determine a target number of moving times of the segmentation window according to the length of the target image in a first direction, the length of the segmentation window in the first direction of the target image, and a preset overlapping ratio. Based on the preset overlapping ratio, sequentially move the segmentation window in the first direction of the target image, and segment the target image through the segmentation window until the number of moving times of the segmentation window reaches the target number of moving times. In the case where there is a target sub-image in the segmentation result of the target image whose page size does not meet the specified size, adjust the target sub-image whose page size does not meet the specified size to the specified size.
6. The text detection method according to claim 3, characterized in that, Before moving the segmentation window on the target image based on the preset overlapping ratio to segment the target image, the method further includes: Adjust the resolution of the target image according to the size of the segmented image.
7. The text detection method according to claim 6, characterized in that, The adjusting the resolution of the target image according to the size of the segmented image includes: Calculate a difference value between the length of a first side of the segmented image and the length of a first side of the target image; wherein, the first side of the segmented image is any side in the segmented image, and the first side of the target image is the shorter side in the target image. In the case where the difference value is less than a preset threshold, adjust the resolution of the target image proportionally so that the length of the first side of the segmented image is the same as the length of the adjusted first side of the target image.
8. The text detection method according to claim 1, characterized in that The calculating the probability value of each pixel in the overlapping area according to the probability values of each pixel in the overlapping area in the probability maps corresponding to the multiple target sub-images includes: Determine the magnitudes and the number of a plurality of probability values corresponding to a first pixel in the probability maps corresponding to the multiple target sub-images; wherein, the first pixel is any pixel in the overlapping area. Determine the probability value of the first pixel by voting according to the relationship between the plurality of probability values corresponding to the first pixel and a preset probability threshold. Or, Calculate the average value of the plurality of probability values, and use the average value as the probability value of the first pixel.
9. The text detection method according to claim 1, characterized in that The determining the text area in the target image according to the probability map of the target image includes: According to the probability map of the target image, determine the probability value of each pixel in the target image. According to the probability values of each pixel in the target image, determine the coordinate values of the text area in the target image. According to the coordinate values of the text area in the target image, determine the text area of the target image.
10. A text detection device, characterized in that Includes: A segmentation model, configured to segment a target image to be detected to obtain a plurality of target sub-images; wherein, there is an overlapping area between adjacent target sub-images among the plurality of target sub-images. An input module, configured to input the plurality of target sub-images into a text detection model to obtain probability maps corresponding to the plurality of target sub-images. A merging module, configured to merge the probability maps corresponding to the plurality of target sub-images to obtain a probability map of the target image. A determining module, configured to determine the text area in the target image according to the probability map of the target image. Among them, the probability map is used to describe the probability value that the content of the pixel in the image corresponding to the probability map is text; The merging module is further configured to: Determine the overlapping area and non-overlapping area in the multiple target subgraphs; According to the probability values of each pixel in the overlapping area in the probability maps corresponding to the multiple target subgraphs, calculate the probability value of each pixel in the overlapping area to obtain the probability map corresponding to the overlapping area; According to the probability maps corresponding to the multiple target subgraphs, obtain the probability map corresponding to the non-overlapping area; Merge the probability maps corresponding to the non-overlapping areas in the multiple target subgraphs and the probability map corresponding to the overlapping area in the multiple target subgraphs to obtain the probability map of the target image.
11. An electronic device, characterized in that Including: A processor, a memory and a bus, where The processor and the memory complete communication with each other through the bus; The memory stores program instructions executable by the processor, and the processor can execute the steps of the text detection method according to any one of claims 1 to 9 by invoking the program instructions.
12. A non-transitory computer-readable storage medium, characterized in that The non-transitory computer-readable storage medium stores computer instructions, and the computer instructions cause the computer to execute the steps of the text detection method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Character recognition method and device thereof
CN104951741A
Character region identifying method and device, and computer-readable storage medium
CN108717542A