Image Processing Method, Apparatus, Computer Device, Storage Medium and Program Product
By combining the text response diagram and text gap response diagram of the target image binary image, the candidate area binary image is obtained and the segmentation line is determined, which solves the problem of text box becoming smaller and individual text being disconnected caused by corrosion and expansion operations, and improves the accuracy of the extraction of the text detection box and the accuracy of subsequent text recognition.
Patent Information
- Application Number
- CN202210761538.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-29
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-06-29
AI Technical Summary
In the existing optical character recognition technology, corrosion expansion operations can easily cause the text box to become smaller and the single text to be disconnected, affecting the accuracy of text area detection.
By combining the text response map and text gap response map of the target image, a binary map of the candidate area is obtained, an initial detection box is extracted, and the entry and exit pixels are obtained from it, the division line between two adjacent lines of text is determined, and the text detection box is then determined.
It improves the extraction accuracy of the text detection box, thereby improving the accuracy of subsequent text recognition, and avoids the problems caused by expansion and corrosion of pixels in the text area.
Smart Images

Figure CN115063823B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the technical field of optical character recognition, and in particular, to an image processing method, apparatus, computer device, storage medium, and program product. Background Art
[0002] OCR (Optical Character Recognition) technology is an important branch of computer vision and can be technically divided into two aspects: text detection and text recognition.
[0003] Among them, text detection aims to detect text regions in an image by processing the image. In related technologies, usually after performing erosion and dilation operations on an image containing text, a text detection frame is extracted from the image.
[0004] However, erosion and dilation operations are prone to problems such as the text box becoming smaller and individual characters being disconnected, which in turn affects the accuracy of text region detection. Summary of the Invention
[0005] Embodiments of the present application provide an image processing method, apparatus, computer device, storage medium, and program product. The technical solutions are as follows:
[0006] According to one aspect of the present application, an image processing method is provided, and the method includes:
[0007] Identify text regions and text gap regions in a target image to obtain a text response map and a text gap response map of the target image; the text response map is used to indicate the regions where text is located in the target image; the text gap response map is used to indicate the regions where text gaps are located in the target image;
[0008] Perform binary merging on the text response map and the text gap response map to obtain a candidate region binary map of the target image;
[0009] Extract an initial detection frame from the candidate region binary map; the initial detection frame is a rectangular frame corresponding to having at least one line of text in the target image;
[0010] Obtain each entry pixel and each exit pixel from the initial detection frame; the entry pixel is the pixel where the upper boundary of the text is located, and the exit pixel is the pixel where the lower boundary of the text is located;
[0011] Determine a dividing line between adjacent lines of text according to each of the entry pixels and each of the exit pixels;
[0012] Determine a text detection box in the candidate region binary map according to the dividing line between two adjacent lines of text; the text detection box is a rectangular box corresponding to a single line of text in the target image.
[0013] According to another aspect of the present application, there is provided an image processing apparatus, the apparatus comprising:
[0014] An identification module, configured to identify a text region and a text gap region in a target image, and obtain a text response map and a text gap response map of the target image; the text response map is used to indicate the region where the text in the target image is located; the text gap response map is used to indicate the region where the text gap in the target image is located;
[0015] A binarization processing module, configured to perform binarization merging on the text response map and the text gap response map to obtain a candidate region binary map of the target image;
[0016] An initial detection box extraction module, configured to extract an initial detection box from the candidate region binary map; the initial detection box is a rectangular box corresponding to at least one line of text in the target image;
[0017] An entrance and exit pixel acquisition module, configured to acquire each entrance pixel and each exit pixel from the initial detection box; the entrance pixel is the pixel where the upper boundary of the text is located, and the exit pixel is the pixel where the lower boundary of the text is located;
[0018] A dividing line determination module, configured to determine a dividing line between two adjacent lines of text according to each of the entrance pixels and each of the exit pixels;
[0019] A text detection box determination module, configured to determine a text detection box in the candidate region binary map according to the dividing line between two adjacent lines of text; the text detection box is a rectangular box corresponding to a single line of text in the target image.
[0020] According to another aspect of the present application, there is provided a computer device, the computer device comprising a processor and a memory, and at least one computer instruction is stored in the memory, and the computer instruction is loaded and executed by the processor to implement the method for displaying content provided in the embodiments of the present application.
[0021] According to another aspect of the present application, there is provided a computer-readable storage medium, and at least one computer instruction is stored in the storage medium, and the computer instruction is loaded and executed by a processor to implement the method for displaying content provided in the embodiments of the present application.
[0022] According to another aspect of the present application, a computer program product is provided. The computer program product includes computer instructions stored in a computer-readable storage medium. The computer instructions are read and executed by a processor of a computer device, so that the computer device executes the display content method provided in the implementation of the present application.
[0023] The beneficial effects brought by the technical solutions provided in the embodiments of the present application may include:
[0024] By binarizing and merging the text response map and the text gap response map of the target image to obtain a candidate region binary map, determining an initial detection box from the candidate region binary map, determining the pixels where the upper boundary of the text in the target image is located and the pixels where the lower boundary is located from the initial detection box, and then determining the dividing line between adjacent two lines of text, and determining a text detection box in the candidate region binary map according to the dividing line; the above solution realizes the segmentation of the text detection box by detecting the dividing line between adjacent two lines of text in the image, can avoid the problems caused by dilating and eroding the pixels of the text region, improves the extraction accuracy of the text detection box, and further improves the accuracy of subsequent text recognition. Description of the Drawings
[0025] In order to more clearly introduce the technical solutions in the embodiments of the present application, the drawings required to be used in the description of the embodiments of the present application will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained according to these drawings without creative efforts.
[0026] Figure 1 is a block diagram of a structure of a terminal provided in an exemplary embodiment of the present application;
[0027] Figure 2 is another block diagram of a structure of a terminal provided in an exemplary embodiment of the present application;
[0028] Figure 3 is a schematic diagram of an unfolded state of a terminal provided in an exemplary embodiment of the present application;
[0029] Figure 4 is a schematic diagram of a folded state of a terminal provided in an exemplary embodiment of the present application;
[0030] Figure 5 is a flowchart of an image processing method provided in an exemplary embodiment of the present application;
[0031] Figure 6 is a flowchart of an image processing method provided in another exemplary embodiment of the present application;
[0032] Figure 7 is Figure 6 a schematic diagram of the text image to be detected involved in the illustrated embodiment;
[0033] Figure 8 is Figure 6 a schematic diagram of the text response diagram involved in the illustrated embodiment;
[0034] Figure 9 is Figure 6 a schematic diagram of the text gap response diagram involved in the illustrated embodiment;
[0035] Figure 10 is Figure 6 a schematic diagram of the initial detection box involved in the illustrated embodiment;
[0036] Figure 11 is Figure 6 a schematic diagram of the traversal result involved in the illustrated embodiment;
[0037] Figure 12 is Figure 6 a schematic diagram of the dividing line involved in the illustrated embodiment;
[0038] Figure 13 is Figure 6 a schematic diagram of the text detection box involved in the illustrated embodiment;
[0039] Figure 14 is Figure 6 a schematic diagram of the execution flow of the text detection box involved in the illustrated embodiment;
[0040] Figure 15 It is a structural block diagram of an image processing device provided by an exemplary embodiment of the present application. Specific embodiments
[0041] To make the objectives, technical solutions and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.
[0042] When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0043] In the description of this application, it should be understood that terms such as "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance. In the description of this application, it should be noted that unless otherwise clearly specified and defined, the terms "connected" and "coupled" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium. For those of ordinary skill in the art, the specific meanings of the above terms in this application can be understood according to specific circumstances. In addition, in the description of this application, unless otherwise stated, "a plurality of" means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships can exist. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.
[0044] For the ease of understanding the solutions shown in the embodiments of this application, several nouns that appear in the embodiments of this application are introduced below.
[0045] OCR (Optical Character Recognition) technology is an important branch of computer vision and can be technically divided into two aspects: text detection and text recognition. Text detection aims to detect the text regions in an image by processing the image. Text recognition aims to process the text image to obtain the corresponding text content, usually outputting one line of content for one line of text. If the text regions output by text detection are inaccurate or include multiple lines of text, the text recognition model will not be able to obtain accurate text content. Therefore, how to obtain single-line text regions is the key to text detection. The text detection method based on image segmentation is becoming the main method for text detection because it can segment each character and can solve the problem of curved text detection through corresponding post-processing. However, more or less, there will be a problem of "multiple-line adhesion" of text in the segmented image, which seriously affects the accuracy of text recognition.
[0046] The solutions provided in the various embodiments of this application can be used to solve the above "multiple-line adhesion" problem of text while avoiding dilation and erosion of the pixels in the text region, improving the extraction accuracy of the text detection box, and thus improving the accuracy of subsequent text recognition.
[0047] The solutions shown in the subsequent embodiments of this application can be executed by a computer device, which can be a terminal or a server. Taking the above computer device as a terminal as an example, refer to Figure 1 and Figure 2As shown, it shows a structural block diagram of a terminal 100 provided by an exemplary embodiment of the present application. The terminal 100 can be a smart phone, a tablet computer, an e-book, etc. The terminal 100 in the present application can include one or more of the following components: a processor 110, a memory 120, and a display screen 130.
[0048] The processor 110 may include one or more processing cores. The processor 110 connects various parts within the entire terminal 100 using various interfaces and lines. By running or executing instructions, programs, code sets, or instruction sets stored in the memory 120, and by calling data stored in the memory 120, it executes various functions of the terminal 100 and processes data. Optionally, the processor 110 can be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 110 can integrate a combination of one or several of a central processing unit (CPU), a graphics processing unit (GPU), and a modem, etc. Among them, the CPU mainly processes the operating system, user interface, application programs, etc.; the GPU is responsible for the rendering and drawing of the content to be displayed on the touch display screen 130; the modem is used to process wireless communications. It can be understood that the above modem may not be integrated into the processor 110 and can be implemented separately by a single chip.
[0049] The memory 120 can include random access memory (RAM) and can also include read-only memory (ROM). Optionally, the memory 120 includes a non-transitory computer-readable storage medium. The memory 120 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 120 can include a program storage area and a data storage area. Among them, the program storage area can store instructions for implementing the operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing each of the following method embodiments, etc.; the data storage area can store data created according to the use of the terminal 100 (such as audio data, phone book, etc.).
[0050] Taking the Android operating system as an example, the programs and data stored in the memory 120 are as follows Figure 1 As shown, the memory 120 stores a Linux kernel layer 220, a system runtime library layer 240, an application framework layer 260, and an application layer 280. The Linux kernel layer 220 provides underlying drivers for various hardware of the terminal 100, such as display drivers, audio drivers, camera drivers, Bluetooth drivers, Wi-Fi drivers, power management, etc. The system runtime library layer 240 provides main feature support for the Android system through some C / C++ libraries. For example, the SQLite library provides database support, the OpenGL / ES library provides 3D drawing support, and the Webkit library provides browser kernel support, etc. An Android Runtime is also provided in the system runtime library layer 240, which mainly provides some core libraries that allow developers to write Android applications using the Java language. The application framework layer 260 provides various APIs (Application Programming Interfaces) that may be used when building application programs. Developers can also build their own application programs by using these APIs, such as activity management, window management, view management, notification management, content provider, package management, call management, resource management, location management. At least one application program runs in the application layer 280. These application programs can be the contact program, SMS program, clock program, camera application, etc. that come with the operating system; they can also be application programs developed by third-party developers, such as instant messaging programs, photo beautification programs, etc.
[0051] Taking the IOS operating system as an example, the programs and data stored in the memory 120 are as follows Figure 2As shown in the figure, the IOS system includes: the Core OS layer 320, the Core Services layer 340, the Media layer 360, and the Cocoa Touch Layer 380. The Core OS layer 320 includes the operating system kernel, drivers, and underlying program frameworks, which provide functions closer to the hardware for the program frameworks in the Core Services layer 340 to use. The Core Services layer 340 provides the system services and / or program frameworks required by applications, such as the Foundation framework, the Accounts framework, the Advertising framework, the Data Storage framework, the Network Connection framework, the Location framework, the Motion framework, and so on. The Media layer 360 provides interfaces related to audio-visual aspects for applications, such as interfaces related to graphics and images, interfaces related to audio technology, interfaces related to video technology, and the AirPlay interface for wireless playback of audio and video transmission technology. The Cocoa Touch Layer 380 provides various commonly used interface-related frameworks for application development, and the Cocoa Touch Layer 380 is responsible for the touch interaction operations of users on the terminal 100. Such as local notification services, remote push services, advertising frameworks, game tool frameworks, message user interface (User Interface, UI) frameworks, the UIKit framework for the user interface, map frameworks, and so on.
[0052] In Figure 2 Among the frameworks shown, the frameworks related to most applications include, but are not limited to: the Foundation framework in the Core Services layer 340 and the UIKit framework in the Cocoa Touch Layer 380. The Foundation framework provides many basic object classes and data types, providing the most basic system services for all applications, regardless of the UI. The classes provided by the UIKit framework are the basic UI class libraries, used to create touch-based user interfaces. iOS applications can provide the UI based on the UIKit framework, so it provides the infrastructure for applications to build user interfaces, draw, process, and handle user interaction events, respond to gestures, and so on.
[0053] Taking the display screen 130 as a foldable display screen as an example, the foldable display screen is a screen with a folding function, used to display the user interfaces of various applications; when the foldable display screen also has a touch function, it is also used to receive touch operations of users using any suitable objects such as fingers and styli on or near it.
[0054] Optionally, the foldable display screen includes a first display area 131 and a second display area 132, and in the unfolded state, as Figure 3 shown, the first display area 131 and the second display area 132 are in the same plane; in the folded state, as Figure 4As shown, the first display area 131 and the second display area 132 are in different planes.
[0055] Optionally, at least one other component is also provided in the terminal 100. The at least one other component includes: a camera, a fingerprint sensor, a proximity light sensor, a distance sensor, etc. In some embodiments, the at least one other component is provided on the front, side, or back of the terminal 100. For example, the fingerprint sensor is provided on the back cover or side, and the camera is provided on one side of the display screen 130.
[0056] In some optional embodiments, an edge touch sensor is provided on a single side, or two sides (such as the left and right sides), or four sides (such as the upper, lower, left, and right sides) of the middle frame of the terminal 100. The edge touch sensor is used to detect at least one of the touch operations, click operations, press operations, and slide operations of the user on the middle frame. The edge touch sensor can be any one of a touch sensor, a sensor, a pressure sensor, etc. The user can apply an operation on the edge touch sensor to control the application programs in the terminal 100.
[0057] In addition, those skilled in the art can understand that the structure of the terminal 100 shown in the above drawings does not constitute a limitation on the terminal 100. The terminal may include more or fewer components than shown in the drawings, or combine certain components, or have different component arrangements. For example, the terminal 100 also includes components such as a radio frequency circuit, an input unit, an audio circuit, a Wireless Fidelity (WiFi) module, a power supply, and a Bluetooth module, which will not be elaborated here.
[0058] Please refer to Figure 5 , which is a flowchart of an image processing method provided by an exemplary embodiment of the present application. The image processing method can be applied to the computer device shown above. In Figure 5 , the image processing method includes:
[0059] Step 501, identify the text area and the text gap area in the target image to obtain a text response map and a text gap response map of the target image; the text response map is used to indicate the area where the text in the target image is located; the text gap response map is used to indicate the area where the text gap in the target image is located.
[0060] Among them, the above text response map and text gap response map can specifically be a heat map, a grayscale map, a binary map, etc.
[0061] Step 502, perform binary merging on the text response map and the text gap response map to obtain a candidate area binary map of the target image.
[0062] In an embodiment of the present application, after a computer device obtains a text response map and a text gap response map that respectively indicate the positions of text and text gaps, the text response map and the text gap response map can be merged, and the resulting merged result is a binary map. This binary map uses two different pixel values to distinguish the positions of text + text gaps and other positions in the map.
[0063] Step 503: Extract the initial detection boxes in the candidate region binary map; the initial detection boxes are rectangular boxes corresponding to at least one line of text in the target image.
[0064] After using the candidate region binary map to indicate the positions of text + text gaps and other positions, the computer device can extract the rectangular boxes containing text + text gaps in the candidate region binary map as the initial detection boxes.
[0065] Step 504: Obtain each entry pixel and each exit pixel from the initial detection boxes; the entry pixel is the pixel where the upper boundary of the text is located, and the exit pixel is the pixel where the lower boundary of the text is located.
[0066] In the image, each text occupies a region in the image. In this step, the computer device can locate the pixels where the upper and lower boundaries of each text are located in the initial detection box, and use this as the basis for dividing different text lines.
[0067] Step 505: Determine the dividing line between adjacent lines of text according to each entry pixel and each exit pixel.
[0068] In an embodiment of the present application, after locating the pixels where the upper and lower boundaries of each text are located in the initial detection box, the computer device can determine the dividing line between adjacent lines of text according to the pixels where the upper and lower boundaries of each text are located. That is to say, this dividing line does not pass through the pixels of the same line of text.
[0069] Step 506: Determine text detection boxes in the candidate region binary map according to the dividing line between adjacent lines of text; the text detection boxes are rectangular boxes corresponding to single lines of text in the target image.
[0070] In an embodiment of the present application, the above-mentioned text detection boxes are rectangular boxes containing single lines of text. That is to say, each text detection box contains only one line of text.
[0071] In summary, the solution shown in the embodiments of the present application merges the text response map and the text gap response map of the target image through binarization to obtain a candidate region binary map, determines the initial detection box from the candidate region binary map, determines the pixels where the upper boundary and the lower boundary of the text in the target image are located through the opening operation of the initial detection box, and then determines the dividing line between adjacent two lines of text, and determines the text detection box in the candidate region binary map according to the dividing line; the above solution realizes the segmentation of the text detection box by detecting the dividing line between adjacent two lines of text in the image, can avoid the problems caused by the dilation and erosion of the text region pixels, improves the extraction accuracy of the text detection box, and further improves the accuracy of subsequent text recognition.
[0072] Please refer to Figure 6 , which is a flowchart of an image processing method provided by an exemplary embodiment of the present application. This image processing method can be applied to the computer device shown above. In Figure 6 , this image processing method includes:
[0073] Step 601, identify the text region and the text gap region in the target image, and obtain the text response map and the text gap response map of the target image.
[0074] Among them, the text response map is used to indicate the region where the text in the target image is located; the text gap response map is used to indicate the region where the text gap in the target image is located.
[0075] Taking the above text response map and text gap response map as heatmaps as an example, in the embodiments of the present application, the computer device can perform pre-processing on the input text image to be detected (i.e., the above target image, and the schematic diagram of the text image to be detected is as Figure 7 shown), perform pre-processing on it (including but not limited to operations such as image scaling, normalization, standardization, dilation, erosion, etc.), input it into the trained text detection network, and obtain the text response heatmap TextMap and the text gap response heatmap LinkMap. The schematic diagram of the text response heatmap TextMap is as Figure 8 shown, and the schematic diagram of the text gap response heatmap LinkMap is as Figure 9 shown.
[0076] Among them, the above text detection network is a machine learning model that takes the target image after pre-processing as the input and TextMap and LinkMap as the output. Among them, the text response heatmap is a heatmap with high response centered on the text, while the text gap response heatmap is a heatmap with high response centered on the gap between texts.
[0077] Step 602: Binarize the text response map according to the first pixel value threshold to obtain a text response binary map; binarize the text gap response map according to the second pixel value threshold to obtain a text gap response binary map.
[0078] Set Threshold 1 (i.e., the above-mentioned first pixel threshold) and perform threshold segmentation on the TextMap. Among them, pixels with pixel values less than Threshold 1 are set to 0, and pixels with pixel values not less than Threshold 1 are set to 1 to obtain a binary TextMap. Set Threshold 2 (i.e., the above-mentioned second pixel threshold) and perform threshold segmentation on the LinkMap. Among them, pixels with pixel values less than Threshold 2 are set to 0, and pixels with pixel values not less than Threshold 2 are set to 1 to obtain a binary LinkMap.
[0079] Among them, the pixel meanings of the above two maps are different. Among them, TextMap is the response map of the text center, and LinkMap is the response map of the text gap. Therefore, after binarization, in the TextMap, pixels with a pixel value of 1 represent the center of a text, and pixels with a pixel value of 0 indicate that the pixel is not the center of a text; while in the LinkMap, pixels with a pixel value of 1 represent the area between two texts, and the area with a pixel value of 0 represents the area that is not the text gap area.
[0080] Step 603: Take the union of the text response binary map and the text gap response binary map to obtain a candidate region binary map.
[0081] In the embodiment of the present application, in the candidate region binary map, when the pixel values of a certain pixel in the text response binary map and the text gap response binary map are both 0, the pixel value of the pixel in the candidate region binary map is 0; when the pixel value of the pixel in either the text response binary map or the text gap response binary map is 1, the pixel value of the pixel in the candidate region binary map is 1.
[0082] In the embodiment of the present application, after taking the union of the TextMap and the LinkMap, there will be a certain connection between the text center and the text gap. Through post-processing (such as dilation), the text in a row can be connected together.
[0083] Step 604: Extract the initial detection box in the candidate region binary map; the initial detection box is a rectangular box corresponding to at least one row of text in the target image.
[0084] In a possible implementation manner, extracting the initial detection box in the candidate region binary map includes:
[0085] Obtain the connected regions with an extraction value of 1 in the candidate region binary map to obtain each second connected region;
[0086] Eliminate the interference regions in each second connected region;
[0087] Take the minimum bounding rectangle of the remaining second connected regions after removing the interference regions to obtain the initial detection frame.
[0088] Wherein, the above-mentioned connected region refers to a region in which all the included pixels have the same value.
[0089] The above-mentioned interference regions can be regions where text cannot be accurately recognized, such as regions with small areas or regions with blurred text.
[0090] In the embodiments of the present application, the computer device can perform connectivity analysis on the candidate region binary map, extract individual connected regions, and then perform corresponding interference region removal operations (optionally calculating the area, the highest gray value in the corresponding TextMap region within the current connected region, etc.). Finally, calculate the minimum bounding rectangle for each connected region to obtain the initial detection frame. The schematic diagram of the initial detection frame 1001 can be as Figure 10 shown.
[0091] Wherein, the above-mentioned removal operation of calculating the area includes: calculating all connected regions, calculating the areas of each connected region, setting an area threshold, and if the area of a connected region is less than this threshold, then this connected region is considered an interference region and is removed. That is to say, in the embodiments of the present application, the computer device can discard the regions with smaller areas (usually not belonging to the text region or difficult to recognize text) among the above-mentioned second connected regions, and retain the regions with larger areas and easy-to-recognize text.
[0092] The above-mentioned removal operation based on the highest gray value in the corresponding TextMap region within the current connected region includes: calculating the highest gray value within each connected region, setting a gray value threshold, and if the highest gray value within the current connected region is less than this threshold, then this connected region is considered an interference region and is removed. That is to say, in the embodiments of the present application, the computer device can discard the regions with lower gray values (possibly difficult to recognize text) among the above-mentioned second connected regions, and retain the regions with larger gray values.
[0093] Step 605, obtain each entrance pixel and each exit pixel from the initial detection frame; the entrance pixel is the pixel where the upper boundary of the text is located, and the exit pixel is the pixel where the lower boundary of the text is located.
[0094] In a possible implementation manner, perform an opening operation on the region where the initial detection frame is located in the candidate region binary map to obtain each entrance pixel and each exit pixel, including:
[0095] Perform an opening operation on the region where the initial detection box is located in the candidate region binary map to obtain a dilation result map; the opening operation includes an operation of eroding first and then dilating; for example, the computer device can first erode the pixels in the initial detection box in the candidate region binary map, and then dilate the pixels in the initial detection box in the candidate region binary map through a structuring element to obtain a dilation result map;
[0096] Traverse the pixels in the dilation result map in the column direction to obtain each entrance pixel and each exit pixel; the column direction is the extension direction of the short side of the initial detection box.
[0097] In the embodiment of the present application, the computer device can first perform an opening operation on the pixels in the initial detection box in the candidate region binary map, and then, in the column direction, find the pixels on the upper and lower edges of the text through the method of image morphological processing.
[0098] Image dilation refers to expanding the boundary points of a binary object, merging all background points in contact with the object into the object, and expanding the boundary to the outside. Image dilation can be implemented through Matlab: that is, adding pixels to the object boundary in the image. In the operation, the state of a given pixel in the output image is determined by using certain rules for the corresponding pixel and its neighborhood in the input image. During the dilation operation, the output pixel value is the maximum value of each pixel value in the neighborhood of the corresponding pixel in the input image. In a binary image, if any pixel value in the neighborhood of a certain pixel is 1, then the output pixel value of the corresponding pixel is 1. Image dilation can be performed using the imdilate function, and the imdilate function requires two basic input parameters, namely the input image to be processed and the structuring element (also called the structuring element).
[0099] In a possible implementation manner, traversing the pixels in the dilation result map in the column direction to obtain each entrance pixel and each exit pixel includes:
[0100] For the first pixel in the dilation result map, when the pixel value of the first pixel is 0 and the pixel value of the next pixel in the column direction of the first pixel is 1, the first pixel is determined as an entrance pixel; the first pixel is any pixel in the dilation result map;
[0101] For the second pixel in the dilation result map, when the pixel value of the second pixel is 1 and the pixel value of the next pixel in the column direction of the second pixel is 0, the second pixel is determined as an exit pixel; the second pixel is any pixel in the dilation result map.
[0102] In the embodiments of the present application, traversing the pixels in the dilation result map in the column direction means that after traversing each pixel in the previous column in the column direction in sequence, then traversing each pixel in the next column. Among them, before the dilation operation, the computer device may also first perform an erosion operation on the pixels in the initial detection frame.
[0103] In a possible implementation, the structuring element is a linear structuring element.
[0104] In the embodiments of the present application, in addition to using a linear structuring element, the above structuring element may also use other types of structuring elements, such as a diamond structuring element, a rectangular structuring element, and so on.
[0105] In a possible implementation, the angle and size of the structuring element are the same as the angle and size of the initial detection frame.
[0106] In the embodiments of the present application, a linear structuring element with the same angle and size as the initial detection frame can be used to perform the dilation operation.
[0107] Among them, the problem of text line adhesion may occur in the initial candidate detection frame, such as Figure 10 As shown, the two lines of text "Qinhuangdao Station" and "Qinhuangdao" are framed in an initial detection frame.
[0108] Considering that line adhesion is mainly caused by taking the union of the binary TextMap graph and the binary LinkMap graph, because the union of the two is required to connect the text in the same line into a line. Therefore, the solution shown in the present application is processed based on the binary graph of the candidate region of the text image. First, the computer device selects a structuring element with an appropriate size and shape to perform an opening operation on the binary graph of the candidate region. An optional solution is to first estimate the angular direction, size, etc. of the current initial detection frame, calculate a linear structuring element with the same angle and size based on these elements, and perform an opening operation on the binary graph of the candidate region based on the linear structuring element. Since the opening operation will break the text in the same line, a linear structuring element with the same angular direction as the initial detection frame can be selected to perform the dilation operation, and the result is denoted as I.
[0109] To break different line directions, in the embodiments of the present application, it can be assumed that the long direction of the text is the same line direction, and the short direction is the different line direction (i.e., the above-mentioned column direction). Therefore, it is necessary to cut and break it in the short direction. An optional solution is to traverse and scan along the short direction (assuming that the short side direction of the initial detection frame is the column direction and the long side direction is the row direction. Here, traversing and scanning along the short direction means scanning one column and then the next column, and the purpose is to break in the column direction). If the detection frame has M rows and N columns of pixels, and M > N, then the short side direction is N columns of pixels. Then, N M*1 image pixels can be scanned. If the i-th pixel is 0 and the i+1-th pixel is 1, then the current pixel (the i-th pixel) is recorded as the "entrance pixel". Please see Figure 11 In the schematic diagram of the traversal result shown, for pixel 1101, if the i-th pixel is 1 and the i+1-th pixel is 0, then the current pixel is recorded as the "exit pixel". Please see Figure 11 pixel 1102 in Figure 11 In
[0110] Step 606, determine the dividing line between adjacent two rows of text according to each entrance pixel and each exit pixel.
[0111] In a possible implementation manner, determining the dividing line between adjacent two rows of text according to each entrance pixel and each exit pixel includes:
[0112] Merge each entrance pixel in the row direction to obtain each entrance band; the row direction is the extending direction of the long side of the initial detection frame;
[0113] Merge each exit pixel in the row direction to obtain each exit band;
[0114] Sort each entrance band and each exit band according to the minimum value of the vertical coordinates of the pixels included in each entrance band and each exit band respectively; the vertical coordinate is the coordinate in the extending direction of the short side of the initial detection frame;
[0115] Remove the first entrance band and the last exit band according to the sorting order;
[0116] Determine the dividing line between adjacent two rows of text according to the vertical coordinate intervals of the remaining entrance bands and the remaining exit bands.
[0117] In the embodiments of the present application, an entrance band is equivalent to the set of the upper boundaries of each character in the same row of text. Correspondingly, an exit band is equivalent to the set of the lower boundaries of each character in the same row of text. Obtaining the entrance band and the exit band is equivalent to obtaining the upper and lower boundaries of the characters in each row of text. The computer device can determine the dividing line between each row of text according to the upper and lower boundaries of the characters in each row of text.
[0118] In a possible implementation, each entry pixel is merged in the row direction to obtain each entry band, including:
[0119] Adding the entry pixels in which the difference between the vertical coordinates is less than the first difference threshold among all the entry pixels into the same entry band;
[0120] Each exit pixel is merged in the row direction to obtain each exit band, including:
[0121] Adding the exit pixels in which the difference between the vertical coordinates is less than the second difference threshold among all the exit pixels into the same exit band.
[0122] In the embodiments of the present application, when the column direction is the vertical coordinate direction, for two entry pixels, if the difference in the vertical coordinates of these two entry pixels is small (i.e., the difference is less than a certain threshold), it can be considered that these two entry pixels are the pixels where the upper boundaries of the characters in the same row of text are located. At this time, it can be determined that these two entry pixels belong to the same entry band.
[0123] Similarly, when the column direction is the vertical coordinate direction, for two exit pixels, if the difference in the vertical coordinates of these two exit pixels is small (i.e., the difference is less than a certain threshold), it can be considered that these two exit pixels are the pixels where the lower boundaries of the characters in the same row of text are located. At this time, it can be determined that these two exit pixels belong to the same exit band.
[0124] In a possible implementation, the dividing line between two adjacent rows of text is determined according to the vertical coordinate intervals of the remaining entry band and the remaining exit band, including:
[0125] Obtaining the minimum value of the vertical coordinates of the pixels included in the first exit band and the maximum value of the vertical coordinates of the pixels included in the first entry band; the first exit band and the first entry band belong to the remaining entry band and the remaining exit band; the first exit band and the first entry band are adjacent, and the first exit band is located above the first entry band; the above "above" is the positive direction of the pixel coordinate system, and the pixel coordinate system is a coordinate system with the extension direction of the long side of the initial detection frame as the abscissa and the extension direction of the short side as the vertical coordinate;
[0126] Taking the middle vertical coordinate between the minimum value of the vertical coordinates of the pixels included in the first exit band and the maximum value of the vertical coordinates of the pixels included in the first entry band;
[0127] Obtaining a dividing line as a straight line passing through the middle vertical coordinate and having the same abscissa.
[0128] In an embodiment of the present application, for adjacent entrance bands and exit bands, when the exit band is located above the entrance band, it indicates that there is a dividing line for two lines of text between the exit band and the entrance band. At this time, the exit band represents the lower boundary of the upper line of text, and the entrance band represents the upper boundary of the lower line of text. By the minimum value of the vertical coordinate of the exit band (i.e., the lowest point of the lower boundary of the upper line of text) and the maximum value of the vertical coordinate of the entrance band (i.e., the highest point of the upper boundary of the lower line of text), the range of the vertical coordinate of the dividing line for the two lines of text can be determined. Taking a straight line parallel to the line direction within this range can be used as the dividing line between the above two lines of text.
[0129] In a possible implementation, the middle vertical coordinate is the average value of the minimum value of the vertical coordinates of the pixels included in the first exit band and the maximum value of the vertical coordinates of the pixels included in the first entrance band.
[0130] In a possible implementation of the embodiment of the present application, the computer device can take the average value of the lowest point of the lower boundary of the upper line of text and the vertical coordinate of the lowest point of the lower boundary of the upper line of text as the vertical coordinate of the dividing line.
[0131] Optionally, the computer device can also take other vertical coordinates between the lowest point of the lower boundary of the upper line of text and the lower boundary of the upper line of text as the vertical coordinate of the dividing line, and the embodiment of the present application does not limit this.
[0132] Merge all "entrance" pixels to obtain pixels of the same edge "entrance", and the same applies to all "exit" pixels. Figure 11 In this way, 2 "exit bands" 1103 and two "entrance bands" 1104 will be obtained. Among them, the above-mentioned merging of pixels can form a set with "entrance pixels" whose y values have intersections with each other with a floating threshold (optional value is 1), and the same applies to the merging of exit pixels.
[0133] In an embodiment of the present application, the computer device can sort the "entrance bands" and "exit bands", and the sorting rule can be: sort in descending order according to the minimum vertical coordinates of the pixels in each entrance band and exit band.
[0134] Then it is necessary to find the dividing line among these "bands". First, discard the first "entrance band" and the last "exit band", and traverse the "bands" in turn to calculate the dividing line between the "exit band" and the "entrance band". For example, please refer to Figure 12, which shows a schematic diagram of the dividing line involved in the embodiments of the present application. Assume that the interval where the ordinates of the pixels in the "exit band" are located is [y1, y2], and the interval where the ordinates of the pixels in the "entrance band" are located is [y3, y4]. Then the dividing line 1201 is a straight line parallel to the row direction between y2 and y3 in the ordinate. For example, an optional value of the ordinate of the dividing line 1201 is (y2 + y3) / 2.
[0135] Step 607: Determine a text detection frame in the candidate region binary map according to the dividing line between two adjacent lines of text; the text detection frame is a rectangular frame corresponding to a single line of text in the target image.
[0136] In the embodiments of the present application, after obtaining the dividing line between two adjacent lines of text, the computer device can divide the initial detection frame in the candidate region binary map according to the dividing line, that is, the text detection frame corresponding to each line of text in the initial detection frame can be obtained.
[0137] In a possible implementation manner, determining a text detection frame in the candidate region binary map according to the dividing line between two adjacent lines of text includes:
[0138] Cut the dilated result map according to the dividing line between two adjacent lines of text;
[0139] Extract the connected regions with a value of 1 in the cut dilated result map to obtain each first connected region;
[0140] Take the minimum bounding rectangle for each first connected region to obtain the text detection frame.
[0141] In the embodiments of the present application, the computer device can cut the dilated result map I according to the dividing line, then perform connected region detection in each of the cut sub-maps, extract each first connected region, and then calculate the minimum bounding rectangle for the extracted first connected regions, that is, the text detection frame corresponding to each line of text in the initial detection frame can be obtained.
[0142] In the embodiments of the present application, the computer device can use the dividing line to cut in I, then extract each connected region, and calculate the minimum bounding rectangle for each connected region to obtain the final text detection frame. Please refer to Figure 13 , which shows a schematic diagram of the text detection frame involved in the embodiments of the present application. In Figure 13 , the computer device can identify the text detection frame 1301 and the text detection frame 1302, and these two text detection frames respectively contain a line of text.
[0143] After obtaining the above text detection frame, the computer device can extract the regional image corresponding to the area where the text detection frame is located in the target image, and perform OCR recognition on the text in the regional image through a text recognition model.
[0144] Please refer to Figure 14 , which shows a schematic diagram of the recognition process of the text detection frame involved in the embodiments of the present application. As Figure 14 shown, the recognition process of the text detection frame involved in the embodiments of the present application may include the following steps:
[0145] S1401, The computer device obtains the input image and inputs the input image into the text detection network for processing to obtain the text response map and the text gap response map output by the text detection network.
[0146] S1402, Perform post-detection processing according to the text response map and the text gap response map, and obtain the minimum bounding rectangle as the initial detection frame.
[0147] Among them, the execution process of this step S1402 can refer to the description under the above steps 602 to 604, and will not be elaborated here.
[0148] S1403, Structural element setting.
[0149] In this step, the computer device can calculate the direction and size of the initial detection frame, and set a linear structural element according to the direction and size of the initial detection frame.
[0150] S1404, Morphological processing.
[0151] The computer device can perform morphological processing of the opening operation on the pixels (binary pixels) in the initial detection frame according to the set linear structural element, that is, first perform an erosion operation, and then perform a dilation operation according to the linear structural element.
[0152] S1405, Determine the dividing line.
[0153] In this step, the computer device can execute the following sub-steps:
[0154] S1405a, Statistically analyze the entrance pixels and exit pixels based on the short side direction;
[0155] S1405b, Merge the entrance pixels and exit pixels to obtain the entrance band and the exit band;
[0156] S1405c, Determine the dividing line;
[0157] S1405d, Perform connectivity analysis to obtain the minimum bounding rectangle to obtain the text detection frame.
[0158] Among them, the execution processes of the above steps S1405a to S1405d can refer to the descriptions under the above steps 605 to 607, and will not be elaborated here.
[0159] S1406, output a text detection frame.
[0160] After the computer device outputs the text detection frame, subsequently, according to the text detection frame, the region images corresponding to each line of text can be extracted from the input image, and the text in the region images corresponding to each line of text can be recognized by means of OCR recognition.
[0161] The solution shown in this application proposes a method. By processing the output of the text detection model, morphological processing such as opening operation and dilation is performed on the initial text detection frame using a structural element with an angular direction, and the "exit pixels" and "entry pixels" are counted, and they are clustered to obtain the "exit band" and "entry band", thereby determining the dividing line, cutting the text with line adhesion, and obtaining a more refined text detection frame.
[0162] In summary, the solution shown in the embodiments of this application binarizes and merges the text response map and the text gap response map of the target image to obtain a candidate region binary map, determines the initial detection frame from the candidate region binary map, determines the pixels where the upper boundary and the lower boundary of the text in the target image are located by performing an opening operation on the initial detection frame, and then determines the dividing line between adjacent lines of text, and determines the text detection frame in the candidate region binary map according to the dividing line; the above solution realizes the segmentation of the text detection frame by detecting the dividing line between adjacent lines of text in the image, can avoid the problems caused by dilating and eroding the pixels in the text region, improves the extraction accuracy of the text detection frame, and further improves the accuracy of subsequent text recognition.
[0163] The following is an embodiment of the apparatus of this application, which can be used to execute the method embodiment of this application. For the details not disclosed in the embodiment of the apparatus of this application, please refer to the method embodiment of this application.
[0164] Please refer to Figure 15 , which shows a structural block diagram of an image processing apparatus provided by an exemplary embodiment of this application. This image processing apparatus can be used to implement the image processing methods involved in the above various method embodiments. The apparatus includes:
[0165] An identification module 1501, configured to identify the text region and the text gap region in the target image, and obtain the text response map and the text gap response map of the target image; the text response map is used to indicate the region where the text in the target image is located; the text gap response map is used to indicate the region where the text gap in the target image is located;
[0166] The binarization processing module 1502 is configured to perform binarization merging on the text response map and the text gap response map to obtain a binary map of candidate regions of the target image;
[0167] The initial detection box extraction module 1503 is configured to extract an initial detection box from the binary map of candidate regions; the initial detection box is a rectangular box corresponding to at least one line of text in the target image;
[0168] The entrance and exit pixel acquisition module 1504 is configured to acquire each entrance pixel and each exit pixel from the initial detection box; the entrance pixel is the pixel where the upper boundary of the text is located, and the exit pixel is the pixel where the lower boundary of the text is located;
[0169] The dividing line determination module 1505 is configured to determine a dividing line between adjacent lines of text according to each of the entrance pixels and each of the exit pixels;
[0170] The text detection box determination module 1506 is configured to determine a text detection box in the binary map of candidate regions according to the dividing line between adjacent lines of text; the text detection box is a rectangular box corresponding to a single line of text in the target image.
[0171] In a possible implementation manner, the entrance and exit pixel acquisition module 1504 is configured to,
[0172] Perform an opening operation on the region where the initial detection box in the binary map of candidate regions is located to obtain a dilated result map; the opening operation includes an operation of erosion first and then dilation;
[0173] Traverse the pixels in the dilated result map in the column direction to obtain each of the entrance pixels and each of the exit pixels; the column direction is the extending direction of the short side of the initial detection box.
[0174] In a possible implementation manner, the entrance and exit pixel acquisition module 1504 is configured to,
[0175] For a first pixel in the dilated result map, when the pixel value of the first pixel is 0 and the pixel value of the next pixel in the column direction of the first pixel is 1, determine the first pixel as the entrance pixel; the first pixel is any pixel in the dilated result map;
[0176] For a second pixel in the dilated result map, when the pixel value of the second pixel is 1 and the pixel value of the next pixel in the column direction of the second pixel is 0, determine the second pixel as the exit pixel; the second pixel is any pixel in the dilated result map.
[0177] In a possible implementation manner, the structuring element is a linear structuring element.
[0178] In a possible implementation manner, the angle and size of the structuring element are the same as those of the initial detection frame.
[0179] In a possible implementation manner, the text detection frame determination module 1506 is configured to
[0180] Cut the dilated result image according to the dividing line between two adjacent lines of text;
[0181] Extract the connected regions with a value of 1 in the cut dilated result image to obtain each first connected region;
[0182] Take the minimum bounding rectangle for each of the first connected regions to obtain the text detection frame.
[0183] In a possible implementation manner, the dividing line determination module 1505 is configured to
[0184] Merge each of the entrance pixels in the row direction to obtain each entrance band; the row direction is the extending direction of the long side of the initial detection frame;
[0185] Merge each of the exit pixels in the row direction to obtain each exit band;
[0186] Sort each of the entrance bands and each of the exit bands according to the minimum value of the vertical coordinates of the pixels included in each of the entrance bands and each of the exit bands; the vertical coordinate is the coordinate in the extending direction of the short side of the initial detection frame;
[0187] Remove the first entrance band and the last exit band in the sorting order;
[0188] Determine the dividing line between two adjacent lines of text according to the vertical coordinate intervals of the remaining entrance bands and the remaining exit bands.
[0189] In a possible implementation manner, the dividing line determination module 1505 is configured to
[0190] Add the entrance pixels with a difference between vertical coordinates less than a first difference threshold among each of the entrance pixels to the same entrance band;
[0191] Add the exit pixels with a difference between vertical coordinates less than a second difference threshold among each of the exit pixels to the same exit band.
[0192] In a possible implementation manner, the dividing line determination module 1505 is configured to
[0193] Obtain the minimum ordinate of the pixels included in the first exit band and the maximum ordinate of the pixels included in the first entrance band; the first exit band and the first entrance band belong to the remaining entrance bands and the remaining exit bands; the first exit band and the first entrance band are adjacent, and the first exit band is located above the first entrance band; the above is the positive direction of the pixel coordinate system, and the pixel coordinate system uses the extension direction of the long side of the initial detection frame as the abscissa and the extension direction of the short side as the ordinate coordinate system;
[0194] Take the middle ordinate between the minimum ordinate of the pixels included in the first exit band and the maximum ordinate of the pixels included in the first entrance band;
[0195] Obtain a line passing through the middle ordinate and having the same abscissa as a segmentation line.
[0196] In a possible implementation manner, the middle ordinate is the average value of the minimum ordinate of the pixels included in the first exit band and the maximum ordinate of the pixels included in the first entrance band.
[0197] In a possible implementation manner, the binarization processing module 1502 is used to,
[0198] Perform binarization processing on the text response map according to a first pixel value threshold to obtain a text response binary map;
[0199] Perform binarization processing on the text gap response map according to a second pixel value threshold to obtain a text gap response binary map;
[0200] Take the union of the text response binary map and the text gap response binary map to obtain the candidate region binary map.
[0201] In a possible implementation manner, the initial detection frame extraction module 1503 is used to,
[0202] Obtain the connected regions with an extraction value of 1 in the candidate region binary map to obtain each second connected region;
[0203] Eliminate the interference regions in each of the second connected regions;
[0204] Take the minimum bounding rectangle of the remaining second connected regions after eliminating the interference regions to obtain the initial detection frame.
[0205] In summary, the solution shown in the embodiments of the present application binarizes and merges the text response map and the text gap response map of the target image to obtain a candidate region binary map, determines an initial detection frame from the candidate region binary map, determines the pixels where the upper boundary and the lower boundary of the text in the target image are located through an opening operation on the initial detection frame, and then determines the dividing line between adjacent lines of text. The text detection frame is determined in the candidate region binary map according to the dividing line. The above solution realizes the segmentation of the text detection frame by detecting the dividing line between adjacent lines of text in the image, can avoid the problems caused by dilation and erosion of the text region pixels, improves the extraction accuracy of the text detection frame, and further improves the accuracy of subsequent text recognition.
[0206] The embodiments of the present application also provide a computer-readable medium, which stores at least one computer instruction, and the at least one computer instruction is loaded and executed by the processor to implement the image processing method described in each of the above embodiments.
[0207] The embodiments of the present application also provide a computer program product, which includes computer instructions stored in a computer-readable storage medium; the computer instructions are read and executed by a processor of a computer device, so that the terminal executes the image processing method described in each of the above embodiments.
[0208] The embodiments of the present application also provide a computer program, which includes computer instructions stored in a computer-readable storage medium; the computer instructions are read and executed by a processor of a computer device, so that the terminal executes the image processing method described in each of the above embodiments.
[0209] It should be noted that when the device for displaying content provided in the above embodiments executes the method for displaying content, only the above division of each functional module is used as an example for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device for displaying content provided in the above embodiments and the method embodiments for displaying content belong to the same concept, and the specific implementation process can be seen in the method embodiments, which will not be repeated here.
[0210] The serial numbers of the embodiments of the present application above are only for description and do not represent the advantages and disadvantages of the embodiments.
[0211] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above embodiments can be completed by hardware, or can be completed by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and the storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disc, etc.
[0212] The above are only exemplary embodiments that can be implemented in this application, and are not intended to limit this application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of this application shall be included within the protection scope of this application.
Claims
1. An image processing method, characterized in that, The method includes: Identifying a text region and a text gap region in a target image to obtain a text response map and a text gap response map of the target image; the text response map is used to indicate the region where the text in the target image is located; the text gap response map is used to indicate the region where the text gaps in the target image are located; Performing binary merging on the text response map and the text gap response map to obtain a binary map of candidate regions of the target image; Extracting an initial detection box from the binary map of candidate regions; the initial detection box is a rectangular box corresponding to having at least one line of text in the target image; Obtaining each entry pixel and each exit pixel from the initial detection box; the entry pixel is the pixel where the upper boundary of the text is located, and the exit pixel is the pixel where the lower boundary of the text is located; Merging each of the entry pixels in the row direction to obtain each entry band; the row direction is the extending direction of the long side of the initial detection box; Merging each of the exit pixels in the row direction to obtain each exit band; Sorting each of the entry bands and each of the exit bands according to the minimum value of the ordinate of the pixels included in each of the entry bands and each of the exit bands; the ordinate is the coordinate in the extending direction of the short side of the initial detection box; Removing the first entry band and the last exit band in the sorting order; Determining a dividing line between adjacent two lines of text according to the ordinate intervals of the remaining entry bands and the remaining exit bands; Determining a text detection box in the binary map of candidate regions according to the dividing line between adjacent two lines of text; the text detection box is a rectangular box corresponding to a single line of text in the target image.
2. The method according to claim 1, wherein The obtaining each entry pixel and each exit pixel from the initial detection box includes: Performing an opening operation on the region where the initial detection box is located in the binary map of candidate regions to obtain a dilated result map; the opening operation includes an operation of eroding first and then dilating; Traversing the pixels in the dilated result map in the column direction to obtain each of the entry pixels and each of the exit pixels; the column direction is the extending direction of the short side of the initial detection box.
3. The method according to claim 2, characterized in that, The traversing the pixels in the dilated result map in the column direction to obtain each of the entry pixels and each of the exit pixels includes: For a first pixel in the dilated result map, when the pixel value of the first pixel is 0 and the pixel value of the next pixel in the column direction of the first pixel is 1, determining the first pixel as the entry pixel; the first pixel is any pixel in the dilated result map; For a second pixel in the dilated result map, when the pixel value of the second pixel is 1 and the pixel value of the next pixel in the column direction of the second pixel is 0, determining the second pixel as the exit pixel; the second pixel is any pixel in the dilated result map.
4. The method according to claim 2, characterized in that, The determining a text detection box in the binary map of candidate regions according to the dividing line between adjacent two lines of text includes: Cut the dilated result image according to the dividing line between two adjacent lines of text; Extract the connected regions with a value of 1 from the cut dilated result image to obtain each first connected region; Take the minimum bounding rectangle for each of the first connected regions to obtain the text detection frame.
5. The method according to claim 1, wherein, The merging of the respective entrance pixels in the row direction to obtain each entrance band includes: Adding the entrance pixels among the respective entrance pixels, where the difference between the vertical coordinates is less than a first difference threshold, to the same entrance band; The merging of the respective exit pixels in the row direction to obtain each exit band includes: Adding the exit pixels among the respective exit pixels, where the difference between the vertical coordinates is less than a second difference threshold, to the same exit band.
6. The method according to claim 1, characterized in that The determination of the dividing line between two adjacent lines of text according to the vertical coordinate intervals of the remaining entrance bands and the remaining exit bands includes: Obtaining the minimum value of the vertical coordinates of the pixels included in a first exit band and the maximum value of the vertical coordinates of the pixels included in a first entrance band; the first exit band and the first entrance band belong to the remaining entrance bands and the remaining exit bands; the first exit band and the first entrance band are adjacent, and the first exit band is located above the first entrance band; the above is the positive direction of the pixel coordinate system, and the pixel coordinate system is a coordinate system with the extension direction of the long side of the initial detection frame as the abscissa and the extension direction of the short side as the ordinate; Taking the intermediate vertical coordinate between the minimum value of the vertical coordinates of the pixels included in the first exit band and the maximum value of the vertical coordinates of the pixels included in the first entrance band; Obtaining a straight line passing through the intermediate vertical coordinate and having the same abscissa as one of the dividing lines.
7. The method according to claim 6, wherein The intermediate vertical coordinate is the average value of the minimum value of the vertical coordinates of the pixels included in the first exit band and the maximum value of the vertical coordinates of the pixels included in the first entrance band.
8. The method according to claim 1, characterized in that The binary merging of the text response image and the text gap response image to obtain the candidate region binary image of the target image includes: Performing binary processing on the text response image according to a first pixel value threshold to obtain a text response binary image; Performing binary processing on the text gap response image according to a second pixel value threshold to obtain a text gap response binary image; Taking the union of the text response binary image and the text gap response binary image to obtain the candidate region binary image.
9. The method according to claim 1, wherein The extraction of the initial detection frame from the candidate region binary image includes: Obtaining the connected regions with a value of 1 extracted from the candidate region binary image to obtain each second connected region; Eliminating the interference regions in each of the second connected regions; Taking the minimum bounding rectangle for the remaining second connected regions after eliminating the interference regions to obtain the initial detection frame.
10. An image processing apparatus, characterized in that, The device includes: An identification module, configured to identify text regions and text gap regions in a target image, and obtain a text response map and a text gap response map of the target image; the text response map is used to indicate the regions where the text in the target image is located; the text gap response map is used to indicate the regions where the text gaps in the target image are located; A binarization processing module, configured to perform binarization merging on the text response map and the text gap response map to obtain a candidate region binary map of the target image; An initial detection box extraction module, configured to extract an initial detection box from the candidate region binary map; the initial detection box is a rectangular box corresponding to at least one line of text in the target image; An entrance and exit pixel acquisition module, configured to acquire each entrance pixel and each exit pixel from the initial detection box; the entrance pixel is the pixel where the upper boundary of the text is located, and the exit pixel is the pixel where the lower boundary of the text is located; A dividing line determination module, configured to merge each of the entrance pixels in the row direction to obtain each entrance band; the row direction is the extending direction of the long side of the initial detection box; merge each of the exit pixels in the row direction to obtain each exit band; sort each of the entrance bands and each of the exit bands according to the minimum value of the vertical coordinates of the pixels included in each of the entrance bands and each of the exit bands; the vertical coordinate is the coordinate in the extending direction of the short side of the initial detection box; remove the first entrance band and the last exit band according to the sorting order; determine the dividing line between adjacent lines of text according to the vertical coordinate intervals of the remaining entrance bands and the remaining exit bands; A text detection box determination module, configured to determine a text detection box in the candidate region binary map according to the dividing line between adjacent lines of text; the text detection box is a rectangular box corresponding to a single line of text in the target image.
11. A computer device, characterized in that, The computer device includes a processor and a memory, and at least one computer instruction is stored in the memory, and the computer instruction is loaded and executed by the processor to implement the image processing method according to any one of claims 1 to 9.
12. A computer-readable storage medium, characterized in that, At least one computer instruction is stored in the storage medium, and the computer instruction is loaded and executed by a processor to implement the image processing method according to any one of claims 1 to 9.
13. A computer program product, characterized in that, The computer program product includes computer instructions, and the computer instructions are stored in a computer-readable storage medium; the computer instructions are read and executed by a processor of a computer device, so that the computer device executes the image processing method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Document image information extraction method and system
CN111611933A
Text region determination method and device, equipment and readable storage medium
CN113076814A