Text detection method and device, computer equipment and readable storage medium
By combining self-attention processing and segmentation features, enhanced features are determined to improve the accuracy of text detection, solving the problems of false detection and multiple detection in complex scenarios, and achieving higher detection accuracy and robustness.
Patent Information
- Application Number
- CN202510832573.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-09-26
AI Technical Summary
Existing text detection methods are prone to false detection or multiple detection when encountering non-standard formats or poor quality scanned images, making it difficult to achieve accurate segmentation and recognition, affecting the accuracy of text recognition.
By obtaining the features of the target image, self-attention processing is performed, and the enhanced features are determined by combining the segmentation features and attention features. The text detection area is determined based on the enhanced features.
Improved the accuracy and robustness of text detection, especially in complex scenarios, enhancing the accuracy and user convenience of text detection.
Smart Images

Figure CN120708230A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a text detection method, apparatus, computer device, computer-readable storage medium, and computer program product. Background Art
[0002] With the advancement of computer and internet technologies, the advent of the 5G era, and the widespread use and development of digital documents, the intelligent processing of unstructured digital documents has become a highly sought-after research direction. For example, applications in areas such as document digitization, bill processing, handwriting entry, intelligent transportation, document retrieval, and information extraction all rely on text recognition processing in images. Text recognition includes text detection and text recognition. Text detection locates the text area in an image, then crops this area and feeds it into a text recognition model to determine the text content. In practice, text recognition accuracy is highly susceptible to the output of text detection, making text detection a crucial cornerstone module of text recognition.
[0003] However, current text detection methods usually rely on simple geometric features or preset rules, and are prone to false detection or multiple detection when encountering scanned images with non-standard formats or poor quality (such as distortion, ink noise, and complex typesetting). For example, the target image to be identified is a scanned document image. Especially for special scanned images (due to improper scanning operation resulting in document distortion, ink noise, or complex typesetting), accurate segmentation and recognition are difficult to achieve. Therefore, how to improve the accuracy of text detection in this niche field has become an urgent problem to be solved. Summary of the Invention
[0004] Based on this, it is necessary to provide a text detection method, device, computer equipment, computer-readable storage medium and computer program product to address the above technical problems, which can effectively improve the accuracy of text detection and bring convenience to users.
[0005] In a first aspect, the present application provides a text detection method, comprising: obtaining a target image containing text; extracting features of the target image; performing self-attention processing on the features to obtain attention features; determining enhanced features based on segmentation features and the attention features; wherein the segmentation features are obtained by performing semantic segmentation on the text in the target image; and determining a text detection area corresponding to the target image based on the enhanced features.
[0006] In the second aspect, the present application also provides a text detection device, including: an acquisition module for acquiring a target image containing text; an extraction module for extracting features of the target image; a processing module for performing self-attention processing on the features to obtain attention features; a determination module for determining enhanced features based on segmentation features and the attention features; wherein the segmentation features are obtained by semantic segmentation of the text in the target image; and the text detection area corresponding to the target image is determined based on the enhanced features.
[0007] In a third aspect, the present application also provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the following steps when executing the computer program: obtaining a target image containing text; extracting features of the target image; performing self-attention processing on the features to obtain attention features; determining enhanced features based on segmentation features and the attention features; wherein the segmentation features are obtained by performing semantic segmentation on the text in the target image; and determining a text detection area corresponding to the target image based on the enhanced features.
[0008] In a fourth aspect, the present application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the following steps: obtaining a target image containing text; extracting features of the target image; performing self-attention processing on the features to obtain attention features; determining enhanced features based on segmentation features and the attention features; wherein the segmentation features are obtained by semantically segmenting the text in the target image; and determining a text detection area corresponding to the target image based on the enhanced features.
[0009] In a fifth aspect, the present application also provides a computer program product, comprising a computer program, which, when executed by a processor, implements the following steps: obtaining a target image containing text; extracting features of the target image; performing self-attention processing on the features to obtain attention features; determining enhanced features based on segmentation features and the attention features; wherein the segmentation features are obtained by semantically segmenting the text in the target image; and determining a text detection area corresponding to the target image based on the enhanced features.
[0010] The above-mentioned text detection method, apparatus, computer equipment, computer-readable storage medium and computer program product obtain a target image containing text and extract features of the target image; perform self-attention processing on the features to obtain attention features, and determine enhanced features based on segmentation features and attention features; wherein the segmentation features are obtained by semantically segmenting the text in the target image; further, the text detection area corresponding to the target image is determined based on the enhanced features. Since the enhanced features in this application are determined based on segmentation features and attention features, and the segmentation features are obtained by semantically segmenting the text in the target image, the text detection area corresponding to the target image ultimately determined based on the enhanced features is more accurate, and also improves the robustness of text detection in complex scenarios, thereby bringing convenience to users. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments of the present application or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying any creative work.
[0012] Figure 1 A diagram showing an application environment of a text detection method in one embodiment;
[0013] Figure 2 1 is a flow chart of a text detection method in one embodiment;
[0014] Figure 3 is a schematic diagram of performing orientation correction on an original image in one embodiment;
[0015] Figure 4 A schematic diagram of the network structure of a text detection model provided in one embodiment;
[0016] Figure 5 A schematic diagram of a partial network structure for determining enhanced features based on features of a target image provided in one embodiment;
[0017] Figure 6 A schematic diagram of an ink noise filtering frame according to an embodiment;
[0018] Figure 7 A schematic diagram of the effect of preprocessing an original image in one embodiment;
[0019] Figure 8 A schematic diagram of a document structure mask obtained by identifying different regions in a document image in one embodiment;
[0020] Figure 91 is a schematic diagram of a process for performing distortion correction on a direction correction image in one embodiment;
[0021] Figure 10 A schematic diagram of the overall process of text detection provided in one embodiment;
[0022] Figure 11 A schematic diagram of the structure of the training side of a text detection model provided in one embodiment;
[0023] Figure 12 Schematic diagram of the difference between the real text mask image and the model prediction mask image in one embodiment;
[0024] Figure 13 FIG. 1 is a schematic diagram of a local network structure for outputting a fine-tuned contraction factor r and expansion factor a in one embodiment;
[0025] Figure 14 is a structural block diagram of a text detection device in one embodiment;
[0026] Figure 15 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0027] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0028] It should be noted that in the following description, the terms "first, second and third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first, second and third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.
[0029] The text detection method provided in the embodiment of the present application can be applied to Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store data that the server 104 needs to process. The data storage system can be integrated on the server 104, or it can be placed on the cloud or other network servers. The server 104 can be the background server of the application, and the terminal 102 obtains the target image containing text, and sends the obtained target image containing text to the background server of the application, that is, the server 104, so that the server 104 extracts the features of the target image, performs self-attention processing on the features, obtains the attention features, and determines the enhanced features based on the segmentation features and the attention features; wherein the segmentation features are obtained by semantic segmentation of the text in the target image, and the server 104 determines the text detection area corresponding to the target image based on the enhanced features, and returns the determined text detection area corresponding to the target image to the terminal 102, so that the terminal 102 can visualize the text detection area corresponding to the target image.
[0030] Terminal 102 may include, but is not limited to, various personal computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices may include smart speakers, smart TVs, smart air conditioners, smart car devices, and projectors. Portable wearable devices may include smart watches, smart bracelets, and head-mounted devices. Head-mounted devices may include virtual reality (VR) devices, augmented reality (AR) devices, smart glasses, and the like. Server 104 may be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server providing cloud computing services.
[0031] In an exemplary embodiment, Figure 2 As shown, a text detection method is provided, which can be executed by a server or a terminal alone, or by a server and a terminal together. Figure 1 The terminal in FIG is taken as an example to illustrate the method, including the following steps 202 to 210. Among them:
[0032] Step 202: Acquire a target image containing text.
[0033] The target image refers to the image to be recognized. For example, the target image in this application can be a document image obtained by scanning, or a document image taken in real time, or a document image downloaded from the Internet. That is, the target image in this application contains text (text content), for example, Figure 3 As shown in FIG. 1 , it is a schematic diagram of performing direction correction on the original image. The target image in this application can be as follows: Figure 3 A scanned image of a document containing text is shown in .
[0034] Optionally / exemplarily, the terminal responds to a recognition operation (e.g., an operation for text recognition) triggered by an operator (user) in a social application (or image application) to obtain a target image containing text corresponding to the recognition operation. That is, a user can log in to an application through a trigger operation. Furthermore, the user can trigger an operation to enter a page with the "image recognition or text extraction from image" function on the page displayed by the application. Furthermore, the user can trigger a text recognition operation on the page with the "image recognition or text extraction from image" function. The terminal then responds to the text recognition operation, obtains a target image containing text corresponding to the text recognition operation, and displays the obtained target image containing text, so that the user can reconfirm whether the obtained target image is correct.
[0035] For example, Figure 3 As shown in FIG, a schematic diagram of correcting the orientation of the original image is shown. User A (operator) can start application A in the terminal by triggering an operation and log in to application A by entering the account and password. Furthermore, in the application page displayed on the terminal, user A can first select the target image to be recognized and click the icon for initiating a text recognition request. The terminal responds to the above text recognition operation triggered by user A and obtains the following text: Figure 3 The target image containing text corresponding to the text recognition operation is shown in the document image (i.e. Figure 3 ), and can display the obtained document image containing text on the page (i.e. Figure 3 ), so that the user can confirm again whether the acquired target image is correct.
[0036] Step 204: extract features of the target image.
[0037] Here, the feature refers to the feature extracted from the target image. For example, the feature in this application can be a feature map, and the feature map can be obtained by splicing feature maps of different levels. Different level feature maps can be extracted from the target image according to different size ratios. For example, Figure 4 As shown in the figure, it is a schematic diagram of the network structure of the text detection model provided by this application. Figure 4 The structure in the rectangular box is the backbone network in the text model structure. The target image can be input as an input parameter. Figure 4 The pre-trained text detection model shown in , and by Figure 4The backbone network in the text model structure shown in the figure, namely the deep convolutional network, extracts features of different scales from the input target image to obtain feature maps of different scales and different semantic depths. Then, through upsampling (such as bilinear interpolation) and lateral connections, the deep high-semantic but low-resolution feature map is fused with the shallow low-semantic but high-resolution feature map to obtain the following: Figure 4 The fused feature map F is output by the backbone network shown in the figure, and the fused feature map F is used as the feature of the target image (Figure).
[0038] Step 206: Perform self-attention processing on the features to obtain attention features.
[0039] Among them, self-attention processing refers to processing the features of the target image through a multi-head self-attention mechanism, for example Figure 5 As shown in the figure, it is a schematic diagram of a partial network structure for determining enhanced features based on the features of the target image provided by this application. Assuming that Figure 4 After the backbone network in the text model structure shown in the figure extracts the fusion feature map F of the target image, it can be further processed as follows: Figure 5 The "self-attention" mechanism shown in the figure performs self-attention processing on the fused feature map F to obtain the attention feature F1. That is, the attention feature in this application refers to the more accurate feature obtained after performing self-attention processing on the features of the target image.
[0040] Step 208: Determine enhanced features based on the segmentation features and the attention features; wherein the segmentation features are obtained by performing semantic segmentation on the text in the target image.
[0041] The segmentation features herein refer to the features obtained after semantic segmentation of the text in the target image. In some cases, the segmentation features in this application may also be referred to as hierarchical features. For example, the segmentation features in this application may be obtained by processing the target image through a segmentation model, i.e., by identifying different regions of the text contained in the target image (such as titles, text, tables, and other different paragraph structures) through the segmentation model, obtaining identified segmentation results (for example, the segmentation results include layout areas and non-layout areas), and using the segmentation results as additional features, i.e., segmentation features.
[0042] Enhanced features refer to features determined based on segmentation features and attention features that are more capable of presenting the text structure contained in the target image. That is, the enhanced features in this application are more text-structure-aware than the fusion features output by the model in the traditional way.
[0043] Specifically, after the terminal obtains the target image containing text, the terminal can call the pre-trained text detection model and extract features of different scales from the target image through the deep convolutional network in the text detection model, so as to obtain feature maps of different scales and different semantic depths. Then, through upsampling (such as bilinear interpolation) and lateral connections, the deep high-semantic but low-resolution feature map is fused with the shallow low-semantic but high-resolution feature map to obtain the following: Figure 4 The fused feature map F shown in is used as the feature F (Figure) of the target image. Furthermore, the terminal can perform self-attention processing on the extracted feature F through the self-attention layer ("self-attention") in the text detection model to obtain the following: Figure 4 , and determine the enhanced feature F2 based on the segmentation feature S1 and the attention feature F1. For example, the terminal can calculate the product F2 between the segmentation feature S1 and the attention feature F1, and use the product F2 as the enhanced feature F2. The segmentation feature S1 in this application can be the terminal calling a pre-trained segmentation model to process the target image, that is, the segmentation model is used to identify different areas of the text contained in the target image (such as different paragraph structures such as titles, text, tables, etc.), and obtain the identified segmentation result S1 (for example, the segmentation result includes layout areas and non-layout areas), and use the segmentation result S1 as an additional feature, namely, the segmentation feature S1.
[0044] For example, Figure 7 The figure shows the effect of pre-processing the original image. User A (operator) can start the social application A in the terminal by triggering the operation and log in to the social application A by entering the account and password. Furthermore, in the application page displayed on the terminal, user A can select the target image to be recognized and click the icon for initiating text recognition request. The terminal responds to the above text recognition operation triggered by user A and obtains the following information: Figure 3 The target image containing text corresponding to the text recognition operation is document image A (i.e. Figure 3 ), the obtained document image A is preprocessed to obtain the preprocessed document image A (i.e. Figure 7 (2)), and the pre-processed document image A (i.e. Figure 7 (2)) as input parameters, input as follows Figure 4 In the pre-trained text detection model shown in , so that Figure 4The deep convolutional network (the network in the dotted rectangle) in the text detection model shown in the figure extracts features of different scales from the target image to obtain feature maps of different scales and different semantic depths. Then, through upsampling (such as bilinear interpolation) and lateral connections, the deep high-semantic but low-resolution feature map is fused with the shallow low-semantic but high-resolution feature map to obtain the following: Figure 4 The fused feature map F shown in is used as the feature map F of the preprocessed document image A. Furthermore, the terminal can use the feature map F output by the deep convolutional network (the network in the dotted rectangular box) in the text detection model as the input parameter of the next layer, the self-attention layer ("self-attention"), that is, the feature map F output by the deep convolutional network (the network in the dotted rectangular box) is processed by the self-attention layer ("self-attention") in the text detection model, and the result can be as follows: Figure 4 The attention feature F1 shown in , and determine the product F2 between the segmentation feature S1 and the attention feature F1, and use the product F2 as the enhanced feature F2. Wherein, the segmentation feature S1 in this application can be the target image after preprocessing (i.e., Figure 7 (2) is processed, that is, the different regions (such as titles, text, tables and other paragraph structures) in the text contained in the preprocessed target image are identified by the segmentation model, and the identified segmentation result S1 is obtained (for example, the segmentation result includes layout areas and non-layout areas), and the segmentation result S1 is used as an additional feature, namely, the segmentation feature S1. For example, Figure 8 As shown in FIG, a schematic diagram of a document structure mask obtained by identifying different regions in a document image, that is, the segmentation model in this application can be as follows Figure 8 The segmentation mask shown in (2) above, where the layout area in the segmentation mask is white and the non-layout area is black.
[0045] Step 210: Determine a text detection area corresponding to the target image based on the enhanced features.
[0046] The text detection area refers to the detection area corresponding to the text contained in the detected target image. For example, the text detection area in this application can be a text detection box, such as Figure 6 As shown in FIG. 1 , which is a schematic diagram of filtering ink noise frame, the text detection area in this application can be as follows: Figure 6 The text (detection) box shown in .
[0047] Specifically, after the terminal determines the enhanced features based on the segmentation features and the attention features, the terminal can generate a probability map based on the determined enhanced features, and obtain a dynamically determined expansion factor, and perform expansion processing on the probability map based on the obtained expansion factor to obtain an expanded probability map; further, the terminal can determine the text detection area corresponding to the target image based on the expanded probability map, for example, the terminal can determine the text detection box corresponding to the target image based on the expanded probability map. Figure 6 As shown in (1), the expansion factor in this application is dynamically determined by the target layer in the text detection model.
[0048] For example, Figure 4 As shown in , the terminal will pre-process the document image A (i.e. Figure 7 (2)) as input parameters, input as follows Figure 4 In the pre-trained text detection model shown in , so that Figure 4 The deep convolutional network (the network in the dotted rectangular box) in the text detection model shown in the figure outputs the feature map F of the target image, and uses the feature map F as the input parameter of the next layer, the self-attention layer ("self-attention"), that is, the feature map F output by the deep convolutional network (the network in the dotted rectangular box) is processed by the self-attention layer ("self-attention") in the text detection model, and the following is obtained: Figure 4 The attention feature F1 shown in FIG is obtained, and the product F2 between the segmentation feature S1 and the attention feature F1 is determined. After the product F2 is used as the enhanced feature F2, the terminal can generate a probability map based on the determined enhanced feature F2 and obtain the probability map obtained by Figure 4 The expansion factor a and the contraction factor r (dynamically determined by the model) output by the MLP layer in the text detection model shown in are obtained, and the probability map is expanded based on the obtained expansion factor a to obtain an expanded probability map; further, the terminal can determine the target image, i.e., the preprocessed document image A, based on the expanded probability map (i.e., Figure 7 For example, the terminal can traverse the probability value corresponding to each pixel position in the expanded probability map to determine the text detection area corresponding to the target image, that is, when the probability value corresponding to a certain pixel position in the expanded probability map is greater than a preset threshold, the pixel position corresponding to the probability value is used as the text detection area (generating a text detection frame); when the probability value corresponding to a certain pixel position in the expanded probability map is not greater than the preset threshold, the pixel position corresponding to the probability value is used as the non-text detection area (no text detection frame is generated).
[0049] In this embodiment, a target image containing text is obtained and features of the target image are extracted; self-attention processing is performed on the features to obtain attention features, and enhanced features are determined based on the segmentation features and attention features; wherein the segmentation features are obtained by semantically segmenting the text in the target image; further, the text detection area corresponding to the target image is determined based on the enhanced features. Since the enhanced features in this application are determined based on the segmentation features and attention features, and the segmentation features are obtained by semantically segmenting the text in the target image, the text detection area corresponding to the target image ultimately determined based on the enhanced features is more accurate, and also improves the robustness of text detection in complex scenarios, thereby bringing convenience to users.
[0050] In an exemplary embodiment, before acquiring a target image containing text, the method includes:
[0051] Preprocessing the target image to obtain a preprocessed target image;
[0052] The step of obtaining a target image containing text includes:
[0053] Get the preprocessed target image.
[0054] Specifically, the document image is used as an example for explanation. Figure 3 As shown, the terminal responds to the text recognition operation triggered by user A and obtains Figure 3 The document image corresponding to the text recognition operation is shown in Figure 3 The original image shown in ), and the obtained document image containing text can be displayed on the page (i.e. Figure 3 ), so that the user can confirm whether the selected target image is correct. Figure 3 The original image shown in ) is preprocessed to obtain the following Figure 7 The pre-processed target image shown in (2) is the correction image, and the correction image is used as the input parameter in the subsequent text detection model. As a result, by introducing direction correction, the document image can be accurately corrected, ensuring that the text area can be accurately extracted even when the document is tilted or deformed, thereby improving the robustness of text detection in complex scenes.
[0055] In one exemplary embodiment, the step of preprocessing the target image to obtain the preprocessed target image includes:
[0056] Performing color space conversion on the target image to obtain a target color space image;
[0057] Direction correction is performed on the target color space image to obtain a target color space image that conforms to the target direction;
[0058] The target color space image that meets the target direction is distorted and corrected to obtain a corrected image, and the corrected image is used as the preprocessed target image.
[0059] The corrected image in the embodiment of the present application does not include non-text areas.
[0060] Specifically, the document image is used as an example for explanation. Figure 3 As shown, the terminal responds to the text recognition operation triggered by user A and obtains Figure 3 The document image corresponding to the text recognition operation is shown in Figure 3 After that, the terminal can process the obtained document image containing text (i.e. Figure 3 The original image shown in ) is pre-processed, that is, the terminal can pre-process the target image (i.e. Figure 3 The original image shown in ) is converted into a color space to obtain a target color space image (such as an HSV color space image), and the target color space image is oriented and corrected by the direction correction model to obtain the following: Figure 3 The target color space image (i.e., the direction correction map) that meets the target direction is shown on the right side of the figure. Then, the target color space image (i.e., the direction correction map) that meets the target direction is distorted and corrected to obtain the following: Figure 7 The corrected image shown in (2) is used as the target image after preprocessing. The detailed processing process may include:
[0061] (1) Color space conversion: The terminal converts the RGB color space of the input target image into the HSV color space because the HSV space can better separate color attributes and is more robust to lighting changes.
[0062] (2) Direction correction: The document image (target image) in the HSV color space is fed into a direction correction model to obtain a document image with the text facing in the positive direction. The direction correction model is a classification model that is divided into four categories: [0 degrees, 90 degrees, 180 degrees, 270 degrees], i.e., the text facing upward, downward, left, and right respectively. For different categories, the corresponding angles are rotated to obtain a document image with the text facing upward (in line with the target direction). Figure 3 As shown in the figure, the original angle of the image is 90 degrees and it is rotated to 0 degrees.
[0063] (3) Distortion correction: The correction model uses a semantic segmentation-based method and is divided into two categories, where the entire text content in the document image is regarded as the foreground 1 and the document edge is regarded as the background 0. Figure 9As shown in FIG, it is a flow chart of the process of performing distortion correction on the direction correction image. Figure 9 (1) shows the direction correction image obtained after the original document image is oriented, as shown in Figure 9 (2) shows the mask image (binary image) predicted by the model. That is, the terminal can use the maximum contour search algorithm of computer vision to find the maximum contour of the predicted mask image (binary image) based on the mask image predicted by the model, and then fit the polygon into a quadrilateral based on the contour point of the maximum contour and find 4 vertices. Since the 4 vertices obtained are in disorder, the terminal needs to first sort them according to the x-axis coordinate position to determine the left and right groups, and then determine the front and back groups according to the y-axis coordinates, and finally obtain the coordinates sorted in clockwise order. The terminal then performs distortion correction on the document image based on perspective transformation to obtain the following: Figure 7 The correction image shown in (2) only shows the text area. Since the proportion of text pixels in the image increases after the non-text area is removed, it can solve the problem of poor text detection effect caused by the small proportion of text pixels. It can also remove the interference of noise in some non-text areas, thereby effectively improving the accuracy of text detection.
[0064] In one exemplary embodiment, the step of performing distortion correction on the target color space image that conforms to the target direction to obtain a corrected image includes:
[0065] The target color space image that matches the target direction is processed by the segmentation model to obtain a predicted binary image;
[0066] Determine the target contour points of the predicted binary image;
[0067] The predicted binary image is transformed based on the target contour points to obtain a transformed binary image; wherein the transformed binary image is a quadrilateral that meets the preset image shape condition;
[0068] The transformed binary image is used as the rectified image.
[0069] Among them, the predicted binary image in the embodiment of the present application can be a polygonal image.
[0070] Specifically, the document image is used as an example for explanation. Figure 9 As shown in , the terminal converts the RGB color space of the input target image into the HSV color space, and performs direction correction processing on the document image (target image) in the HSV color space through the direction correction model, and the output is as follows Figure 9 After the text direction shown in (1) is corrected to the positive direction, the terminal can use the segmentation model to Figure 9 The target color space image (i.e., the direction correction image) shown in (1) is processed to obtain the following Figure 9The predicted binary image (i.e., mask image) shown in (2) is obtained; further, the terminal can determine the target contour points of the predicted binary image and perform a nonlinear transformation on the predicted binary image based on the target contour points to obtain the following: Figure 9 The transformed binary image (i.e., quadrilateral image) shown in (3) is filled with text as the corrected image; wherein the transformed binary image is a quadrilateral that meets the preset image shape conditions. The specific search method for determining the target contour points of the predicted binary image can be as follows:
[0071] Input a set of contour points M of the maximum contour and output a set of quadrilateral points N:
[0072] Step 1: Set the threshold t, take the starting point A and the end point B of A and add them to N;
[0073] Step 2: Take a point C in M so that it is farthest from the straight line connecting A and B.
[0074] Step 3: If the distance is greater than the threshold, add C to N;
[0075] Step 4, recursively AC and CB respectively;
[0076] step5. Output result set N.
[0077] Since the four vertices obtained are in disorder, the terminal needs to first sort them according to the x-axis coordinate position to determine the left and right groups, and then determine the front and back groups according to the y-axis coordinates, and finally obtain the coordinates sorted clockwise. The terminal then uses perspective transformation (which projects an image to a new viewing plane. The process includes: 1. Converting the two-dimensional coordinate system to a three-dimensional coordinate system; 2. Projecting the three-dimensional coordinate system to a new two-dimensional coordinate system. This process is a nonlinear transformation process. A rhombus becomes a quadrilateral after nonlinear transformation, but it is no longer parallel) to correct the distortion of the document image to obtain the following Figure 7 The correction image shown in (2) only shows the text area. Since the proportion of text pixels in the image increases after the non-text area is removed, it can solve the problem of poor text detection effect caused by the small proportion of text pixels. It can also remove the interference of noise in some non-text areas, thereby effectively improving the accuracy of text detection.
[0078] In an exemplary embodiment, the feature includes a fused feature map; and the step of extracting the feature of the target image includes:
[0079] Extract hierarchical feature maps of the target image through the text detection model; hierarchical feature maps are extracted feature maps of different sizes;
[0080] Multiple levels of feature maps are concatenated into a fused feature map.
[0081] Specifically, the document image is used as an example for explanation. Figure 4 As shown in , the terminal responds to the text recognition operation triggered by user A and obtains Figure 3 The target image containing text corresponding to the text recognition operation is document image A (i.e. Figure 3 ), the obtained document image A is preprocessed to obtain the preprocessed document image A (i.e., Figure 7 (2)), the terminal can call the pre-trained text detection model and use the pre-processed document image A (i.e. Figure 7 (2)) as input parameters, input as follows Figure 4 In the pre-trained text detection model shown in , so that Figure 4 The deep convolutional network (the network in the dotted rectangle) in the text detection model shown in the figure extracts features of different scales from the target image, and obtains feature maps of different scales and different semantic depths (i.e., hierarchical feature maps). Then, through upsampling (such as bilinear interpolation) and lateral connections (Lateral Connections), the deep high-semantic but low-resolution feature map is spliced with the shallow low-semantic but high-resolution feature map (i.e., hierarchical feature map), and the following is obtained: Figure 4 The fused feature map F shown in FIG is used as the feature map F of the preprocessed document image A. This makes the subsequent text detection area corresponding to the target image determined based on the enhanced features more accurate, and also improves the robustness of text detection in complex scenarios, thereby bringing convenience to users.
[0082] In an exemplary embodiment, the attention feature includes an attention feature map; and the step of performing self-attention processing on the feature to obtain the attention feature includes:
[0083] Perform self-attention processing on the fused feature map through the text detection model to obtain the attention feature map;
[0084] The method further comprises: performing semantic segmentation on the text in the target image to obtain segmentation features;
[0085] The determining of the enhanced feature based on the segmentation feature and the attention feature includes:
[0086] Determine the product between the segmentation feature and the attention feature map and use the product as the enhanced feature.
[0087] Specifically, the document image is used as an example for explanation. Figure 4 As shown in , the terminal will pre-process the document image A (i.e. Figure 7 (2)) as input parameters, input as follows Figure 4 In the pre-trained text detection model shown in , so that Figure 4 The deep convolutional network (the network in the dotted rectangular box) in the text detection model shown in the figure outputs the fused feature map F of the target image, and uses the fused feature map F as the input parameter of the next layer, the self-attention layer ("self-attention"), that is, the fused feature map F output by the deep convolutional network (the network in the dotted rectangular box) is processed by the self-attention layer ("self-attention") in the text detection model, and the following is obtained: Figure 4 The attention feature F1 shown in , and the product F2 between the segmentation feature (ie, the hierarchical mask image) S1 and the attention feature F1 is calculated, and the product F2 is used as the enhanced feature F2. The segmentation feature S1 in this application can be the terminal calling the pre-trained segmentation model for the pre-processed target image (ie, Figure 7 (2) is processed, that is, the different regions (such as different paragraph structures such as titles, text, tables, etc.) in the text contained in the preprocessed target image are identified through the segmentation model, and the identified segmentation result S1 is obtained (for example, the segmentation result is a hierarchical mask map, and the hierarchical mask map includes layout areas and non-layout areas), and the segmentation result S1 is used as an additional feature, namely, the segmentation feature S1. For example, Figure 8 As shown in FIG, a schematic diagram of a document structure mask obtained by identifying different regions in a document image, that is, the segmentation model in this application can be as follows Figure 8 The segmentation mask shown in (2) above shows the layout area in white and the non-layout area in black. This makes the text detection area corresponding to the target image determined based on the enhanced features more accurate, and also improves the robustness of text detection in complex scenes, thereby bringing convenience to users.
[0088] In an exemplary embodiment, the step of determining a text detection area corresponding to a target image based on enhanced features includes:
[0089] determining a probability map based on the enhanced features;
[0090] Expanding the probability map based on the expansion factor to obtain an expanded probability map; wherein the expansion factor is dynamically determined by the target layer in the text detection model;
[0091] The text detection area corresponding to the target image is determined based on the expanded probability map.
[0092] Specifically, if Figure 4 As shown in , the terminal will pre-process the document image A (i.e. Figure 7 (2)) as input parameters, input as follows Figure 4 In the pre-trained text detection model shown in , so that Figure 4 The self-attention layer ("self-attention") in the text detection model shown in the figure performs self-attention processing on the feature map F output by the deep convolutional network (the network in the dotted rectangle), and obtains the following Figure 4 The attention feature F1 shown in FIG is obtained, and the product F2 between the segmentation feature S1 and the attention feature F1 is determined. After the product F2 is used as the enhanced feature F2, the terminal can generate a probability map based on the determined enhanced feature F2 and obtain the probability map obtained by Figure 4 The expansion factor a and the contraction factor r (dynamically determined by the model) output by the MLP layer in the text detection model shown in are obtained, and the probability map is expanded based on the obtained expansion factor a to obtain an expanded probability map; further, the terminal can determine the target image, i.e., the preprocessed document image A, based on the expanded probability map (i.e., Figure 7 (2) The text detection frame corresponding to the correction image) is shown in FIG. For example, the terminal can determine the text detection frame corresponding to the target image based on the expanded probability map. Figure 6 As shown in (1). As a result, the text detection area corresponding to the target image determined based on the enhanced features is more accurate, and the robustness of text detection in complex scenes is improved, thereby bringing convenience to users.
[0093] In one exemplary embodiment, determining the text detection area corresponding to the target image based on the expanded probability map includes:
[0094] When the probability value in the expanded probability map is greater than the preset threshold, the pixel position corresponding to the probability value is used as the text detection area;
[0095] When the probability value in the expanded probability map is not greater than a preset threshold, the pixel position corresponding to the probability value is taken as the non-text detection area.
[0096] Specifically, the terminal determines the target image, i.e., the pre-processed document image A (i.e., Figure 7When determining the text detection area corresponding to the correction image shown in (2), the terminal can traverse the probability value corresponding to each pixel position in the expanded probability map to determine the text detection area corresponding to the target image. That is, when the probability value corresponding to (a certain) pixel position in the expanded probability map is greater than a preset threshold, the pixel position corresponding to the probability value is used as the text detection area (generating a text detection frame); when the probability value corresponding to (a certain) pixel position in the expanded probability map is not greater than the preset threshold, the pixel position corresponding to the probability value is used as the non-text detection area (no text detection frame is generated). As a result, the text detection area corresponding to the target image determined based on the enhanced features is more accurate, and the robustness of text detection in complex scenes is also improved, thereby bringing convenience to users.
[0097] In an exemplary embodiment, the text detection area includes a text detection frame; after determining the text detection area corresponding to the target image based on the enhanced features, the method further includes:
[0098] Noise filtering is performed on the text detection frame to obtain a filtered text detection frame.
[0099] Specifically, if Figure 10 As shown in FIG, it is a schematic diagram of the overall process of text detection provided by this application. Figure 10 As shown in , the terminal determines the target image, i.e., the pre-processed document image A, based on the expanded probability map (i.e., Figure 7 After the text detection area (text detection frame) corresponding to the correction image shown in (2) is detected, the terminal can further Figure 6 The text detection box shown in (1) is filtered for noise, and the following is obtained: Figure 6 The filtered text detection box shown in (2) has filtered out Figure 6 The false positive text boxes shown in (1) (i.e., the false positive text boxes that have been filtered out) Figure 6 The arrow in (1) points to the misdetected ink noise. Thus, by combining the noise filtering method of image and NLP, unnecessary interference items are effectively reduced, and the quality and reliability of the final output results are improved.
[0100] In one exemplary embodiment, the step of performing noise filtering on the text detection frame to obtain a filtered text detection frame includes:
[0101] Obtain the preprocessed target image;
[0102] Perform morphological processing on the preprocessed target image to obtain an optimized image;
[0103] Extract the text area corresponding to the optimized image;
[0104] The classification confidence of the pixels in the text area is determined, and the text detection frame is filtered based on the classification confidence to obtain a filtered text detection frame.
[0105] Specifically, if Figure 10 As shown in , the terminal determines the target image, i.e., the pre-processed document image A, based on the expanded probability map (i.e., Figure 7 After the text detection area (text detection frame) corresponding to the correction image shown in (2) is detected, the terminal can further Figure 6 The text detection frame shown in (1) is used to filter out noise. For example, the terminal can obtain the pre-processed target image (i.e. Figure 7 ), and perform morphological processing on the pre-processed target image to obtain an optimized image; further, the terminal can extract the text area corresponding to the optimized image, determine the classification confidence of the pixels in the text area, and filter the text detection frame based on the classification confidence, and obtain the following Figure 6 The filtered text detection box shown in (2) is shown in Figure 2. Thus, by combining the noise filtering method of image and NLP, unnecessary interference items are effectively reduced, and the quality and reliability of the final output result are improved.
[0106] In one exemplary embodiment, the step of performing morphological processing on the pre-processed target image to obtain an optimized image includes:
[0107] Binarize the preprocessed target image to obtain a binary image;
[0108] Remove isolated pixels from the binary image to obtain a corroded image;
[0109] The dilation operation is used to restore the edges of the eroded text in the eroded image to obtain an optimized image.
[0110] Specifically, the terminal obtains the pre-processed target image (i.e., Figure 7 After obtaining the correction image shown in ), the terminal can binarize the preprocessed image, converting it to black and white, with the text area in white and the background in black, to obtain a binary image. The terminal can then use an erosion operation to remove isolated points in the binary image. Since the erosion operation shrinks the foreground area (white area), it can remove small, isolated white pixels. Finally, the terminal can use a dilation operation to restore the eroded text edges. This dilation operation expands the foreground area, thereby restoring some text edges lost due to erosion, resulting in an optimized image. This effectively reduces unnecessary interference by combining image and NLP noise filtering methods, improving the quality and reliability of the final output.
[0111] In an exemplary embodiment, after extracting the text area corresponding to the optimized image, the method further includes:
[0112] Remove non-text information in the text area to obtain a cleaned text area;
[0113] Determining a first word and a second word in the cleaned text area; wherein the appearance frequency of the first word is greater than a preset frequency, and the appearance frequency of the second word is less than the preset frequency;
[0114] The determining of the classification confidence of the pixels in the text area and filtering the text detection frame based on the classification confidence to obtain the filtered text detection frame includes:
[0115] The cleaned text area is processed through the document structure model to obtain the classification confidence of the pixels in the text area;
[0116] The text detection frame is filtered based on the first vocabulary, the second vocabulary, and the classification confidence to obtain a filtered text detection frame.
[0117] Specifically, if Figure 10 As shown in , the terminal determines the target image, i.e., the pre-processed document image A, based on the expanded probability map (i.e., Figure 7 After the text detection area (text detection frame) corresponding to the correction image shown in (2) is detected, the terminal can further Figure 6 The text detection frame shown in (1) is used to filter out noise. For example, the terminal can obtain the pre-processed target image (i.e. Figure 7 ), and perform morphological processing on the preprocessed target image to obtain an optimized image; further, the terminal can extract the text area corresponding to the optimized image, and remove the non-text information in the text area to obtain a cleaned text area; further, the terminal can determine the first word and the second word in the cleaned text area; wherein the frequency of occurrence of the first word is greater than the preset frequency, and the frequency of occurrence of the second word is less than the preset frequency; the terminal can process the cleaned text area through the document structure model to obtain the classification confidence of the pixel points in the text area, and filter the text detection box based on the first word, the second word and the classification confidence, and obtain the following Figure 6 The filtered text detection box shown in (2) is shown in Figure 2. Thus, by combining the noise filtering method of image and NLP, unnecessary interference items are effectively reduced, and the quality and reliability of the final output result are improved.
[0118] In an exemplary embodiment, the step of determining a text detection area corresponding to a target image based on enhanced features includes:
[0119] Predict a binary image based on the enhanced features;
[0120] Generate the text detection area corresponding to the target image based on the binary image.
[0121] Specifically, if Figure 11 The figure shows the structure diagram of the training side of the text detection model provided by this application. During the model training phase, the text outline needs to be shrunk and then the probability map needs to be expanded to obtain the true text box outline. The shrinkage calculation formula is as follows (1):
[0122] (1)
[0123] The expanded calculation formula is as follows:
[0124] (2)
[0125] Where, L in the above formulas (1) and (2) is the perimeter of the annotation box, A is the area of the annotation box, r is the shrinkage factor, usually defined as 0.4, and a is the expansion factor, usually defined as 1.5. Figure 12 As shown, it is a schematic diagram of the difference between the real text mask map and the model prediction mask map, that is, assuming that there are three text boxes with lengths of short, medium and long, Figure 12 (1) is the real text mask map of three different lengths (labels of the probability map P during training), Figure 12 (2) is the mask image after shrinking and then expanding (similar to the model prediction result). Figure 12 (3) is the difference between the real text mask image and the model prediction mask image, that is, Figure 12 The wider the white border line shown in (3), the greater the difference. Figure 12 As can be seen from (3), the fixed expansion factor and contraction factor are not universal for text boxes of different lengths (a longer text box cannot be completely expanded by shrinking it first and then expanding it, and an incomplete text box will seriously affect the text recognition effect). Therefore, an adaptive enhancement module is proposed in the embodiment of the present application. The adaptive enhancement module dynamically adjusts the algorithm parameters (contraction factor r and expansion factor a) according to the specific characteristics of the input document image to adapt to different types of data sources. This module can be based on Figure 4 The feature F in is connected to an MLP layer (convolutional layer, in order to output two output values), and the output of the MLP layer is 2 output values (contraction factor r and expansion factor a).
[0126] That is, during the model training phase, such as Figure 11 The model shown in regards the shrinkage factor r and the expansion factor a as fine-tuning of the original shrinkage factor (0.4) and expansion factor (1.5), so it can be directly regarded as a regression problem. Figure 11 In the model structure shown in , feature F is a backbone structure followed by an MLP layer with outputs r and a. The loss function can be selected as the mean square error (MSE) to measure the gap between the predicted value and the target value. For each sample, the loss function can be defined as the following formula (3):
[0127] (3)
[0128] in, represents the loss value, represents the predicted shrinkage factor (i.e., the shrinkage factor after fine-tuning the original shrinkage factor), represents the original shrinkage factor (which can be 0.4), represents the predicted expansion factor (i.e., the expansion factor after fine-tuning the original expansion factor), Indicates the original expansion factor (can be 1.5).
[0129] In this embodiment, the adaptive enhanced text detection algorithm (which uses a neural network to derive an adaptive expansion scaling factor based on a variety of documents) has stronger generalization compared to the adaptive adjustment of the expansion scaling factor based on the length ratio or area single-dimensional attribute. It is proposed to improve the text detection effect of complex typesetting documents based on document structure perception, enhance the generalization ability and scope of application of the algorithm, and can maintain high performance in more diverse application scenarios.
[0130] In one exemplary embodiment, the step of predicting a binary image based on the enhanced features includes:
[0131] determining a probability map and a threshold map based on the enhanced features;
[0132] Expanding the probability map based on the expansion factor to obtain an expanded probability map; wherein the expansion factor is dynamically determined by the target layer in the text detection model;
[0133] A binary image is predicted based on the expanded probability map and the threshold map.
[0134] Specifically, if Figure 11 As shown in , during the model training phase, the terminal inputs the sample image as Figure 11 In the initial text detection model shown in , so that Figure 11The backbone network in the initial text detection model shown in the figure, namely the deep convolutional network, extracts features of different scales from the input sample image to obtain feature maps of different scales and different semantic depths. Then, through upsampling (such as bilinear interpolation) and lateral connections, the deep high-semantic but low-resolution feature map is fused with the shallow low-semantic but high-resolution feature map to obtain the following: Figure 11 The sample fusion feature map F output by the backbone network shown in is processed by the self-attention layer (“self-attention”) in the initial text detection model on the extracted sample fusion feature (map) F, and the following can be obtained: Figure 11 The sample attention feature F1 shown in , and the sample enhancement feature F2 is determined based on the sample segmentation feature S1 and the sample attention feature F1. Further, the terminal can determine the probability map and the threshold map based on the sample enhancement feature, expand the probability map based on the expansion factor to obtain the expanded probability map, and predict the binary map based on the expanded probability map and the threshold map, and finally determine the text detection area corresponding to the sample image based on the binary map; wherein the expansion factor is obtained by Figure 11 The target layer (MLP layer) in the initial text detection model shown in the figure is dynamically determined. This makes the adaptively enhanced text detection algorithm (which uses a neural network to derive adaptive expansion scaling factors based on a variety of documents) more generalizable than adaptive scaling factors based on single-dimensional attributes such as length ratios or area. The proposed method improves text detection in complex document layouts based on document structure awareness, enhancing the algorithm's generalization and applicability, enabling it to maintain high performance in a wider range of application scenarios.
[0135] This application also provides an application scenario, in which the above-mentioned text detection method is applied. Specifically, the application of the text detection method in this application scenario is as follows:
[0136] In the process of user interaction with an information platform (such as a social application), the above-mentioned text detection method can be used. When the user (operator) wants to recognize the text content in the image, the user can open the information platform (such as a social application) on the terminal through a trigger operation and trigger a text recognition operation in the information platform (such as a social application). The terminal responds to the text recognition operation triggered by the user (operator), obtains a target image containing text corresponding to the text recognition operation, and calls Figure 4The pre-trained text detection model shown in [1] is used to extract features of the target image using this text detection model. The extracted features are then subjected to self-attention processing to obtain attention features. Furthermore, the terminal can use a segmentation model to semantically segment the text in the target image (paragraph segmentation) to obtain segmentation features. Enhancement features are determined based on the segmentation features and attention features, and the text detection region (i.e., text detection box) corresponding to the target image is determined based on the enhanced features. This allows for adaptive detection of text content in the target image, effectively improving the processing efficiency and accuracy of text detection. The determined text detection region corresponding to the target image is then visually displayed, effectively improving the user experience and providing convenience.
[0137] The method provided in the embodiment of the present application can be applied to various text detection scenarios. The following uses the scenario of user interaction with a social application as an example to illustrate the text detection method provided in the embodiment of the present application.
[0138] Among them, OCR: Optical Character Recognition refers to the process of analyzing and identifying images to obtain text and layout information, which is a typical computer vision task.
[0139] DBNET: Differentiable binarization network (differentiable binarization network) is a segmentation-based text detection algorithm.
[0140] DeepLabv3-plus: An advanced deep learning model for semantic segmentation tasks. The DeepLabV3plus model uses an encoder-decoder structure to extract image features through the encoder and then map these features back to the original image size through the decoder to achieve pixel-level classification.
[0141] With the widespread use and development of digital documents, intelligent processing of unstructured digital documents has become a highly sought-after research area. Applications in areas such as document digitization, bill processing, handwriting capture, intelligent transportation, document retrieval, and information extraction all rely on optical character recognition (OCR) technology. Text recognition involves both text detection and text recognition. Text detection locates the text region within an image, then crops this region and feeds it into a text recognition model to extract the text content. In practice, text recognition accuracy is highly susceptible to the output of text detection, making text detection a crucial cornerstone module. While OCR technology has matured thanks to the rapid development of deep learning over the past decade, specific areas face numerous challenges. For example, when scanning documents are being recognized, existing technologies struggle to accurately segment and recognize documents due to distortion, ink noise, or complex layouts caused by improper scanning. Therefore, improving text detection accuracy in this niche area is a pressing challenge.
[0142] Disadvantages of traditional methods:
[0143] Traditional text detection algorithms often rely on simple geometric features or preset rules, and are prone to false detection or multiple detection when encountering scanned documents with non-standard formats or poor quality (such as distortion, ink noise, and complex typesetting). These problems limit the widespread adoption of OCR technology in practical applications.
[0144] Therefore, this application proposes an adaptive enhanced text detection algorithm based on document structure perception, which aims to significantly improve the recognition accuracy of text areas in various types of documents.
[0145] On the technical side, the overall algorithm flow of the technical solution provided by this application is as follows Figure 10 As shown:
[0146] (1) Image preprocessing
[0147] Many scanned documents are not placed properly during the scanning process, resulting in document distortion, incorrect document orientation, a very small proportion of the text area in the document, and increased difficulty in text detection under different light intensities.
[0148] 1.1 Module function: Correct distorted documents and extract the entire text area to remove document edges and increase the percentage of text pixels.
[0149] 1.2 Module Detailed Explanation:
[0150] Module input: document image.
[0151] Module output: Corrected image of the entire text area.
[0152] Detailed processing process:
[0153] (1) Color space conversion: Convert the input RGB color space to HSV color space because HSV space can better separate color attributes and is more robust to lighting changes.
[0154] (2) Direction correction: Figure 3 As shown, the document is first fed into an orientation correction model to obtain a document with positive text orientation. The orientation correction model is a classification model with four categories: [0 degrees, 90 degrees, 180 degrees, and 270 degrees], representing text facing up, down, left, and right. For each category, the corresponding angle is rotated to produce an image with text facing upward. The following image shows an image with an original angle of 90 degrees rotated to 0 degrees.
[0155] (3) Distortion correction: Figure 9 As shown in , this model uses a semantic segmentation-based method to divide the document into two categories, with the entire text content in the document as the foreground 1 and the document edge as the background 0. Figure 1 For the original document image, Figure 2 The mask image (binary image) predicted by the model is used. Based on the mask image obtained by the model, the maximum contour search algorithm of computer vision is used to find the maximum contour of the mask image. Then, based on the contour points of the maximum contour, the polygon is fitted into a quadrilateral and the four vertices are found. The method is as follows:
[0156] Input a set of contour points M of the maximum contour and output a set of quadrilateral points N:
[0157] Step 1: Set the threshold t, take the starting point A and the end point B of A and add them to N;
[0158] Step 2: Take a point C in M so that it is farthest from the straight line connecting A and B.
[0159] Step 3: If the distance is greater than the threshold, add C to N;
[0160] Step 4, recursively AC and CB respectively;
[0161] step5. Output result set N.
[0162] Since the four vertices obtained are in disorder, they need to be sorted according to the x-axis coordinate position to determine the left and right groups, and then the front and back groups are determined according to the y-axis coordinate to finally obtain the coordinates sorted in a clockwise order. Then based on the perspective transformation (which is to project an image onto a new viewing plane, the process includes: 1. Converting the two-dimensional coordinate system to a three-dimensional coordinate system; 2. Projecting the three-dimensional coordinate system to a new two-dimensional coordinate system. This process is a nonlinear transformation process. A rhombus becomes a quadrilateral after a nonlinear transformation, but it is no longer parallel), the document is corrected and a correction image with only text areas is obtained. After removing the non-text areas, the proportion of text pixels in the image becomes larger, which can solve the problem of poor text detection results caused by the small proportion of text pixels, and can also remove the interference of noise in some non-text areas.
[0163] (2) Document structure-aware adaptive enhanced text detection algorithm
[0164] 1.1 Module function: Perform text detection on the corrected whole text area image.
[0165] 1.2 Module Detailed Explanation:
[0166] Module input: Corrected overall text area image
[0167] Module output: detection box of the text block (x, y, w, h)
[0168] Detailed processing process: The current mainstream text detection method in the industry is DBNET, so this application mainly improves the text detection algorithm based on DBNET. That is, in traditional text detection algorithms, images are usually input into the network, and then subjected to feature extraction, upsampling, fusion, and concat operations to obtain the following: Figure 4 The feature map in is called F (dashed box area), and then F is used to predict the probability map (probability map) called P and the threshold map (thresholdmap) is predicted using F and called T. Finally, the approximate binary map is calculated by P and T, and the text detection box is determined based on the approximate binary map.
[0169] The technical solution provided in this application mainly improves detection accuracy by combining adaptive enhancement and enhanced features F based on contextual semantic understanding and document structure perception (text detection algorithms cannot fully utilize high- and low-level information when performing multi-scale feature fusion, resulting in missed text detections), especially for document images with complex layouts.
[0170] 1) Adaptive enhancement:
[0171] During training, the text outline needs to be shrunk and then the probability map needs to be expanded to obtain the true text box outline. The shrinkage calculation formula is shown in the above formula (1), and the expansion calculation formula is shown in the above formula (2). Figure 12 As shown, it is a schematic diagram of the difference between the real text mask map and the model prediction mask map, that is, assuming that there are three text boxes with lengths of short, medium and long, Figure 12 (1) is the real text mask map of three different lengths (labels of the probability map P during training), Figure 12 (2) is the mask image after shrinking and then expanding (similar to the model prediction result). Figure 12 (3) is the difference between the real text mask image and the model prediction mask image, that is, Figure 12 The wider the white border line shown in (3), the greater the difference. Figure 12 As can be seen from (3), the fixed expansion factor and contraction factor are not universal for text boxes of different lengths (a longer text box cannot be completely expanded by shrinking it first and then expanding it, and an incomplete text box will seriously affect the text recognition effect). Therefore, an adaptive enhancement module is proposed in the embodiment of the present application. The adaptive enhancement module dynamically adjusts the algorithm parameters (contraction factor r and expansion factor a) according to the specific characteristics of the input document image to adapt to different types of data sources. This module can be based on Figure 4 The feature F in is connected to an MLP layer (convolutional layer, in order to output two output values), and the output of the MLP layer is 2 output values (contraction factor r and expansion factor a), as shown in Figure 13 As shown in FIG, a schematic diagram of the local network structure of the contraction factor r and expansion factor a after output fine-tuning.
[0172] That is, during the model training phase, such as Figure 11 The model shown in regards the shrinkage factor r and the expansion factor a as fine-tuning of the original shrinkage factor (0.4) and expansion factor (1.5), so it can be directly regarded as a regression problem. Figure 11 In the model structure shown in , feature F is a backbone structure followed by an MLP layer with outputs r and a. The loss function can be the mean squared error (MSE) to measure the difference between the predicted value and the target value. For each sample, the loss function can be defined as shown in the above formula (3).
[0173] 2) Integrating contextual semantic understanding: A self-attention module (such as Self-Attention) is added to the feature F obtained from the original network backbone. Through the multi-head self-attention mechanism, the model can learn different representations in different subspaces, enhancing the model's expressiveness and robustness, and obtaining feature F1.
[0174] 3) Document structure perception: For documents with complex layout, we can combine the semantic segmentation technology DeepLabv3-plus to identify different areas in the document (such as title, text, table, etc.), and then use the segmentation mask results as Figure 11 (The layout area is white, and the non-layout area is black) as an additional weight feature S1, by multiplying the segmentation feature S1 with the F1 feature, we get an enhanced feature F2 that is more aware of the document structure, and then use F2 to predict the probability map P and the threshold map T, and finally calculate the approximate binary map B through P and T. Among them, the network structure of the enhanced feature F is as follows Figure 5 As shown in .
[0175] in, Figure 5 and Figure 13 Shown are the network structures for two independent tasks. Figure 5 The shrunk text box obtained from Figure 13 The expansion factor obtained in the expansion is used to perform post-processing operations on the text box (expansion calculation formula) to obtain the final complete text box.
[0176] (3) Noise filtering
[0177] 1.1 Module function: Filter out falsely detected text boxes.
[0178] 1.2 Module Detailed Explanation:
[0179] Module input: detection box of text block (x, y, w, h)
[0180] Module output: text box after filtering false positives.
[0181] 1) Use morphological operations to remove isolated points: First, binarize the preprocessed image to convert it to black and white, with the text area in white and the background in black. Then, use the erosion operation to remove isolated points. Erosion shrinks the foreground area (white region), removing small, isolated white pixels. Finally, use the dilation operation to restore the edges of the text that were eroded. Dilation expands the foreground area, restoring some of the edges lost by erosion.
[0182] 2) Use NLP technology to evaluate the quality of text content: Extract text areas from morphologically processed images and convert them into text strings. Perform preliminary cleaning on the extracted text to remove non-text information such as spaces and special characters. Then count the word frequencies in the text and identify high-frequency and low-frequency words. High-frequency words usually represent important text content, while low-frequency words may be the result of noise or misrecognition. By combining the text area output by the document structure model (the classification confidence of all pixels in the area), the false detection boxes can be further filtered out. Figure 6 Shown in the figure is text with falsely detected ink noise (pointed by the arrow) filtered out.
[0183] The essential and auxiliary technical points in the technical solution provided in this application are:
[0184] Necessary technical points include:
[0185] 1. The text detection algorithm mentioned in this application.
[0186] 2. The semantic segmentation algorithm mentioned in this application: It is understood that the method provided in this application can also be based on other text box contraction and expansion methods.
[0187] The beneficial effects of the technical solution of this application include:
[0188] 1. By introducing orientation correction, document images can be accurately corrected, ensuring accurate text extraction even when the document is tilted or deformed, improving the robustness of text detection in complex scenarios.
[0189] 2. The adaptive enhanced text detection algorithm (which uses a neural network to derive an adaptive expansion scaling factor based on a variety of documents) has stronger generalization than adaptively adjusting the expansion scaling factor based on single-dimensional attributes such as length ratio or area. It proposes to improve the text detection effect of complex typesetting documents based on document structure perception, enhances the algorithm's generalization ability and scope of application, and can maintain high performance in more diverse application scenarios.
[0190] 3. By combining image and NLP noise filtering methods, unnecessary interference items are effectively reduced, improving the quality and reliability of the final output results.
[0191] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.
[0192] Based on the same inventive concept, the present application also provides a text detection device for implementing the aforementioned text detection method. The implementation solution provided by this device is similar to the implementation solution described in the aforementioned method. Therefore, the specific limitations of one or more text detection device embodiments provided below can be found in the above-mentioned limitations of the text detection method and will not be repeated here.
[0193] In an exemplary embodiment, Figure 14 As shown, a text detection device is provided, including: an acquisition module 1402, an extraction module 1404, a processing module 1406 and a determination module 1408, wherein:
[0194] The acquisition module 1402 is configured to acquire a target image containing text.
[0195] The extraction module 1404 is configured to extract features of the target image.
[0196] The processing module 1406 is used to perform self-attention processing on the features to obtain attention features.
[0197] Determination module 1408 is used to determine the enhanced feature based on the segmentation feature and the attention feature; wherein the segmentation feature is obtained by semantic segmentation of the text in the target image; and determine the text detection area corresponding to the target image based on the enhanced feature.
[0198] In one embodiment, the processing module is further used to preprocess the target image to obtain the preprocessed target image; the acquisition module is further used to acquire the preprocessed target image.
[0199] In one embodiment, the processing module is also used to perform color space conversion on the target image to obtain a target color space image; perform direction correction on the target color space image to obtain the target color space image that conforms to the target direction; perform distortion correction on the target color space image that conforms to the target direction to obtain a corrected image, and use the corrected image as the preprocessed target image.
[0200] In one embodiment, the features include a fused feature map; the extraction module is also used to extract the hierarchical feature map of the target image through a text detection model; the hierarchical feature map is an extracted feature map of different sizes; the device also includes: a splicing module, used to splice multiple hierarchical feature maps into the fused feature map.
[0201] In one embodiment, the attention feature includes an attention feature map; the processing module is also used to perform self-attention processing on the fusion feature map through the text detection model to obtain the attention feature map; the device also includes: a segmentation module, used to perform semantic segmentation on the text in the target image to obtain the segmentation feature; the determination module is also used to determine the product between the segmentation feature and the attention feature map, and use the product as the enhanced feature.
[0202] In one embodiment, the determination module is further used to determine a probability map based on the enhanced features; the processing module is further used to perform expansion processing on the probability map based on an expansion factor to obtain an expanded probability map; wherein, the expansion factor is dynamically determined by the target layer in the text detection model; the determination module is further used to determine the text detection area corresponding to the target image based on the expanded probability map.
[0203] Each module in the above-mentioned text detection device can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in the computer device in the form of software, so that the processor can call and execute the corresponding operations of each module.
[0204] In an exemplary embodiment, a computer device is provided. The computer device may be a terminal or a server. In this embodiment, the computer device is described as a terminal. The internal structure diagram thereof may be as follows: Figure 15As shown. The computer device includes a processor, memory, an input / output interface, a communication interface, a display unit, and an input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals via wired or wireless means, and the wireless means can be implemented via Wi-Fi, a mobile cellular network, near field communication (NFC), or other technologies. When executed by the processor, the computer program implements a text detection method. The display unit of the computer device is used to form a visually visible image, and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device casing, or an external keyboard, touchpad or mouse.
[0205] Those skilled in the art will understand that Figure 15 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0206] In an exemplary embodiment, a computer device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above-mentioned method embodiments when executing the computer program.
[0207] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0208] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.
[0209] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0210] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), quantum computing-based data processing logic devices, artificial intelligence (AI) processors, and the like.
[0211] The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0212] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.
Claims
1. A text detection method, characterized in that: The method comprises: Get the target image containing text; extracting features of the target image; Performing self-attention processing on the features to obtain attention features; Determining an enhanced feature based on the segmentation feature and the attention feature; wherein the segmentation feature is obtained by semantically segmenting the text in the target image; A text detection area corresponding to the target image is determined based on the enhanced features.
2. The method according to claim 1, characterized in that The features include a fusion feature map; The extracting the features of the target image includes: Extracting a hierarchical feature map of the target image through a text detection model; the hierarchical feature map is an extracted feature map of different sizes; The multiple hierarchical feature maps are spliced into the fused feature map.
3. The method according to claim 2, characterized in that The attention feature includes an attention feature map; the self-attention processing of the feature to obtain the attention feature includes: Performing self-attention processing on the fused feature map through the text detection model to obtain the attention feature map; The method further includes: performing semantic segmentation on the text in the target image to obtain the segmentation features; The determining of the enhanced feature based on the segmentation feature and the attention feature includes: A product between the segmentation feature and the attention feature map is determined, and the product is used as the enhanced feature.
4. The method according to claim 1, wherein The determining of the text detection area corresponding to the target image based on the enhanced feature includes: determining a probability map based on the enhanced features; Expanding the probability map based on an expansion factor to obtain an expanded probability map; wherein the expansion factor is dynamically determined by a target layer in a text detection model; A text detection region corresponding to the target image is determined based on the expanded probability map.
5. The method according to claim 1, wherein Before acquiring the target image containing text, the method further includes: Performing color space conversion on the target image to obtain a target color space image; Performing direction correction on the target color space image to obtain the target color space image that conforms to the target direction; Performing distortion correction on the target color space image that conforms to the target direction to obtain a corrected image, and using the corrected image as the preprocessed target image; The step of obtaining a target image containing text includes: Acquire the preprocessed target image.
6. The method according to claim 1, characterized in that The text detection area includes a text detection frame; after determining the text detection area corresponding to the target image based on the enhanced features, the method further includes: Acquiring the preprocessed target image; Performing morphological processing on the pre-processed target image to obtain an optimized image; Extracting a text area corresponding to the optimized image; The classification confidence of the pixels in the text area is determined, and the text detection frame is filtered based on the classification confidence to obtain the filtered text detection frame.
7. The method according to claim 1, characterized in that The determining of the text detection area corresponding to the target image based on the enhanced features includes: Predicting a binary image based on the enhanced features; A text detection area corresponding to the target image is generated based on the binary image.
8. A text detection device, characterized in that: The device comprises: An acquisition module, used to acquire a target image containing text; An extraction module, used for extracting features of the target image; a processing module, configured to perform self-attention processing on the features to obtain attention features; A determination module is used to determine an enhanced feature based on a segmentation feature and the attention feature; wherein the segmentation feature is obtained by semantically segmenting the text in the target image; and to determine a text detection area corresponding to the target image based on the enhanced feature.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
11. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.