Multi-stage deep learning map character erasing method applicable to 3D loading

Through a multi-stage deep learning algorithm combined with OCR and mask optimization technology, a text erase model suitable for tile maps was designed, which solved the problems of inefficient and poor results of map text erase in the prior art, and achieved efficient and accurate text removal and background completion.

CN119991720AInactive Publication Date: 2025-05-13TIANJIN RICHSOFT ELECTRIC POWER INFORMATION TECH +1
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510459831.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art has problems of inefficiency and poor results in map text erasing, especially in high-scale scenarios and complex background images. The text mask generated by traditional deep learning algorithms is not accurate enough, and it is easy to leave unnatural textures or color deviations.

Method used

A multi-stage deep learning algorithm is used, combined with optical character recognition (OCR) technology and mask optimization technology, a set of text erasing models suitable for tile maps are designed. The model is processed through multiple steps, including map text area detection, refined text mask generation, map background completion and text erasing, achieving efficient text removal and background completion.

Benefits of technology

The efficiency and accuracy of text erasing are significantly improved, and the average structural similarity (MSSIM) of the generated wordless map and the real map reaches 0.973, which is better than the traditional method and is suitable for batch processing of large-scale images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991720A_ABST
    Figure CN119991720A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-stage deep learning map character erasing method applicable to 3D loading, which comprises the following steps of: detecting a map text region: outputting an initial rectangular text mask based on an optical character recognition (OCR) method; a refined text mask is generated, wherein the preliminarily generated rectangular text mask is optimized, and a refined mask attached to the character edge is output; map background completion: based on an image completion pre-training model, taking a word-free map and a text mask as input, and obtaining a final model through confrontation process learning of a generator and a discriminator; character erasing: performing mask mapping on the map with characters through the generated refined text mask to generate a text region missing image; and outputting a character-free map of which the text area background is complemented in the map image complementation model so as to erase map characters. According to the method, a multi-stage deep learning algorithm combination mode of a text image processing scene is adopted, and an area needing character removal is efficiently erased.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of digital image processing, and in particular to a map text erasing method applicable to 3D loading through multi-stage deep learning. Background Art

[0002] With the rapid development of three-dimensional GIS visualization, maps, as important carriers of geographic information, and text annotations play a vital role in helping users understand map content and obtain geographic location information. However, in some application scenarios, such as map aesthetics requirements, privacy protection of specific areas, or base map editing requirements, it is necessary to remove text annotations on the map. Existing processing technologies are mainly achieved by hiding or deleting the text layer of the map. However, for the sake of data security and commercial interests, some map service providers only provide tile maps with text, and do not provide unlabeled base map data. For example, in actual project applications of power systems, when Cesium is used to call Siji Map to realize three-dimensional GIS visualization, only map services with tile maps with text can be obtained. The text erasing method of map tile maps in the prior art has the following shortcomings:

[0003] 1. Using image editing software to manually erase text is not only inefficient, but also almost infeasible when large-scale data processing is required at high map zoom levels;

[0004] 2. Using simple image processing technology, the text detection and erasure effect is poor for complex background images such as maps;

[0005] 3. With the help of artificial intelligence technology, when using traditional deep learning algorithms for processing, it is usually a single-stage model based on an encoder-decoder network. The generated text masks are often not accurate enough. After erasing the text, it is easy to leave unnatural textures or color deviations on the map, affecting the overall visual effect of the map. Summary of the invention

[0006] The present invention aims to solve at least one of the technical problems existing in the prior art. To this end, one purpose of the present invention is to propose a multi-stage deep learning map text erasing method applicable to 3D loading, which is based on artificial intelligence technology, adopts a multi-stage deep learning algorithm combination method for text image processing scenarios, and uses optical character recognition (hereinafter referred to as OCR) technology and mask optimization technology, combined with image completion deep learning network algorithm to develop a set of text erasing models suitable for tile maps, which can efficiently erase the area where the text needs to be removed, thereby improving the erasing efficiency and map background completion accuracy.

[0007] To this end, the present invention proposes a multi-stage deep learning map text erasing method applicable to 3D loading, comprising the following steps:

[0008] S1. Dataset creation: Collect map tiles with text for text region detection and text mask generation; Collect map tiles without text as training and validation sets for the text erasure model;

[0009] S2. Map text area detection: Based on the optical character recognition method OCR, use the text detection pre-trained model to locate the text and obtain the minimum bounding rectangle of the text area outline, and output the initial rectangular text mask;

[0010] S3. Refined text mask generation: Optimize the initially generated rectangular text mask, combine the text morphology, and output a refined mask that fits the text edge;

[0011] S4. Map background completion: Based on the image completion pre-training model, the wordless map and text mask are used as input, and the final model is learned through the adversarial process of the generator and the discriminator, so that the completion result conforms to the texture and color characteristics of the map image;

[0012] S5. Text erasure: Mask mapping is performed on the map with words through the generated refined text mask to generate an image of the missing text area; in the map image completion model, a wordless map with background completion of the text area is output to erase the map text.

[0013] Preferably, in S2, when detecting map text areas, the segmentation-based text detection algorithm DBNet is first used to detect text lines, and the text probability threshold and text area probability threshold requirements are reduced.

[0014] Preferably, S3 includes the following sub-steps:

[0015] S3.1. Binarize the text map to generate a text mask map;

[0016] S3.2. Perform a morphological dilation operation on the initial rectangular text mask to make the mask completely cover the potential area of ​​the text edge;

[0017] S3.3. Performing a superposition operation on the expanded text mask map and the tile map to generate a local map of only the mask area map;

[0018] S3.4. Calculate the pixel mean of the map in the area within the mask and fill it with the background value of the area outside the mask to weaken the background interference and enhance the saliency of the text outline in the mask;

[0019] S3.5. Perform grayscale processing and edge detection. Based on the text edge detection results, perform a closing operation of first dilation and then erosion, and finally generate an irregular mask that highly matches the actual shape of the text.

[0020] Preferably, in S4, a LaMa-based pre-training model is used for customized training to learn the complex texture and structural information of the image during the training process.

[0021] Preferably, in S5, a batch of map tiles with text to be erased are first input, and the above-mentioned text area detection process is used to generate a text mask map; then the image missing area is determined by mask optimization and mapping, and the above-mentioned generated map image completion model is input to obtain a batch reasoning output of a complete wordless tile map.

[0022] The advantages of the present invention compared with the prior art are:

[0023] The multi-stage deep learning method of the present invention is applicable to a 3D loaded map text erasure method, which solves the text erasure problem for tile map application scenarios without independent text layers; it uses artificial intelligence technology to automatically detect and locate text areas through a deep learning algorithm to achieve text removal, without relying on the original text layer information of the map, thereby significantly improving the level of intelligence.

[0024] The present invention innovatively designs and optimizes the text mask acquisition method, abandons the traditional rectangular mask, makes full use of the background information of the text area itself, combines multi-process image processing technology, and generates an irregular mask that fits closely to the edge of the text, providing more precise guidance for the subsequent image completion process of the text mask removal area, thereby improving the accuracy of image content missing completion and greatly improving the text erasure effect.

[0025] A multi-stage deep learning algorithm combination is adopted, and each stage performs more refined processing, which improves the overall processing ability of complex background images; using deep learning technology capabilities, batch processing of large-scale images is achieved in the model reasoning stage, which significantly improves efficiency. It is especially suitable for actual application scenarios that need to process a large number of map tiles. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or technical descriptions will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0027] Figure 1 It is the overall flow chart of the method of the present invention;

[0028] Figure 2 This is a flowchart of text mask optimization of the present invention. DETAILED DESCRIPTION

[0029] The embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and cannot be understood as limiting the present application.

[0030] In the description of this application, it should be noted that, unless otherwise clearly specified and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium, it can be the internal connection of two elements or the interaction relationship between two elements. For ordinary technicians in this field, the specific meanings of the above terms in this application can be understood according to specific circumstances.

[0031] The present invention will be further described in detail below with reference to the accompanying drawings. Figure 1~Figure 2 In order to solve the above problems, the present invention provides a multi-stage deep learning method for erasing map text applicable to 3D loading:

[0032] like Figure 1 As shown, it is a schematic diagram of the method flow of this embodiment, which mainly includes two stages: map text erasure model training and text erasure method application. It is specifically divided into four parts: map tile sample set production, text detection and mask generation, text erasure model training and text erasure method application. Each implementation part contains a unique technical algorithm. The specific process principle steps involved are as follows:

[0033] (1) Sample set preparation stage: collect tile maps with text and corresponding tile maps without text. In this embodiment, a total of 6442 tile maps were collected, and the data set was randomly divided into a training set and a validation set in an 8:2 ratio.

[0034] (2) In the text detection and mask generation stage, based on the tile map samples with words corresponding to the validation set, a refined irregular text mask map is generated to guide parameter optimization during the text erasure model training process. This step is the core innovation of the present invention. Based on the text region detection results, through multi-step morphological operations, edge optimization and region merging, the refined generation of irregular text masks is achieved, which significantly improves the fit between the mask and the actual shape of the text. The core algorithm steps involved are as follows: Figure 1 The specific implementation process of the ① and ② marked in the figure is as follows: Figure 2 .

[0035] (3) In the text erasure model training phase, the training set and validation set generated above are used to train the image completion model after the text area is removed. The key strategies and model effect verification are as follows:

[0036] Normally, the model training process requires the data set to input the real original image and target image (such as a single-stage text erasure model, which uses the image with text as the initial input and the image without text as the training target. First, the image without text after the text is erased is inferred through model training, and then the similarity is compared with the real original image without text to measure the learning ability of the model); however, the present invention fully considers the increased difficulty of model training caused by the complexity of the map image background, and finally adopts a three-stage deep learning model of map text area detection + text mask optimization generation + map background completion. At this time, the model is trained by inputting a real tile map without text and its corresponding text mask map. Since the LaMa algorithm itself has a random mask generation strategy, the training set only needs to input the tile map without text, and the verification set needs to input the tile map without text and the text mask map. During the training process, the model completes the map effect of the text mask area to debug and learn the optimal model parameters.

[0037] Important parameter settings: Use Adam optimizer, generator learning rate is set to 0.0005, discriminator learning rate is set to 0.0001, batch_size is set to 15, and train for 50 rounds.

[0038] Model effect verification: 645 tile maps were randomly selected to evaluate the text erasure effect of this embodiment. The experimental results show that the average structural similarity (MSSIM for short) between the word-free map generated by the method of this embodiment and the real map is 0.973, which is higher than the EraseNet method (a single-stage deep learning text erasure method, the MSSIM value of the experimental sample is 0.936) and the text erasure method under rectangular mask (a two-stage text erasure method of text detection + image completion, the MSSIM value of the experimental sample is 0.959).

[0039] (4) Text erasure model application stage. The specific application process is as follows:

[0040] ① Input any tile map with text (256*256);

[0041] ② Based on DBNet, the text area detection is performed, and the threshold thresh for determining whether a pixel is a text is set to 0.2, and the text area probability threshold box_thresh is lowered to 0.1 to improve the text missed detection rate;

[0042] ③ Generate refined text masks based on the mask optimization model;

[0043] ④ Perform text mask mapping on the tile map with text, and output the tile map with missing text area;

[0044] ⑤ Input the tile map after text mask mapping into the LaMa model to complete the map background of the text area and achieve seamless erasure of map text.

[0045] According to the above model application process, the present invention is actually applied to Siji Map for verification, and the text erasing effect is good, which is an example effect.

[0046] The concept of "MSSIM" mentioned above is defined as: mean structural similarity, that is, average structural similarity, which is an indicator that measures image similarity from three aspects: brightness, contrast, and structure. Its value range is between 0 and 1. The closer the value is to 1, the higher the similarity between the reconstructed image and the original image, and the better the image quality.

[0047] The innovation of the above step (2) of the invention for optimizing and generating text masks based on the text region detection results is described below. Figure 2 As shown:

[0048] 1. First, input a tile map containing text information (such as area names, landmark labels, etc.), use the DBNet-based text line detection model to detect all text lines on the tile map, and calculate and output the text minimum enclosing rectangular area Box for each text line. Several text areas can be detected in each image, and a preliminary list of rectangular text area candidate boxes is generated;

[0049] 2. Process a single text area independently to ensure that subsequent operations are fine-tuned for a single text. Considering that small lines, symbols and other elements may be misdetected as text due to the complex map background, if the area of ​​the text area is less than 50 (the dimension size of the DBNet pre-processed image is 800*800), it will be removed to achieve noise filtering, otherwise it will be retained;

[0050] 3. Generate a binary image based on the tile map pixel size (256*256), where the text area pixels are marked as 255 (white) and the non-text area pixels are marked as 0 (black), forming a rectangular text mask image that covers the rough position of the text;

[0051] 4. Perform a morphological dilation operation on the initial rectangular text mask, using a 5*5 convolution kernel to expand the mask boundary to ensure that the mask completely covers the potential area at the edge of the text;

[0052] 5. Superimpose the expanded text mask map with the tile map to generate a local map containing only the mask area map;

[0053] 6. Local map background optimization: Calculate the pixel mean of the map in the area within the mask and fill it with the background value of the area outside the mask to weaken background interference and improve the significance of the text outline in the mask, so as to improve the accuracy of subsequent text edge detection;

[0054] 7. Grayscale processing and edge detection: grayscale the global image, input the Canny edge detection algorithm, use the Sobel operator (kernel size is 3*3) to calculate the gradient, mark the pixels with gradient intensity greater than 200 as strong edges, and mark the pixels with gradient intensity between 150 and 200 as weak edges. Keep the strong edge pixels and the weak edge pixels adjacent to the strong edge pixels as the final detected edge results;

[0055] 8. Irregular mask generation based on text contour: Based on the text edge detection results, a closing operation of dilation followed by erosion is performed. First, a large-size 7*7 convolution kernel is used to dilate the text edge to achieve complete coverage of the text contour, and then a small-size 3*3 convolution kernel is used to erode the text contour to eliminate redundant fine line edges and extract the fine contour of the text. Finally, an irregular mask that highly matches the actual shape of the text is generated;

[0056] 9. Multi-text area mask merging: Traverse all optimized single text masks, merge them, and generate a text mask map with the same pixel size as the tile map.

[0057] The above is the production process of the map text erasure model of the present invention, which integrates text detection technology, mask optimization technology and image completion technology, and obtains an efficient and accurate text erasure model through multi-stage training optimization.

[0058] Finally, all the parts not described in the present invention adopt mature products and mature technical means in the prior art.

[0059] It should be understood that those skilled in the art can make improvements or changes based on the above description, and all these improvements and changes should fall within the scope of protection of the appended claims of the present invention.

Claims

1. A multi-stage deep learning method for erasing map text applicable to 3D loading, characterized in that: The following steps are involved: S1. Dataset creation: Collect map tiles with text for text region detection and text mask generation; Collect map tiles without text as training and validation sets for the text erasure model; S2. Map text area detection: Based on the optical character recognition method OCR, use the text detection pre-trained model to locate the text and obtain the minimum bounding rectangle of the text area outline, and output the initial rectangular text mask; S3. Refined text mask generation: Optimize the initially generated rectangular text mask, combine the text morphology, and output a refined mask that fits the text edge; S4. Map background completion: Based on the image completion pre-trained model, the wordless map and text mask are used as input, and the final model is learned through the adversarial process of the generator and the discriminator, so that the completion result conforms to the texture and color characteristics of the map image; S5. Text erasure: Mask mapping is performed on the map with words through the generated refined text mask to generate an image of the missing text area; in the map image completion model, a wordless map with background completion of the text area is output to erase the map text.

2. According to claim 1, a multi-stage deep learning method for erasing map text applicable to 3D loading, characterized in that: In S2, when detecting map text areas, the segmentation-based text detection algorithm DBNet must first be used to detect text lines, and the text probability threshold and text area probability threshold requirements must be lowered.

3. According to claim 1, a multi-stage deep learning method for erasing map text applicable to 3D loading, characterized in that: The S3 includes the following sub-steps: S3.

1. Binarize the text map to generate a text mask map; S3.

2. Perform a morphological dilation operation on the initial rectangular text mask to make the mask completely cover the potential area of ​​the text edge; S3.

3. Performing a superposition operation on the expanded text mask map and the tile map to generate a local map of only the mask area map; S3.

4. Calculate the pixel mean of the map in the area within the mask and fill it with the background value of the area outside the mask to weaken the background interference and enhance the saliency of the text outline in the mask; S3.

5. Perform grayscale processing and edge detection. Based on the text edge detection results, perform a closing operation of first dilation and then erosion, and finally generate an irregular mask that highly matches the actual shape of the text.

4. According to claim 1, a multi-stage deep learning method for erasing map text applicable to 3D loading, characterized in that: In S4, a LaMa-based pre-training model is required to perform customized training so as to learn the complex texture and structural information of the image during the training process.

5. According to claim 1, a multi-stage deep learning method for erasing map text applicable to 3D loading, characterized in that: In S5, a batch of map tiles with text to be erased are first input, and the text area detection process is used to generate a text mask map; then the image missing area is determined by mask optimization and mapping, and the map image completion model generated above is input to obtain a batch reasoning output of a complete wordless tile map.

Citation Information

Patent Citations

  • Universal detection algorithm for optical printing text and scene text

    CN112861794A

  • Character detection and recognition method and device, electronic equipment and storage medium

    CN113920295A

  • OCR training sample generation method, device and system

    CN114419632A

  • Method and device for processing text in image, electronic equipment and readable storage medium

    CN116740227A

  • Medical record character recognition method and system based on deep learning

    CN117218672A