Method and device for rubbing image number identification and region matching

By using neural network model detection and character filtering and splicing, the numbered area is obscured and the content area of ​​the rubbing is extracted. This solves the problems of low OCR accuracy and inaccurate region segmentation in rubbing image processing, and achieves efficient matching of the number and the content area of ​​the rubbing, thus improving the accuracy and stability of extraction.

CN120976955APending Publication Date: 2025-11-18HUAZHONG UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510902518.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-06-09
Filing Date
2025-07-01
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing rubbing image processing technologies suffer from low OCR accuracy, inaccurate segmentation of rubbing content areas, and low efficiency of manual matching, lacking an end-to-end solution.

Method used

A preset neural network model is used to detect rubbing images, perform character filtering and splicing, cover the numbered area and extract the content area of ​​the rubbing, and achieve matching between the number and the content area of ​​the rubbing.

Benefits of technology

It improves the accuracy and stability of rubbing outline extraction, replaces manual comparison, achieves efficient one-to-one matching, removes interference factors, and improves the accuracy of extracting the content area of ​​the rubbing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120976955A_ABST
    Figure CN120976955A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image recognition and artificial intelligence, in particular to a rubbing image number recognition and region matching method and device, and the method comprises the steps: detecting a rubbing image through a preset neural network model, and obtaining a detection result; performing character screening and splicing processing on the detection result to obtain number information; performing shielding processing on the identified number on the copy of the rubbing image according to the number information, and reserving a rubbing content area to obtain a number shielding image; performing rubbing content area detection and extraction on the numbered occlusion image to obtain a rubbing content area, and matching the numbered information with the rubbing content area to output a matched image; one-to-one matching of numbers and rubbing content areas can be efficiently completed, and manual comparison is replaced; a serial number area shielding strategy is adopted, interference factors are removed before the rubbing content area is extracted, the serial number is prevented from being mistakenly recognized as a part of the rubbing, and the rubbing contour extraction accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image recognition and artificial intelligence, and particularly relates to a method and device for rubbing image number recognition and region matching. BACKGROUND

[0002] In the prior art, there are many challenges in digitizing and processing rubbing images. First, rubbing images usually have numbers (such as serial numbers or front and back marks) attached for identification, but the positions and formats of these numbers are not fixed. The number characters can appear on the edges, corners or other blank areas of the rubbing image, and the font sizes and styles are different, and sometimes contain Chinese characters (such as "positive" and "negative" for identifying the front or back). Some rubbings are old, and the number area may be faded, damaged, or disturbed by uneven lighting and noise during shooting. These factors result in low accuracy of traditional optical character recognition (OCR) technology in processing rubbing numbers: image noise, uneven ink and complex background can interfere with the stability of the text recognition algorithm.

[0003] Existing solutions often treat OCR recognition and image region segmentation as different steps, and lack integration, resulting in process breaks: the text information obtained by OCR cannot be directly used to locate the corresponding image region, and the regions obtained by image segmentation cannot be automatically associated with the numbers. This fragmented processing method cannot meet the application requirements of batch and structured archiving of rubbing images. For example, a museum may want to automatically process hundreds of rubbing photos, crop the rubbing content in each image and label its number, and the existing technology lacks an end-to-end solution to complete this task.

[0004] Therefore, it is urgent to overcome the defects of the prior art in the technical field. SUMMARY

[0005] The technical problem to be solved by the present application is how to solve the problems of low OCR accuracy, inaccurate rubbing content region segmentation and low efficiency of manual matching in the existing rubbing image processing technology.

[0006] The present application adopts the following technical solutions: In a first aspect, a method for rubbing image number recognition and region matching is provided, comprising: detecting a rubbing image by a preset neural network model to obtain a detection result; performing character screening and splicing processing on the detection result to obtain number information; performing occlusion processing on the recognized number according to the number information and retaining the rubbing content region on a copy of the rubbing image to obtain a number occlusion image; performing detection and extraction of the content region of the expanded sheet on the numbered occlusion image to obtain a content region of the expanded sheet, and matching the numbering information with the content region of the expanded sheet to output a matching image.

[0007] Preferably, the detection on the expanded sheet image by the preset neural network model to obtain a detection result specifically comprises: automatically locating all text regions possibly containing text in the expanded sheet image, and outputting position coordinates in the form of a polygon or a rectangular frame; performing text content recognition on each detected text region, and outputting a corresponding numbering string and recognition confidence to obtain the detection result.

[0008] Preferably, the character screening and splicing processing on the detection result to obtain numbering information specifically comprises: applying a preset regular expression or format rule to filter irrelevant text of non-numbering type according to the detection result, and only keeping numbering strings conforming to a specific pattern; for the same numbering segmented into multiple independent parts in the expanded sheet image, splicing and fusing adjacent characters according to their spatial distribution to reconstruct a complete numbering identifier, so as to obtain the numbering information.

[0009] Preferably, the occlusion processing on the recognized numbering according to the numbering information and the reservation of the content region of the expanded sheet on the copy of the expanded sheet image to obtain a numbered occlusion image specifically comprises: traversing all effective numbering region corresponding rectangular boundary frames in the copy of the expanded sheet image according to the numbering information, and performing occlusion block filling operation on the rectangular boundary frames one by one to obtain the numbered occlusion image. Preferably, the detection and extraction of the content region of the expanded sheet on the numbered occlusion image to obtain a content region of the expanded sheet specifically comprises: performing a closing operation on the numbered occlusion image by a preset size of a structural element; performing external contour extraction on the numbered occlusion image after the closing operation by a contour detection algorithm, and screening the extracted external contour set according to a preset screening rule to obtain the content region of the expanded sheet.

[0010] Preferably, the preset screening rule comprises: calculating the area of each external contour and the aspect ratio of its circumscribed rectangle; if the area of the external contour is less than a set threshold, or the aspect ratio of the circumscribed rectangle of the external contour deviates from a preset reasonable range, the external contour is removed; If the circumscribed rectangle of a region of an outer contour is completely contained in the circumscribed rectangle of a region of another outer contour, only the region corresponding to the outer contour of the outer layer is retained.

[0011] Preferably, the matching of the numbering information with the content region of the map sheet to output a matching image specifically comprises: The center coordinates of each numbered region and each content region of the map sheet are calculated, and their relative positions are compared. If the center point of the number is within the horizontal range of the content region of the map sheet, and the vertical distance between the bottom edge of the numbered region and the top edge of the content region of the map sheet is within a preset tolerance threshold, it is determined that the two belong to a corresponding relationship, and a matching link is established to output a matching image.

[0012] Preferably, the method further comprises: for the unpaired content region of the map sheet, calculating the intersection-over-union ratio between it and the paired content region of the map sheet, and if the corresponding intersection-over-union ratio exceeds a preset threshold, merging processing is performed, and a number identification is automatically assigned.

[0013] In a second aspect, a device for map sheet image numbering recognition and region matching is provided, which comprises a processor and a memory for storing processor-executable instructions. The processor is configured to execute the method for map sheet image numbering recognition and region matching.

[0014] In a third aspect, a non-volatile computer storage medium is provided, which stores computer executable instructions, which are executed by one or more processors to complete the method for map sheet image numbering recognition and region matching of the first aspect.

[0015] In a fourth aspect, a chip is provided, which comprises a processor and an interface for calling and running a computer program stored in a memory to execute the method for map sheet image numbering recognition and region matching of the first aspect.

[0016] In a fifth aspect, a computer program product comprising instructions is provided, which, when executed on a computer or processor, causes the computer or processor to execute the method for map sheet image numbering recognition and region matching of the first aspect to the fourth aspect and any one thereof.

[0017] Compared with the prior art, the present application has the following beneficial effects: The application detects the decoction piece image by a preset neural network model to obtain a detection result; then performs character screening and splicing processing on the detection result to obtain numbering information; then performs occlusion processing on the recognized number according to the numbering information on a copy of the decoction piece image and retains a decoction piece content area to obtain a number-occluded image; finally, performs detection and extraction of the decoction piece content area on the number-occluded image to obtain the decoction piece content area, and matches the numbering information with the decoction piece content area to output a matched image; the one-to-one matching of the number and the decoction piece content area can be efficiently completed, replacing manual comparison, on the other hand, the number area occlusion strategy is adopted, the interference factor is removed before the decoction piece content area is extracted, the number is avoided from being misrecognized as part of the decoction piece, and the accuracy and stability of the decoction piece contour extraction are greatly improved. BRIEF DESCRIPTION OF DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or the prior art description. Obviously, the drawings in the following description only constitute some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor based on these drawings.

[0019] Figure 1 is a flowchart of a method for decoction piece image number recognition and area matching provided by an embodiment of the present application; Figure 2 is a flowchart of a method for obtaining a detection result provided by an embodiment of the present application; Figure 3 is a flowchart of a method for obtaining numbering information provided by an embodiment of the present application; Figure 4 is a flowchart of a method for obtaining a decoction piece content area provided by an embodiment of the present application; Figure 5 is a schematic diagram of a segmentation completion provided by an embodiment of the present application; Figure 6 is a schematic diagram of a decoction piece area and a number area before matching provided by an embodiment of the present application; Figure 7 is a schematic diagram after successful matching provided by an embodiment of the present application; Figure 8 is a structural schematic diagram of a device for decoction piece image number recognition and area matching provided by an embodiment of the present application. DETAILED DESCRIPTION

[0020] In order to make the objects, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and should not be used to limit the present application.

[0021] Unless otherwise required by context, the term "comprises" or "comprising" as used in this specification is taken to mean the inclusion since but not limited to. In the description of the specification, the terms "one embodiment", "some embodiments", "exemplary embodiment", "example", "specific example" or "some examples" are intended to indicate that the described implementation, implementation or example is included in at least one embodiment or example of the present disclosure. The illustrative representation of the above terms does not necessarily mean the same embodiment or example. In addition, the specific features, structures, materials or characteristics described can be included in any one or more embodiments or examples in any appropriate manner, that is, although they are carried in the embodiment or example of the above terms due to the order of appearance and location, they are not limited to the combination of one embodiment or example.

[0022] In the description of the present application, the terms "first", "second" are only used for description purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, the features limited by "first", "second" can be explicitly or implicitly included in one or more features. In the description of the embodiments of the present disclosure, unless otherwise stated, the meaning of "multiple" is two or more. In addition, for example, in the description, the same type of nouns can also be described as two independent individuals by adding "A", "B" at the end, in which case the features limited by "A", "B" are only used for the purpose of distinguishing the same type of individual description, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated.

[0023] In describing some embodiments, "coupled", "coupled" and "connected" and their derivatives can be used. For example, the term "connected" can be used to describe some embodiments to indicate that two or more components have direct physical or electrical contact with each other. For example, the term "coupled" can be used to describe some embodiments to indicate that two or more components have direct physical or electrical contact. However, the term "connected" or "coupled" can also refer to two or more components that do not have direct contact with each other, but still cooperate or interact with each other, such as "optical coupling", "wireless connection", etc. The embodiments disclosed herein are not necessarily limited to the content of the present application.

[0024] In addition, the technical features involved in each embodiment of the application described below can be combined with each other as long as there is no conflict.

[0025] Embodiment 1: To solve the problems of the prior art, the embodiment provides a method for number recognition and region matching of a rubbing image, which comprises the following steps: detecting a rubbing image by using a preset neural network model to obtain a detection result; performing character screening and splicing processing on the detection result to obtain number information; performing occlusion processing on the recognized number on a copy of the rubbing image according to the number information and reserving a rubbing content region to obtain a number-occluded image; performing detection and extraction of the rubbing content region on the number-occluded image to obtain a rubbing content region, and matching the number information with the rubbing content region to output a matched image.

[0026] The embodiment detects a rubbing image by using a preset neural network model to obtain a detection result; then performs character screening and splicing processing on the detection result to obtain number information; then performs occlusion processing on the recognized number on a copy of the rubbing image according to the number information and reserving a rubbing content region to obtain a number-occluded image; finally, performs detection and extraction of the rubbing content region on the number-occluded image to obtain a rubbing content region, and matches the number information with the rubbing content region to output a matched image; the one-to-one matching of the number and the rubbing content region can be efficiently completed, and manual comparison is replaced. On the other hand, the number region occlusion strategy is adopted, the interference factors are removed before the rubbing content region is extracted, the number is avoided from being misrecognized as part of the rubbing, and the accuracy and stability of the rubbing contour extraction are greatly improved.

[0027] Before the method for number recognition and region matching of a rubbing image is executed, a series of preprocessing operations are performed on the input rubbing image, including grayscale, denoising, contrast enhancement, adaptive threshold segmentation and the like. By converting a color image into a grayscale image, the calculation complexity is reduced, high-frequency noise is suppressed by using a smoothing filter (such as Gaussian blur), and local details are enhanced by using adaptive histogram equalization with limited contrast, and finally the image is binarized by using an adaptive threshold such as the Otsu algorithm, the image clarity and edge quality are improved, and a stable foundation is provided for subsequent text detection and region analysis.

[0028] In one embodiment, the user imports the original map image into the system, which can be in a common format such as JPEG (Joint Photographic Experts Group) or PNG (Portable Network Graphics), and the image contains the content of the map and the number label beside it. First, the system performs pre-processing operations on the input image to improve image quality and reliability of subsequent analysis. The pre-processing process includes: grayscale conversion (converts the color map image to a grayscale image, reduces the data dimension while retaining the main brightness information. This step can reduce the computational complexity of subsequent processing and remove the interference of color on text recognition), noise filtering (uses a smoothing filter method such as Gaussian blur to denoise the grayscale image, to weaken the influence of high-frequency details such as paper texture or scanning noise on edge detection. For example, use a 5x5 Gaussian kernel to convolve and smooth the image to remove isolated black points or ink noise), contrast enhancement (for the problem of uneven illumination or low contrast of the map image, apply an adaptive histogram equalization algorithm to enhance the local contrast of the image. This method limits the contrast gain of histogram equalization to avoid excessive enhancement of noise, while making the distinction between map text and background more obvious and highlighting edge and character details), threshold segmentation (binarize the enhanced grayscale image, preferably using an adaptive threshold algorithm such as Otsu's method to automatically calculate the optimal segmentation threshold according to the image grayscale histogram, and divide the pixels into foreground (ink color) and background (paper). The binarized result makes the map text and numbers appear as clear black connected regions, while the background is white, which helps subsequent edge extraction and contour analysis).

[0029] In one embodiment, first, a color image in original RGB format is read, and fixed interference areas (such as scales, labels, text boxes, etc.) that may affect the recognition accuracy in the image are identified, and a solid color pixel value (for example, white) is used for shielding processing. Specifically, in the image coordinate system, the position of the rectangular region is selected and covered with a solid white rectangle to generate the initial image after shielding. Then, the color image after shielding is converted into a grayscale image, the image dimension is reduced, and the subsequent processing process is simplified, representing the preprocessing operation of the grayscale module. Next, Gaussian blur processing is applied to the grayscale image, and a 5x5 size Gaussian kernel function can be used to smooth the image to suppress high-frequency random noise and small texture interference in the image, representing the filtering operation of the image denoising module; the filtered image can effectively improve the stability and edge consistency of the subsequent binarization result. In order to further enhance the local contrast and detail expression ability of the image, an adaptive histogram equalization method with limited contrast is used to enhance the image. The image is divided into small grid regions by the adaptive histogram equalization method, and histogram equalization is performed in each region to prevent over-enhancement and maintain the natural feeling of the image, representing the function of the local contrast enhancement module. Finally, the enhanced grayscale image is subjected to Otsu adaptive threshold segmentation method again, and the image is automatically converted into a black and white binary image, so that the text region and the background region have good contrast and separation effect, representing the output of the image binarization module. The above image preprocessing steps provide a clear and structurally stable image basis for subsequent text recognition and numbering region extraction, which helps to improve the overall recognition accuracy and the robustness of region segmentation of the system.

[0030] After the above preprocessing, an optimized binary image (i.e., the plate image used in the subsequent method) is output. At this time, the noise of the plate image has been significantly reduced, and the edges of the text and patterns have been strengthened, laying a good foundation for the next step of text detection and region segmentation.

[0031] In one embodiment, as shown in Figure 1 , the specific process of the plate image numbering recognition and region matching method includes: Step 101: detecting a plate image through a preset neural network model to obtain a detection result.

[0032] Among them, the numbering region in the plate image can be automatically detected by using a deep learning Optical Character Recognition (OCR) model, and the text content and position coordinates thereof are extracted to obtain the detection result. The OCR model integrates a text detection network and a text recognition network.

[0033] In one embodiment, as shown in Figure 2 , the step 101 specifically includes: Step 1011: Automatically locate all possible text regions in the map image, and output the position coordinates in the form of a polygon or rectangular frame.

[0034] Among them, the text detection network can be constructed based on a deep convolutional neural network, and a differentiable binarization technology is combined to accurately locate the text block region in the image. The text detection network can adapt to the variable and directionless numbered characters in the map image, and realize the detection of text in any direction. The output is a number of rectangular bounding boxes, each of which locates a possible numbered character region.

[0035] Step 1012: Perform character content recognition on each detected text region, and output the corresponding numbered character string and recognition confidence to obtain the detection result.

[0036] Among them, the text recognition network performs character content recognition on each detected text region, and outputs the corresponding numbered character string and recognition confidence, representing the double-channel detection result of the text detection and recognition module.

[0037] In one embodiment, the text recognition network recognizes each numbered region obtained by detection to obtain the actual character content of the number. The text recognition network adopts an architecture combining convolutional neural network and sequence modeling. For example, a convolutional neural network (Convolutional Neural Networks, abbreviated as CNN) can be used to extract the visual feature sequence of the character image, then a bidirectional long short-term memory network or a visual Transformer is used to model the feature sequence, and finally a connectionist temporal classification algorithm is used to decode and output the character sequence. This process does not require character-by-character segmentation of the image, can directly output the numbered character string, and has robustness for continuous character recognition.

[0038] In one embodiment, considering the actual situation that the number plate may appear tilt or rotation, an angle classification module is enabled in the OCR model to determine the direction of each detected text region and automatically correct the angle if necessary to improve the recognition accuracy and stability, representing the function of the angle adaptive recognition module. After obtaining the OCR output, the system performs effectiveness screening on the recognition results. Specifically, the system traverses all the text items returned by the OCR and uses a preset regular expression to determine whether they conform to the number format rules. The number format includes: a string composed of only Arabic numerals (such as "123"), or a number form with "positive", "negative", "molar" and other text labels in its suffix (such as "45 positive", "67 negative", "12 molar"). For text items that do not conform to the above structure, the system marks them as invalid numbers and eliminates them, representing the output strategy of the number legality screening module. During the screening process, the spatial position corresponding to each valid number text is recorded simultaneously, including the top-left corner coordinates (x, y) of the rectangular box, the width w, the height h, and the corresponding number string text, constituting a structured number candidate region set. This set not only contains the text content, but also contains the physical positioning information of the number in the image, providing key data support for the character fusion, region occlusion and number plate pairing operations in the subsequent steps.

[0039] In one embodiment, for the rotation or tilt of the number plate in the image (for example, the number is sometimes handwritten on the side of the paper, with a certain angle), an angle classification module is introduced in the OCR process. This module determines the text direction according to the image features of the number region. If a significant rotation angle is detected, the system will correct the rotation of the region (such as rotating the image by the corresponding angle and then sending it into the recognition network) to ensure that the recognition network always works in the horizontal direction of the text, thereby improving the recognition accuracy and stability.

[0040] Through the above-mentioned OCR model, the system can obtain the detection results of each number region, including: number text string (such as "12" or "37 negative", etc.), number region position (coordinates of the bounding box), and recognition confidence, etc. If the recognition confidence of some number regions is low, the system can also mark them for subsequent manual verification, but generally the OCR results will be directly used for subsequent processing.

[0041] Step 102: performing character screening and splicing processing on the detection results to obtain number information.

[0042] In one embodiment, as shown in Figure 3 , the step 102 specifically includes: Step 1021: filtering non-number irrelevant text according to the detection results using a preset regular expression or format rule, and only keeping the number string that conforms to the specific pattern.

[0043] Among them, the application of pre-set regular expression or format rule filters irrelevant text of non-number class, only retains the number string conforming to the specific mode, such as pure number format (such as "123"), format with positive and negative signs (such as "45 positive", "67 negative" and the like). Such as full number string or combination with "positive", "negative", "mortar" and the like suffix, indicates the function of the text legality screening module.

[0044] In one embodiment, the pre-defined rule or regular expression is first applied to screen the identified text list. Only the content entries that may be the number of the map sheet are retained, which are typically characterized by a string of numbers or numbers plus a small number of Chinese characters. For example, in this embodiment, it is set to retain the string like "number + optional Chinese character", such as pure number "123", "007" or number "45 positive", "45 negative" with "positive", "negative" and the like Chinese characters at the end. Like Chinese characters in the text of the map sheet, disorganized character combinations and noise strings that may be misrecognized by the OCR model will be filtered out and not enter the subsequent processing.

[0045] Step 1022: For the same number that is segmented into multiple independent parts in the map sheet image, the adjacent characters are spliced and fused according to the spatial distribution to reconstruct the complete number identification to obtain the number information.

[0046] Among them, for those in the image that are segmented into multiple independent parts (for example, due to the gap or slight misalignment between characters, which is recognized by the OCR as multiple segments), the adjacent characters are spliced and fused according to the spatial distribution to reconstruct the complete number identification.

[0047] In one embodiment, the geometric threshold condition of the fusion judgment includes that the horizontal distance between adjacent characters is less than twice the average width of the characters, and the vertical alignment deviation is not more than 20% of the height of the characters. The character block that meets the above conditions will be considered as belonging to the same number and be merged. Through this screening and fusion step, the integrity and accuracy of the number recognition result can be significantly improved, and the number of each map sheet can be correctly extracted as a standard format.

[0048] In one embodiment, the spatial fusion processing is performed on the multiple numbered character blocks screened by the screening. The system determines whether the numbered characters belong to the same number according to their spatial positions. The determination criteria include: the alignment error of adjacent characters in the vertical direction is not more than 20% of the height of the character, and the spacing in the horizontal direction is less than twice the average width of the character, indicating the character geometric fusion condition; the system sorts the character blocks in the order from top to bottom and from left to right, and determines whether the current character and the adjacent character on the right side satisfy the fusion condition, if yes, the text is spliced, the right boundary and the bottom boundary of the merged bounding box are updated, and the fused character is marked as processed and no longer participates in the subsequent merging operation. After the fusion is completed, each merged number region is redefined as an independent number object, which includes a unified rectangular bounding box and a complete number string. This processing method can effectively solve the problem of number segmentation caused by large character spacing, position offset or OCR recognition error, and the merging output result of the character fusion processing module; the system records the unified bounding box coordinates (x, y, w, h) and the corresponding number string text for each fused number, and uses the result as the key basic data for subsequent region occlusion and map region pairing, indicating the data structure of the number fusion result; finally, the number list after number fusion is output as the effective output (i.e. number information) of the number region detection and extraction stage, which is used for the subsequent step 103 of occlusion processing.

[0049] For the numbered characters screened by the format, some may be separated from each other in the image and recognized as independent frames by the OCR, and it is necessary to further determine whether they should be combined into a complete number. The system determines whether to splice and fuse adjacent character fragments according to the coordinate information and arrangement rules of the numbered characters. For example, when two recognition frames are very close in the horizontal direction and located on the same line, or slightly misaligned but almost aligned in the vertical direction, they are very likely to belong to the same part of the number. The specific determination conditions are as described above: the horizontal spacing between adjacent characters is less than twice the average width of the character, and the vertical center line deviation is not more than 20% of the height of the character. Two or more character blocks that meet these conditions will be spliced into a number in the order of position. For example, if the OCR outputs two adjacent frame contents as "4" and "5", and they are very close in the horizontal direction and well aligned in the vertical direction, the system will merge them into the number "45". After this step, a standardized number list is output, each number corresponding to a determined text string and position boundary. This fusion process greatly improves the completeness of number extraction and avoids the situation that the number is cut due to improper character segmentation.

[0050] Step 103: performing occlusion processing on the recognized numbers according to the number information and retaining the map content area on the copy of the map image to obtain a number occlusion image.

[0051] In one embodiment, the step 103 specifically comprises traversing all the rectangular bounding boxes corresponding to the valid number regions in the copy of the image of the map according to the number information, and performing the occlusion block filling operation on the rectangular bounding boxes one by one to obtain the number-occluded image.

[0052] In one embodiment, the step 103 specifically comprises traversing all the rectangular bounding boxes corresponding to the valid number regions in the copy of the image of the map according to the number information, and performing the occlusion block filling operation on the rectangular bounding boxes one by one to obtain the number-occluded image.

[0053] In one embodiment, according to the number information output by the step 102, the rectangular bounding box information of all valid number regions is traversed, including the upper left corner coordinates (x, y) and the width and height (w, h), and a black filling operation is performed on these regions one by one on the binary copy of the image, that is, a solid black rectangular block is drawn in each number region, so that the region is completely covered in the image. The occlusion operation is implemented by the OpenCV (Open Source Computer Vision Library) drawing function, which sets the pixel value to 0 (black) and fills the entire number box range, ensuring that the number content no longer participates in the map edge detection and connected region analysis. The occlusion operation is only performed on the image analysis copy, and the original image is not affected, so that the purpose of structuring and occluding the number region is achieved while the complete image is retained for result display and output. The occluded image (i.e., the number-occluded image) retains all the structural information of the map body, while significantly reducing the misidentified non-map regions, improving the accuracy and integrity of subsequent map contour detection, and providing a stable and interference-free basic image environment for the independent extraction of the map region, indicating the execution effect and purpose of the number region occlusion module.

[0054] Step 104: performing detection and extraction of the map content region on the number-occluded image to obtain the map content region, and matching the number information with the map content region to output a matching image.

[0055] In one embodiment, as shown in Figure 4 and Figure 5 performing detection and extraction of the map content region on the number-occluded image to obtain the map content region, specifically comprises: Step 1041: performing a close operation on the numbered occlusion image by a preset size of a structural element.

[0056] In one embodiment, the structural element used is a 10x10 square kernel, and the close operation sequence is first dilation and then erosion. The purpose is to connect the broken parts in the edge of the patch that may be caused by occlusion, image degradation or character covering, so that the originally discontinuous patch contour is reconstructed and closed, thereby improving the detectability of the overall contour, representing the processing logic of the morphological contour enhancement module.

[0057] In one embodiment, the close operation consists of one dilation operation and one erosion operation, which has the effect of filling small holes in the foreground (black pixel) region of the image and connecting adjacent foreground objects. For patch images, the close operation helps to fill small cracks or breaks that may exist in the edge of the patch content. For example, due to the number occlusion, the edge of the patch may be truncated into several segments, as shown in FIG. 10B, where the patch region corresponding to the number "4154" in the right image can be deepened and repaired by the close operation, making the patch region a continuous whole. This actually plays a role in edge preservation and repair, ensuring the integrity of the patch contour. Figure 5

[0058] Step 1042: performing external contour extraction on the numbered occlusion image after the close operation by a contour detection algorithm, and screening the extracted external contour set according to a preset screening rule to obtain the patch content region.

[0059] In one embodiment, the contour detection algorithm in OpenCV is called, and the cv2.findContours function is used to perform external contour extraction on the image after the close operation, and all closed connected regions are obtained. The system calculates the area and the size of the bounding box of the external rectangle of each contour region extracted, and screens the candidate regions according to the preset geometric rules, including that the area must be greater than 7000 pixels to filter small noise interference regions in the image.

[0060] For patch images, the close operation helps to fill small cracks or breaks that may exist in the edge of the patch content. For example, due to the number occlusion, the edge of the patch may be truncated into several segments, and the close operation can reconnect these edges, making the patch region a continuous whole. This actually plays a role in edge preservation and repair, ensuring the integrity of the patch contour.

[0061] ​The preset screening rule includes: calculating the area of each external contour and the aspect ratio of the circumscribed rectangle thereof; if the area of the external contour is less than a set threshold, or the aspect ratio of the circumscribed rectangle of the external contour deviates from a preset reasonable range, the external contour is removed; if the circumscribed rectangle of the region of one external contour is completely contained in the circumscribed rectangle of the region of another external contour, only the region corresponding to the outer external contour is retained.

[0062] In an embodiment, the area (i.e. the number of pixels) of each external contour and the aspect ratio of the circumscribed rectangle thereof are first calculated. If the area of a region is less than a set threshold (which can be empirically set according to the image resolution or the size of the map sheet, for example, 1% of the total pixels or a minimum number of pixels), the region is considered to be too small and is likely to be a noise point or non-map content, and is removed. Then the aspect ratio is checked. Map content usually exhibits certain rectangular characteristics (may be vertical or horizontal, but extremely elongated is unlikely to be a complete map sheet), so if the aspect ratio of the circumscribed rectangle of a contour deviates significantly from a reasonable range (for example, > 10:1 or < 1:10, etc.), it is also excluded.

[0063] After screening the candidate regions that meet the basic shape condition, further nesting detection is performed to determine whether the current candidate rectangular frame is completely contained in other larger regions. If it is an embedded region, it is marked as a redundant region and is excluded to avoid mistakenly identifying detail regions or corner noise inside the map sheet as independent map sheet blocks. Finally, the set of candidate regions that remain are the main regions of the map sheet identified in the image, each region is described by its boundary coordinates (x, y, w, h), and is used as the map content region for subsequent numbering matching.

[0064] In an embodiment, as shown in Figure 6 , each numbering frame (corresponding to the numbering information) and the adjacent map region frame (corresponding to the map content region) need to be matched. As shown in Figure 7 , the matching is successful. The matching of the numbering information and the map content region to output a matching image specifically includes: calculating the center coordinates of each numbering region and each map content region in the numbering information, and comparing their relative positions; if the center point of the numbering is within the horizontal range of the map content region, and the vertical distance between the bottom edge of the numbering region and the top edge of the map content region is within a preset tolerance threshold, it is determined that the two belong to a corresponding relationship, and a matching link is established to output a matching image.

[0065] The one-to-one matching association between the numbering regions and the content regions of the map is automatically completed based on the spatial geometric relationship between the numbering regions and the content regions of the map in the numbering information. The embodiment proposes a matching strategy taking the center point alignment and the vertical distance as the criteria. The matching strategy includes calculating the center coordinates of each numbering region and each content region of the map, and then comparing the relative positions thereof. If the center point of a certain numbering is within the horizontal range (left and right boundaries) of a certain content region of the map, and the vertical distance between the bottom edge of the numbering region and the top edge of the content region of the map is within the preset tolerance threshold, it is determined that the two belong to the corresponding relationship, and a matching link is established.

[0066] The matching condition can be described in the formula form. Let the center coordinates of the numbering region be (x1, x2), the center coordinates of the content region of the map be (x2, y2), and the allowed deviation thresholds in the horizontal and vertical directions be Δx and Δy, respectively. When |x1-x2|<Δx and |y1-y2|<Δy are satisfied, it is considered that the numbering and the content region of the map are matched. Through this spatial matching rule, the system can accurately associate the numbering label in the image with the corresponding content region of the map, and realize the one-to-one corresponding identification of the map and the numbering. In the matching process, if an abnormal situation that a content region of the map may contain multiple numberings (which should not theoretically occur, but if an OCR splits one numbering into two parts and does not fuse them, it can be checked again at this stage) or one numbering corresponds to multiple regions occurs, the system can also select the nearest pair according to the center point distance to match, so as to maintain the reasonable uniqueness of the pairing relationship.

[0067] In an embodiment, the matching process further includes: for each content region of the map, sequentially checking all the numbering regions, calculating the horizontal distance between the center horizontal coordinate of the numbering frame and the center horizontal coordinate of the content frame of the map, and if the center point of the numbering is within the left and right boundary range of the content region of the map, and the vertical distance between the bottom edge coordinate of the numbering region and the top edge of the content region of the map satisfies the set threshold condition (for example, between -20 and 0 pixels), it is considered that the numbering and the content have an effective spatial corresponding relationship, and it is determined that they are a pair of legal matches. After the matching is successful, the system merges the numbering region and the content region of the map into a new image cropping block, takes the minimum value of the boundary coordinates of the two as the top left corner and the maximum value as the bottom right corner, crops the image segment composed of the numbering and the content, and records the numbering string text and the cropping region coordinate information together, representing the output of the numbering and content pairing module. The system adopts a greedy matching strategy, preferentially selects the nearest numbering and content that satisfy the conditions for pairing, and ensures that each numbering and each content can be paired at most once, avoiding repeated matching and ambiguity conflicts; meanwhile, the matched numbering index and the content index are marked, which are excluded in subsequent operations. The pairing mechanism effectively guarantees the logical consistency and physical corresponding relationship between the content regions and the numbering regions of the map, and lays a foundation for the processing of the unpaired regions and the overall management of the numbering regions.

[0068] In one embodiment, in the matching process, if it is found that a number corresponds to multiple regions or a region corresponds to multiple numbers, the matching degree should be distinguished: generally, each number corresponds to only one region, and each region also generally corresponds to only one number. Therefore, a matching score can be calculated according to the vertical distance and horizontal coincidence degree, and the pair with the highest score is selected as the effective match to avoid the occurrence of one-to-many or many-to-one. In addition, if there is a number missing in the image (for example, there is no number in the content of a region), the region will not be able to find a matching number at this step and will be temporarily marked as "unpaired". The output of the matching step is a list of established number-region pairs and a list of possibly unpaired regions.

[0069] After the above steps, most of the content regions of the region should have corresponding number matches. However, for a small number of unpaired regions (possibly due to missing numbers, recognition failure or special formats that do not meet the matching conditions), this embodiment introduces an overlap analysis mechanism to process these regions. In one embodiment, the method for number recognition and region matching of the region image further comprises calculating the intersection over union (IOU) between the unpaired region and the paired region, and if the corresponding IOU exceeds a predetermined threshold, merging and automatically assigning a number identifier.

[0070] Specifically, the intersection over union (IOU) between each unpaired region and any paired region is calculated. The definition of IOU is: IOU(A, B) = Area(A∩B) / Area(A∪B), A and B are two regions, Area(A∩B) is the area of their overlapping part, and Area(A∪B) is the area of their union. When the IOU value of a certain unpaired region and any paired region exceeds a predetermined threshold (for example, 0.4), it can be judged that the two regions are actually adjacent or overlapping fragments on the content of the region, and should be considered as part of the same region. At this time, the system merges these regions: taking the union of their bounding boxes as a new complete region, and considering the merged region as a region. For such a merged region, since there is no independent number, the system automatically generates a new number identifier for it to record and distinguish. For example, it can be named "unnumbered 1", "unnumbered 2", etc. as a marker for the case where the content of the region has no corresponding number. The merging strategy of unpaired regions improves the fault tolerance of the system to special cases, ensuring that each region of the content has a corresponding number or is reasonably marked, and nothing is missed.

[0071] In one embodiment, the present embodiment further processes these special regions by overlap analysis. The present embodiment considers two main cases: (a) Numbering missing: the content region of the map sheet itself is complete, but for various reasons, it does not appear on the image as a recognizable number (possibly the number is missing or not written); (b) Region fragmentation: the content of a map sheet is detected by the algorithm as two or more independent regions, resulting in only part of the region matching the number, while the other regions do not match.

[0072] For case (a) described above, the processing method is to generate a temporary numbering identifier for this region, such as a series named "unnumbered" with a serial number, so as to account for the output results.

[0073] For case (b) described above, it is necessary to determine whether the unmatched region actually belongs to a part of a matched region, so as to be merged. The specific judgment method is to calculate the intersection over union (IoU) of the unmatched region and each matched region, as defined in the aforementioned formula. When the IOU of a certain unmatched region and a certain matched region is higher than a preset threshold (such as 0.4), it is considered that the two regions are actually one whole (possibly due to the blank in the middle of the map sheet, which divides the outline into two segments, or the image preprocessing causes slight segmentation fracture). At this time, the system will merge the unmatched region and the matched region: taking the union of their outlines as the new map sheet region, and updating the number corresponding to the region. If the matched region has a number, the existing number is used after merging; if the matched region also has no number (in theory, this case will not occur, because the matched region must have a number), or multiple unmatched regions overlap each other without any number, then a "unnumbered" identifier is assigned to the merged result. After merging, two regions that were originally separated may now be considered as one whole map sheet content, and correspond to one number (which may be automatically generated). This step ensures that the region list output will not leave obvious cases of belonging to the same map sheet but assigned different numbers or unnumbered, making the final number-map sheet correspondence more complete.

[0074] In one embodiment, after completing the matching of numbers and map sheet regions, the system generates the final output result, including image visualization and data files. First, the system draws annotation boxes and number markers on the original map sheet image: different colors are used to distinguish different types of regions, such as green boxes to mark the boundaries of detected map sheet content regions, red boxes to mark the positions of numbering regions, and purple boxes to mark the new map sheet regions obtained by merging, and the numbers (including automatically generated "unnumbered" identifiers) are annotated beside the corresponding boxes. This visualized pairing result image can intuitively show the correspondence between numbers and map sheet regions.

[0075] Subsequently, each pair of matched numbers and rubbings areas are cropped from the original image to generate independent image files (containing only the content of the rubbings) and named by the number or establish an index relationship between the number and the file name for future retrieval. At the same time, the structured data corresponding to the pairing result is output, such as JSON (English full name: JavaScript Object Notation) or CSV (English full name: Comma Separated Values) format table, recording the position, size and cropped image file name of each number and its corresponding rubbings area in each original rubbings image. These data will facilitate database storage, retrieval and interface with other digital management systems. At this point, the entire rubbings image number identification and area intelligent segmentation matching process is completed, and the results can be used for archiving and further analysis of cultural heritage digital resources.

[0076] In one embodiment, first, the system performs visual annotation processing on the original image, traverses all the number-rubbings area combinations that have been successfully paired or merged by IOU, and draws bounding boxes on the image to show the area position, while adding number label text next to the box to clarify the number correspondence. Different colors are used to distinguish area types in the annotated image, such as green boxes representing paired rubbings areas, red boxes representing number areas, and purple boxes representing unnumbered rubbings areas generated by merging, to facilitate users to quickly identify the pairing status. Subsequently, the system performs image cropping on each pair of number and rubbings area combination, extracts the corresponding image block from the original image according to its rectangular boundary, and generates a separate image file for saving. For image blocks with recognized valid numbers, the file naming rule is "number_original image file name", for example "123_image1.png"; for merged areas that lack numbers, the naming rule is "unnumbered_original image file name_sequence number", such as "unnumbered_image1_3.png", indicating the execution strategy of the output image naming module. To avoid the error of covering the same name file, the system performs a duplicate detection mechanism before saving the image file, and if the file already exists, it automatically adds a sequence number suffix to the file name to generate a unique path. All image blocks are saved to the specified output directory and stored by number type. In addition, the system also exports all the pairing information of numbers and rubbings areas as structured data files, supporting JSON or CSV format, and the data content includes original image name, number text, rubbings area coordinates (x, y, w, h), output image file path, etc., representing the output result of the number pairing structure data module. Through this step, the system realizes the visualization confirmation of the pairing result, the standard output of the rubbings area image and the data storage of the pairing information, which facilitates the integrated use and quick call in the subsequent digital cultural relic management, database retrieval, image archiving and other scenarios.

[0077] In one embodiment, after all the numbers and the scroll regions are determined, the system generates the final visualization results and output files for the user: In which, the result image is labeled: the system labels each scroll region and its number on the original scroll image, and highlights the matching relationship. As mentioned earlier, different types of bounding boxes use different colors for easy differentiation: green border represents the scroll content region, red border represents the number region; if there is a merged scroll region, use purple border. Each scroll region border will be labeled with the corresponding number text (also marked for automatically generated "unnumbered"). In this way, a labeled image can help users intuitively check the matching effect, and facilitate manual review of important information. For example, if a number position is incorrect or missing, it can be discovered in time by observing the labeled image.

[0078] Image cropping output: the system crops each scroll content region in the original image into an independent image file according to the matching relationship. The edges are appropriately expanded to completely contain the scroll content and remove excess white space. The file is named in association with the number, for example, named "number.jpg", or the file name and number correspondence is recorded in the result data. For "unnumbered" regions, the file can be named according to the generated identification (such as "unnumbered_1.jpg"). These cropped small images are convenient for later direct use, such as making a scroll catalog or storing the image of each scroll in the database.

[0079] Structured data output: the system outputs the matching results in table form, including the processing information of each original image. For example, each record includes: original image file name, number content, number region coordinates, scroll region coordinates (or size), cropped image file name, etc. The output format can be CSV text, JSON file or direct writing into the database. Such data is convenient for subsequent batch import into cultural relic management systems or for statistical analysis. For example, managers can quickly count the total number of scrolls, the number of scrolls missing numbers, or retrieve the image corresponding to a specific number.

[0080] The above only describes the preferred embodiments of the present application and is not intended to limit the present application. Any modifications, equivalent replacements and improvements made within the spirit and principle of the present application shall be included in the protection scope of the present application.

[0081] Embodiment 2: In embodiment 1, a method for scroll image number recognition and region matching is provided, and in this embodiment, a device for scroll image number recognition and region matching will be proposed, which comprises a processor and a memory for storing processor executable instructions; wherein the processor is configured to execute the method for scroll image number recognition and region matching described in embodiment 1.

[0082] As Figure 8 shown, the device for scroll image number recognition and region matching comprises a processor 21 and a memory 22, wherein the processor 21 and the memory 22 can be connected through a bus or other means.

[0083] The processor 21 can be a central processing unit (CPU). The processor 21 can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or a combination thereof.

[0084] The memory 22 is a non-transitory computer readable storage medium, which can be used to store non-transitory software programs, non-transitory computer executable programs and modules, such as the program instructions / modules corresponding to the method for scroll image number recognition and region matching in Embodiment 1 of the present application. The processor executes various functions of the processor and training processes by running the non-transitory software programs, instructions and modules stored in the memory.

[0085] The memory 22 can include a program storage area and a training storage area, wherein the program storage area can store an operating system and at least one application required by a function; and the training storage area can store training created by the processor. In addition, the memory can include a high-speed random access memory, and can also include a non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state memory device. In some embodiments, the memory 22 can optionally include a memory remotely arranged with respect to the processor, and these remote memories can be connected to the processor through a network. Examples of the network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof. The one or more modules are stored in the memory 22, and when executed by the processor 21, perform the method for scroll image number recognition and region matching in Embodiment 1 as shown in Figure 1 The specific details of the method for scroll image number recognition and region matching can be understood by referring to the corresponding descriptions and effects in Embodiments shown in Figure 1 , Figure 2 and Figure 3 , and will not be repeated here.

[0086] The embodiment also provides a computer storage medium, which stores a computer program executable by a processor to complete the method for image number identification and region matching of a tuocha image as described in the embodiment 1.

[0087] The computer storage medium stores computer executable instructions, which are executable for the method for image number identification and region matching of a tuocha image in any method embodiment. The storage medium can be a disk, an optical disk, a read-only memory (ROM), a random access memory (RAM), a flash memory, a hard disk drive (HDD) or a solid-state drive (SSD), etc. The storage medium can also include a combination of the above-mentioned types of memories.

[0088] The specific steps of the method for image number identification and region matching of a tuocha image are described in the embodiment 1, which will not be repeated here.

Claims

1. A method for number recognition and region matching of rubbing images, characterized in that, include: The detection results are obtained by detecting the rubbing image through a preset neural network model; The detection results are processed by character filtering and concatenation to obtain the number information; On a copy of the rubbing image, the identified number is masked according to the numbering information while the content area of ​​the rubbing is preserved to obtain a number-masked image; The numbered occluded image is subjected to detection and extraction of the rubbing content region to obtain the rubbing content region, and the number information is matched with the rubbing content region to output a matching image.

2. The method for number recognition and region matching of rubbing images according to claim 1, characterized in that, The method further includes: The system automatically locates all possible text regions in the rubbing image and outputs their position coordinates in the form of polygons or rectangles. For each detected text region, text content recognition is performed, and the corresponding number string and recognition confidence score are output to obtain the detection result.

3. The method for number recognition and region matching of rubbing images according to claim 1, characterized in that, The method further includes: Based on the detection results, preset regular expressions or formatting rules are applied to filter irrelevant text that is not numbered, and only numbered strings that conform to a specific pattern are retained; For the same number that is divided into multiple independent parts in the rubbing image, adjacent characters are spliced ​​and merged according to their spatial distribution to reconstruct the complete number identifier and obtain the number information.

4. The method for number recognition and region matching of rubbing images according to claim 1, characterized in that, The method further includes: Based on the numbering information, traverse all the rectangular bounding boxes corresponding to the valid numbered areas in the copy of the rubbing image, and perform occlusion block filling operations on each of the rectangular bounding boxes to obtain the numbered occluded image.

5. The method for number recognition and region matching of rubbing images according to claim 1, characterized in that, The method further includes: A closing operation is performed on the numbered occluded image using a structuring element of a preset size; The outer contour of the numbered occluded image after the closing operation is performed by a contour detection algorithm, and the extracted outer contour set is filtered according to a preset filtering rule to obtain the content area of ​​the rubbing.

6. The method for number recognition and region matching of rubbing images according to claim 5, characterized in that, The preset filtering rules include: Calculate the area of ​​each outer contour and the aspect ratio of its bounding rectangle; If the area of ​​the outer contour is less than a set threshold, or if the aspect ratio of the outer contour's bounding rectangle deviates from a preset reasonable range, it will be rejected. If the bounding rectangle of one outer contour region is completely contained by the bounding rectangle of another outer contour region, then only the region corresponding to the outer outer contour is retained.

7. The method for number recognition and region matching of rubbing images according to claim 1, characterized in that, The method further includes: Calculate the center coordinates of each numbered area and each rubbing content area in the numbering information, and compare their relative positions; If the center point of the number is within the horizontal range of the rubbing content area, and the vertical distance between the bottom edge of the number area and the top edge of the rubbing content area is within the preset tolerance threshold, then the two are determined to be in a corresponding relationship, and a matching link is established to output a matching image.

8. The method for number recognition and region matching of rubbing images according to claim 1, characterized in that, The method further includes: for unpaired rubbing content areas, calculating the intersection-union ratio between them and paired rubbing content areas; if the corresponding intersection-union ratio exceeds a preset threshold, merging is performed and a number is automatically assigned.

9. A device for identifying and matching the number of rubbings in images, characterized in that, The device for identifying and matching the number of rubbings images includes: a processor and a memory for storing processor-executable instructions; The processor is configured to execute the method for number recognition and region matching of rubbing images as described in any one of claims 1-8.

10. A non-volatile computer storage medium, characterized in that, The computer storage medium stores computer-executable instructions, which are executed by one or more processors to perform the method for number recognition and region matching of rubbing images as described in any one of claims 1-8.