An image alignment method applicable to text images
By extracting the field features in the text image and calculating the similarity, positioning synonyms and performing precise pairing, the error problem caused by distortion in the shooting environment during image alignment in the prior art is solved, and higher alignment accuracy and controllability are achieved.
Patent Information
- Application Number
- CN202111170598.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-08
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2041-10-08
AI Technical Summary
In the prior art, due to distortion and distortion of the shooting environment during image alignment, characteristic points pairing errors are difficult to achieve ideal alignment effects.
Through field feature extraction, including text box location, content and neighborhood information, the similarity between the field features in the template diagram and the image to be aligned is calculated, synonymous fields are positioned, and precise matching position alignment and matching points are performed to achieve image alignment.
It improves the accuracy and controllability of image alignment, and can maintain a good alignment effect in the presence of differences in shooting environment and distortion.
Smart Images

Figure CN113947678B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and specifically provides an image alignment method applicable to text images. Background Art
[0002] With the popularization of information technology, digital office has become inevitable, and the advantages of digital information such as convenience, sharing, and fast retrieval are becoming more and more prominent. In daily production work, a large number of bills, documents, etc. are accumulated, including a large amount of picture data. Effectively extracting the content of these picture data automatically, structuring it, and archiving it in the database has become the demand of the industry.
[0003] Currently, the extraction of image content with specific formats such as bills is mostly processed based on templates and optical character recognition (OCR technology). This method relies on accurate image alignment technology, that is, aligning the corresponding positions of the image to be parsed with the template image. Traditional alignment methods are mostly based on feature points. In actual applications, images taken by mobile phones are affected by the shooting environment and have problems such as distortion and warping, resulting in errors in paired feature points and making it difficult to obtain an ideal alignment effect. Summary of the Invention
[0004] In view of the above-mentioned deficiencies of the prior art, the present invention provides a practical image alignment method applicable to text images.
[0005] The technical solution adopted by the present invention to solve its technical problems is as follows:
[0006] An image alignment method applicable to text images. First, field feature extraction, respectively extract the field features in the template image and the image to be aligned. Secondly, synonymous field alignment, calculate the similarity between pairwise field features in the template image and the image to be aligned, locate the same-name and same-meaning fields in the template image and the image to be aligned, and obtain paired field pairs. Finally, precise paired position alignment and paired point optimization are performed to complete image alignment.
[0007] Further, in the field feature extraction, it further includes:
[0008] S101. Extract the relative position of the field detection frame on the image as the position feature;
[0009] S102. Extract the text content in the field as the content feature;
[0010] S103. Extract the number and content of text boxes in the field neighborhood as the neighborhood feature.
[0011] Further, after the construction of the image position feature, content feature, and neighborhood feature is completed, the field feature of the image is denoted as: F = {f1, f2,..., f n}, fn Represents the feature of the first field in the image, f n ={text pos , text rec , text nerb}, obtain the template image and the field features to be aligned, denoted as: f temp and f eval .
[0012] Furthermore, in step S101, the position feature of the text box, denoted as text pos , is obtained by the text detection algorithm. Through the text detection algorithm, the coordinates of the text bounding boxes of each field in the image are obtained;
[0013] Convert the bounding box coordinates to relative positions, divide the image into four regions, upper left, upper right, lower right, lower left, denoted as [1, 0, 0, 0], [0, 1, 0, 0], [0, 0, 1, 0], [0, 0, 0, 1] respectively. The relative position indicates the position of the current coordinate box in the image.
[0014] Furthermore, in step S102, the content feature of the text box, denoted as text rec , is obtained by the text recognition algorithm, and its content is the text recognition result in the text box.
[0015] Furthermore, in step S103, the neighborhood information, denoted as text nerb , calculates the number of text boxes in the neighborhood of the current text box and their text information. The neighborhood is defined as the number of pixel points between two field text boxes.
[0016] Furthermore, in the alignment of synonymous fields, it further includes:
[0017] S201. Calculate the content matching degree of f temp and f eval , take the text rec feature. The content matching degree is the number of overlapping characters of the text rec feature in the template image and the image to be aligned / the number of characters of the template image text rec ;
[0018] If the similarity is greater than the set threshold, proceed to the next step, indicating that this threshold can control the error introduced by the character recognition algorithm; if the similarity is equal to 1, directly return the paired field pair.
[0019] S202. For the field pairs that meet the threshold, take the text pos feature and calculate the position similarity. The similarity metric space uses the Euclidean distance. If the similarity is greater than the set threshold, proceed to the next step; if the similarity is less than the set threshold, skip this field pair;
[0020] S203. For the field pairs that meet the threshold, take the text nerb feature and calculate the neighborhood similarity. The calculation method of the neighborhood similarity is: the number of repeated fields / the total number of fields in the template neighborhood;
[0021] If the similarity is greater than the set threshold, the pairing is successful and the field pair is recorded; if the similarity is less than the set threshold, the field pair is skipped;
[0022] Furthermore, in the precise pairing position alignment and pairing point optimization, it further includes:
[0023] For the obtained field pairs, perform character segmentation to obtain the center points of individual characters. Among them, the character position information is denoted as char pos , char pos Store the information in the form of key-value pairs;
[0024] Character segmentation can be obtained by a character-level text detection algorithm. The character-level text detection model outputs the character position information in the form of a heat map. Here, through binarization and threshold segmentation, the minimum bounding rectangle of an individual character is finally obtained. Based on the coordinates of the minimum bounding rectangle, cropping is performed and input into the character recognition model to obtain its character content, which is used as the key of char pos The four coordinates of the bounding rectangle are further converted into the center point coordinates, which are used as the value of char pos value.
[0025] Furthermore, calculate the longest matching sequence of the paired field sequences in the field pairs, calculate the center points of the characters in the longest matching sequence respectively, and use their average value as the final pairing point position, and finally obtain multiple groups of paired coordinate pairs.
[0026] Furthermore, optimize the coordinate pairs. The optimization criteria are:
[0027] Traverse the coordinate pairs. Take 4 coordinate pairs as a group and use them as the vertices of a quadrilateral to construct a quadrilateral. The quadrilateral with the largest area formed is the optimal 4 pairing points. Thus, 4 optimal coordinate pairs are obtained;
[0028] Calculate the transformation matrix based on the obtained 4 optimal coordinate pairs to complete the perspective transformation of the image to be aligned and realize image alignment.
[0029] Compared with the prior art, an image alignment method applicable to text images of the present invention has the following outstanding beneficial effects:
[0030] The present invention extracts key points based on character features. Compared with traditional SIFT features, its dimensions are richer and more meaningful. The feature pairing process designed based on character meanings is also more accurate and controllable. The shooting environment of the image is less restricted. Even when there are shooting environment differences and distortions between the template image and the image to be aligned, good accuracy can still be maintained. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0032] FIG Figure 1 is a schematic flowchart of an image alignment method applicable to text images;
[0033] FIG Figure 2 is an example template image of an image alignment method applicable to text images;
[0034] FIG Figure 3 is an example image to be aligned of an image alignment method applicable to text images. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0035] In order to enable those skilled in the art of the present technology to better understand the solution of the present invention, the following further details the present invention in conjunction with specific embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0036] The following gives a best embodiment:
[0037] As Figures 1 - 3 shown, an image alignment method applicable to text images in this embodiment first performs field feature extraction, extracts field features in the template image and the image to be aligned respectively. Each field feature includes: text box content feature, text box coordinate feature, and text box neighborhood position feature.
[0038] Secondly, synonymous field alignment is performed. Calculate the similarity between pairwise field features in the template image and the image to be aligned, locate the same-name and same-meaning fields in the template image and the image to be aligned, and obtain paired field pairs;
[0039] Finally, the precise pairing position alignment and the optimization of pairing points are carried out, and then the image alignment is completed. Through character segmentation and the method of the largest area, the optimal 4 pairing points are determined, and then the transformation matrix is calculated to complete the perspective transformation, and finally the image alignment is realized.
[0040] The specific approach is as follows:
[0041] The first part: Feature extraction of fields in the template image and the image to be aligned
[0042] Define fields: various headings contained in the bill itself, such as fields like the invoicing date, invoice code, etc.
[0043] Define the template image: The image only contains each named field. Taking the invoice as an example, see Figure 1 ;
[0044] Define the image to be aligned, which needs to be transformed into a picture consistent with the template Figure 1 and its text content includes: each named field and the corresponding field information. See Figure 2 .
[0045] (1) Input the template image and the image to be aligned
[0046] (2) Respectively construct the field features of the template image and the image to be aligned, where the field features include 3 dimensions: the position feature, content feature, and neighborhood feature of the text box. The following takes the "invoice code" field in the template image as an example for each feature dimension:
[0047] a. The position feature of the text box, denoted as text pos : Obtained by text detection algorithms, such as algorithms like DB, yolo, etc. Through the text detection algorithm, the coordinates of the text bounding boxes of each field in the image are obtained. Here, the bounding box coordinates are further converted into relative positions, and the image is divided into four regions: upper left, upper right, lower right, and lower left, denoted as [1, 0, 0, 0], [0, 1, 0, 0], [0, 0, 1, 0], [0, 0, 0, 1] respectively. The relative position indicates the position of the current coordinate box in the image.
[0048] Taking 'invoice code' as an example, text pos = [0 1 0 0], indicating that the relative position of the invoice code is in the upper right part of the image.
[0049] Optionally, for the template image, the user can choose to use the text box after manual verification to replace the text calculated by the deep learning model in the above steps. pos .
[0050] b. The content feature of the text box, denoted as text rec , obtained by algorithms such as CRNN, and its content is the text recognition result in the text box.
[0051] Example: text of the 'Invoice Code' field rec = [Invoice Code]
[0052] Optionally, for the template image, the user can choose to replace the calculation result of the deep learning model in the step with the text information verified by themselves.
[0053] c. Neighborhood information, denoted as text nerb : Calculate the number of text boxes and their text information within the neighborhood of the current text box. The neighborhood is defined as the number of pixel points between two field text boxes. The example is as follows:
[0054] text nerb = {0: Invoice Number}, which means that for the 'Invoice Code' field, there is 1 text box in its neighborhood, and the content of the text box is 'Invoice Number'.
[0055] (3) Complete the construction of the field features of the image. Then the field features of the image are denoted as: F = {f1, f2,..., f n},f n represents the feature of the first field in the image, f n = {text pos , text rec , text nerb}.
[0056] So far, the template image and the field features to be aligned are obtained, denoted as: f temp and f eval .
[0057] Part Two: Alignment of Synonymous Fields between the Template Image and the Image to be Aligned
[0058] Here, synonymous fields are defined as: fields with the same meaning in the template image and the image to be aligned. For example, the synonymous field of the taxpayer identification number field of the seller in the template image is the taxpayer identification number field of the seller in the image to be aligned.
[0059] The alignment of synonymous fields mainly realizes the alignment of the 'Invoice Code' field in the image with the 'Invoice Code' field in the image to be aligned, rather than aligning to the 'Invoice Number'.
[0060] (1) Calculate the content matching degree of f temp and f eval , and take the text rec feature. The content matching degree is the number of overlapping characters of the text rec feature in the template image and the image to be aligned / the number of characters of the text rec in the template image.
[0061] If the similarity is greater than the set threshold, proceed to the next step. The threshold can control the error introduced by the character recognition algorithm. If the similarity is equal to 1, the successfully matched field pair is directly returned.
[0062] (2) Further, for the field pairs that meet the threshold, take the text pos Features, calculate position similarity, and the similarity measurement space uses Euclidean distance.
[0063] If the similarity is greater than the set threshold, proceed to the next step. The check here is to prevent errors introduced by fields with the same name but different meanings, such as the seller's 'Taxpayer Identification Number' and the buyer's 'Taxpayer Identification Number'. If position verification is not performed, it may lead to mismatching.
[0064] (3) For the field pairs that meet the threshold, take text nerb Features, calculate neighborhood similarity, the calculation method of neighborhood similarity is: number of repeated fields / total number of template neighborhood fields
[0065] If the similarity is greater than the set threshold, the pairing is successful and the field pair is recorded. The verification here is to further prevent errors introduced by similar fields due to character recognition errors, such as the mismatch caused by "invoice code" being mistakenly recognized as "invoice code" and "invoice number".
[0066] Part 3: Accurately locate and select matching points of the same-meaning field pairs in the template image and the image to be aligned
[0067] (1) For the field pairs obtained in the second part, character segmentation is performed to obtain the center point of a single character. The character position information is recorded as char pos , char pos Information is stored in the form of key-value pairs. Character segmentation can be obtained by character-level text detection algorithms (such as the craft model). Character-level text detection models generally output character position information in the form of heat maps. Here, through binarization and threshold segmentation, the minimum bounding rectangle of a single character can be obtained. Based on the coordinates of the minimum bounding rectangle, it is cropped and input into the character recognition model to obtain its character content as char pos The key converts the four-point coordinates of the circumscribed rectangle into the center point coordinates as char pos The value of .
[0068] Take the 'Invoice Code' field in the template diagram as an example.
[0069] char pos ={F:[cent x , cent y ], Tickets: [cent x , cent y, generation: [cent x , cent y , code: [cent x , cent y , where cent x , cent y represents the center point coordinates of the corresponding character.
[0070] (2) Calculate the longest matching sequence of the paired field sequences, calculate the center points of the characters in the longest matching sequence respectively, and use their average value as the final paired point position, and finally obtain multiple groups of paired coordinate pairs.
[0071] (3) Optimize the coordinate pairs. The optimization criterion is: traverse all the coordinate pairs, take 4 coordinate pairs as a group, as the vertices of a quadrilateral, construct a quadrilateral, and the quadrilateral with the largest area formed is the optimal 4 paired points. Thus, 4 optimal coordinate pairs are obtained.
[0072] (4) Based on the coordinate pairs in step (3), calculate the transformation matrix, complete the perspective transformation of the image to be aligned, and realize image alignment.
[0073] The above specific implementation manners are only specific cases of the present invention. The patent protection scope of the present invention includes but is not limited to the above specific implementation manners. Any implementation that conforms to the claims of an image alignment method for text images of the present invention and any appropriate changes or substitutions made by those of ordinary skill in the art in any of the technical fields shall fall within the patent protection scope of the present invention.
[0074] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirits of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. An image alignment method applicable to text images, characterized in that, First, field feature extraction is performed to extract the field features in the template image and the image to be aligned respectively. Secondly, synonymous field alignment is carried out to calculate the similarity between pairwise field features in the template image and the image to be aligned, locate the fields with the same name and the same meaning in the template image and the image to be aligned, and obtain paired field pairs. Finally, precise paired position alignment and paired point optimization are carried out to complete image alignment; In the field feature extraction, it further includes: S101: Extract the relative position of the field detection box on the image as the position feature; The position feature of the text box, denoted as , is obtained by the text detection algorithm. Through the text detection algorithm, the coordinates of the text bounding boxes of each field in the image are obtained; Furthermore, convert the bounding box coordinates to relative positions, divide the image into four regions: upper left, upper right, lower right, and lower left, denoted as [1,0,0,0], [0,1,0,0], [0,0,1,0], [0,0,0,1] respectively. The relative position indicates the position of the current coordinate box in the image; S102: Extract the text content in the field as the content feature; The content feature of the text box, denoted as , is obtained by the text recognition algorithm, and its content is the text recognition result in the text box; S103: Extract the number and content of text boxes in the field neighborhood as the neighborhood feature; Neighborhood information, denoted as , calculate the number of text boxes and their text information within the neighborhood of the current text box. The neighborhood is defined as the number of pixels between two field text boxes; After the construction of the image position feature, content feature, and domain feature is completed, the field feature of the image is recorded as: , represents the feature of the first field in the image, , obtain the template image and the field feature to be aligned, and record them as: and ; In the synonymous field alignment, it further includes: S201. Calculate and content matching degree, and take features. The content matching degree is the number of overlapping characters of the features in the template image and the image to be aligned / the number of characters of the template image ; If the similarity is greater than the set threshold, proceed to the next step, indicating that this threshold can control the error introduced by the character recognition algorithm; if the similarity is equal to 1, directly return the successfully paired field pairs; S202. For the field pairs that meet the threshold, take features and calculate the location similarity. The similarity metric space uses the Euclidean distance. If the similarity is greater than the set threshold, proceed to the next step; if the similarity is less than the set threshold, skip this field pair. S203. For the field pairs that meet the threshold, take features and calculate the neighborhood similarity. The calculation method of the neighborhood similarity is: the number of repeated fields / the total number of fields in the template neighborhood; If the similarity is greater than the set threshold, the pairing is successful and the field pair is recorded; if the similarity is less than the set threshold, skip the field pair; In the precise paired position alignment and paired point optimization, it further includes: For the obtained field pairs, perform character segmentation to obtain the center points of individual characters, where the character position information is denoted as , store the information in the form of key-value pairs; Character segmentation can be obtained by a character-level text detection algorithm. The character-level text detection model outputs character position information in the form of a heat map. Here, through binarization and threshold segmentation, the minimum bounding rectangle of a single character is finally obtained. Based on the coordinates of the minimum bounding rectangle, cropping is performed and input into the character recognition model to obtain its character content, which is used as the key, and the four-point coordinates of the bounding rectangle are further converted into the center point coordinates, which are used as the value; Calculate the longest matching sequence of the paired field sequences in the field pair, calculate the center points of the characters in the longest matching sequence respectively, and use their average value as the final paired point position, finally obtaining multiple groups of paired coordinate pairs; Optimize the coordinate pairs, and the optimization criteria are: Traverse the coordinate pairs, take 4 coordinate pairs as a group, as the vertices of a quadrilateral, construct a quadrilateral, and the quadrilateral with the largest area formed is the optimal 4 paired points. Thus, 4 optimal coordinate pairs are obtained; Calculate the transformation matrix based on the obtained 4 optimal coordinate pairs, complete the perspective transformation of the image to be aligned, and achieve image alignment.
Citation Information
Patent Citations
Image alignment method and device
CN106503634A
Image feature extraction method for pedestrian re-identification
CN107316031A
Image alignment method, device and equipment
CN110059711A
Field structured output method and device and computer readable storage medium
CN110738203A
Method for detecting characters in complex natural scene image
CN112418216A